Inspiration

Professional firms serving NRIs and Global Citizens face a strange automation problem. The work is repetitive, but the decisions are not. Residency, tax filing, remittances, FEMA questions and supporting evidence appear again and again across clients. Yet the correct answer depends on the client's facts, the applicable rule, effective dates, source quality and professional judgment. A normal chatbot can answer quickly. That is not enough.

The harder question was:

Can an AI agent do the repetitive work of a professional firm without quietly acquiring the authority of the professional?

That became Nicole.

Our design principle is simple: Reasoning can be probabilistic. Authority cannot. What it does? Nicole is a governed professional agent for NRI and cross-border tax advisory. She takes a case from intake through evidence retrieval, reasoning, bounded tool use and a reviewable recommendation.

The professional remains in control of consequential decisions.

Nicole can:

Convert an enquiry into structured case state. Retrieve governed and versioned knowledge. Reason over case facts using Strands Agents. Choose from a small set of bounded tools. Propose facts without confirming them. Ask for missing information instead of guessing. Produce a professional draft with its evidence basis. Record the tools, facts and knowledge used in the decision. Learn from documents and professional feedback through a governed learning loop. Reuse only knowledge that a professional has accepted.

Nicole cannot:

Approve her own work. Promote her own learning into verified institutional knowledge. Confirm a proposed client fact as true. bypass case and service scope. silently convert model output into professional authority.

The result is not another tax chatbot.

It is an operating model for professional AI where the agent can do more work without being given more authority than it should have.

How we built it

Nicole is built around four architectural layers.

Layer 1 is deterministic and controls what the model may know.

Inbound enquiries are persisted into case state. Routing, case scope, context assembly and governed retrieval happen before model reasoning.

Layer 2 is probabilistic.

Nicole runs as a Strands Agent with dynamic reasoning, hooks, skills, steering and four bounded tools:

get_case_context

get_service_knowledge

record_proposed_facts

propose_next_action

Layer 3 is deterministic again.

Applicability rules, evidence quality, authority boundaries and claim validation determine what the model may honestly propose.

A model can suggest a supported answer, but the server can still refuse it if the evidence does not permit that conclusion.

Layer 4 is human authority.

The professional reviews, edits, approves or rejects. Knowledge publication is also a professional act.

The current AWS deployment includes:

Amazon Bedrock AgentCore Runtime

Python 3.12 runtime in VPC mode

MMDSv2

AgentCore workload identity

AgentCore Identity

AWS Secrets Manager

Private Amazon RDS PostgreSQL

Versioned deployment artifacts

Least privilege execution roles

Nicole currently uses Groq through a provider abstraction for inference while remaining deployed on AgentCore Runtime.

The runtime database identity cannot approve work or publish knowledge. Those restrictions are enforced outside the prompt.

Challenges we ran into

The hardest problems were not model prompts.

They were authority, applicability, provenance and evidence.

One early test exposed a serious design flaw.

The system had a verified rule about resident individuals and a confirmed non-resident client. Every other condition looked correct.

The deterministic layer still permitted a supported answer.

The model was not the problem.

The governing system did not yet have a way to express whether the rule actually applied to the person.

That failure changed the architecture.

We introduced three-valued applicability:

TRUE

FALSE

UNKNOWN

FALSE excludes the rule.

UNKNOWN does not guess. It becomes a question for the client.

This was an important lesson: a deterministic system can still be confidently wrong if its vocabulary is too weak to represent the constraint that matters.

We also found several other false greens during the build.

A READY AgentCore Runtime did not prove a complete business workflow.

A private database did not automatically prove least privilege.

A successful model response did not prove safe authority boundaries.

Passing tests did not prove stochastic reliability.

A learning candidate did not mean verified knowledge.

We treated each of those as separate evidence classes instead of collapsing them into one claim of success.

Accomplishments that we are proud of

Nicole is deployed as a READY Amazon Bedrock AgentCore Runtime.

The runtime is attached to private subnets and connects to private Amazon RDS PostgreSQL.

Secrets are resolved at runtime instead of being embedded in the deployment environment.

We proved from inside the AgentCore runtime that the application database credential can be resolved and the private database can be reached.

We also proved that the same runtime database identity is refused when attempting privileged administrative access.

A live invocation against an unknown case fails closed with RUN_REFUSED instead of fabricating context.

The repository currently passes 345 checks from a clean disposable PostgreSQL build created entirely from repository files.

The governed learning loop is also implemented and tested.

A document can produce candidate knowledge.

The candidate is checked against captured evidence.

A professional can accept or reject it.

Accepted knowledge enters a versioned reusable release.

Rejected knowledge remains unavailable to future runs.

Most importantly, Nicole cannot promote her own learning into professional truth.

What we learned

The biggest lesson was that professional AI is not primarily a model selection problem.

It is an authority design problem.

Good agent architecture needs at least three different questions:

What can the model reason about?

What can the system claim?

Who is allowed to make the consequential decision?

Those answers should not come from the same component.

We also learned that human in the loop should not mean adding an approval button at the end.

Human authority needs to be structural.

The runtime should lack the capability to approve its own work.

The model should not be able to transform a proposed fact into a confirmed fact.

Learning should not become reusable merely because the model extracted it confidently.

We also learned to separate different kinds of proof:

Source code proof

Deterministic test proof

Artifact proof

AWS deployment proof

Runtime connectivity proof

Agent behavior proof

Repeated behavioral reliability

Those are not interchangeable.

That discipline changed how we built and how we describe the project.

What's next for Nicole

The current hackathon build establishes the governed foundation.

The next stage is to expand the same architecture without weakening its authority model.

The roadmap includes:

Repeated behavioral evaluation and pass at k reliability measurement. Prompt injection and cross-case adversarial acceptance tests. Richer OpenTelemetry and operational metrics. Stronger document provenance and multimodal evidence. Policy-aware external integrations. Multi-tenant professional firms. Additional cross-border domains beyond NRI tax. Specialized agents only where evaluation evidence shows they are necessary.

The long-term goal is a professional intelligence platform where firms can scale AI capability without scaling unchecked authority.

Nicole is built around one principle that should remain true even as the system becomes more capable:

Reasoning can be probabilistic. Authority cannot.

Built With

Share this project:

Updates

Submission history