Inspiration

An internationally trained nurse, engineer, physician or teacher has to clear a licensing sequence before they can work in their field: credential evaluation, language testing, registration, exams, and sometimes supervised experience. The order, the prerequisites and the validity windows differ by profession and by jurisdiction. Caseworkers at settlement agencies and nonprofits guide people through this largely by hand, one client at a time.

The failure we kept coming back to is timing. A document can be valid when you apply but expired by the registration decision it has to support. A checklist shows both items as done. Catching that takes something that knows when each document has to be valid and can reason over dates.

What it does

A caseworker enters an applicant's name, profession, country trained, target country and region, and gets a dated licensing pathway. Each step has a status, a plain-language detail and a link to the regulator page it came from, next to a panel that logs the agent's reasoning. Three scenario buttons send events to the agent:

  • Simulate a schedule delay: the agent re-reads the current steps, finds a dependency that no longer lines up, marks those steps at-risk, and explains the collision using dates computed by tools, with a concrete recommendation.
  • Simulate a document rejection: the agent marks an early step at-risk, inserts a remediation step, blocks the steps that depend on it, and explains why.
  • Reset the case: the agent rebuilds the clean pathway.

A real example from our live evaluation. Aida Torres is a nurse trained in the Philippines and applying in Ontario. On the delay run, the agent found a collision on its own. The slipped credential assessment pushes her registration decision to 2027-09-15, but a police criminal record check ordered as planned would expire on 2027-09-01. The knowledge base records that check as valid for 6 months and required to be valid at the registration decision. The agent recommended ordering it a month later.

For an unregulated profession such as Software Engineer, the grounding lookup says so. The agent returns a short work-authorization path and states that no practice licence is required, instead of forcing a licensing template.

Coverage: Canada (Ontario, British Columbia, Alberta), the United States (New York, California, Texas), the United Kingdom, Australia and Germany. The compliance knowledge base holds 8 curated, schema-validated rulesets. A compact fallback table covers other pairs, such as Physician and Teacher in Ontario and New York.

How we built it

Strands Agents SDK (Python) on Amazon Bedrock, deployed to Amazon Bedrock AgentCore Runtime.

  • Pathway agent: a Strands Agent running Claude Sonnet 4.6 on Amazon Bedrock (us-west-2), with three tools:
    • get_regulator_rules, which is the grounding tool. It maps the applicant's profile to a jurisdiction and returns the matching ruleset, or flags the profession as unregulated or the region as unknown_region.
    • today() and months_between(), which are deterministic date tools. The model never does date arithmetic in its head.
  • A fresh agent per request: a Strands Agent keeps its message history, so each request builds a new one over a shared model client. No applicant's context leaks into the next.
  • Structured output: responses are parsed straight into the same pydantic model that FastAPI uses as its response_model. If structured output fails, a validated text-parse fallback runs, and nothing unvalidated is returned.
  • Grounding enforced in code: after the model answers, _ground() sets the regulator from the lookup and drops any sourceUrl that isn't in that jurisdiction's curated set. A made-up link can't reach the UI.
  • Compliance knowledge base: a JSON Schema for one ruleset per profession and jurisdiction, plus an ontology with valid_at, validity_months and depends_on. Every source URL was checked by hand.
  • Web UI: static HTML/CSS/vanilla JS. A single adapter (agent.js) maps the UI onto POST /reason and rejects any reply that isn't in the contract shape.
  • Serving: FastAPI with a session orchestrator. Per-session locks and atomic writes let POST /batch run many caseworkers' sessions concurrently.
  • AgentCore Runtime: the same reason() function sits behind a BedrockAgentCoreApp entrypoint, deployed as credential_bridge and verified with a live invoke.
  • Evaluation: eval_run.py sends 7 scenarios to the live endpoint and scores contract validity, grounding, delay semantics and the unregulated-profession judgment.

Challenges we ran into

  • Modelling time, not just documents. The useful fact isn't "a police check is required". It's "the check is valid for 6 months and must still be valid at the registration decision". We added valid_at and validity_months to the schema so the agent has something concrete to reason over.
  • Our model went away mid-build. Claude 3.7 Sonnet reached end-of-life on Bedrock, so we moved to Claude Sonnet 4.6. The first live evaluation passed 5 of 7 scenarios. The rejection scenario hit the output-token limit, and the delay dates weren't in ISO format. After we raised max_tokens and required YYYY-MM-DD dates in the prompt, the second run passed 7 of 7.
  • Not trusting the model with citations. The prompt says "never invent a URL", but a prompt isn't a guarantee, so we moved that rule into code.
  • One contract for two runtimes and a UI built in parallel. FastAPI and the AgentCore entrypoint share the same pydantic models, and a small adapter maps the frontend's view model onto them.
  • Knowing when not to plan. Software engineering isn't a licensed profession in these jurisdictions. We made that an explicit, grounded outcome.

Accomplishments that we're proud of

  • Live evaluation (Claude Sonnet 4.6 on Amazon Bedrock): 7 of 7 scenarios passed every check. 95% of steps (52 of 55) cite a URL from the target jurisdiction's curated ruleset, with 0 invented or off-jurisdiction URLs. The delay, rejection and unregulated-profession judgments passed 3 of 3. This measures provenance, not correctness: it shows each cited link belongs to the right jurisdiction's ruleset, not that each step's content is right.
  • The agent found a real, regulator-sourced expiry collision without being told which pair to look for.
  • It's deployed on Amazon Bedrock AgentCore Runtime.
  • The caseworker experience is complete: intake, a dated pathway with source links, a reasoning log, and one-click scenarios.

What we learned

  • For regulatory reasoning, the schema matters more than the prompt. Encoding validity windows is what lets the agent find collisions.
  • Keep date math in tools and decisions in the model. The output becomes checkable and the prompt gets shorter.
  • Put the rules you can't afford to break in code, not in the prompt.
  • Define the contract once and reuse it in every runtime.

What's next

  • A human-approval step before a case is finalized, plus a deadline watcher.
  • An authenticated proxy so the UI can call AgentCore Runtime directly.
  • More jurisdictions, with human review before any harvested ruleset is used.
  • Durable session storage (DynamoDB or AgentCore Memory) and caseworker sign-off.
  • A pilot with a settlement agency using real caseloads.

Credential Bridge is a research aid for caseworkers, not legal or immigration advice.

Video

Make sure to watch the video in 1080p on YouTube

Built With

  • amazon-bedrock
  • amazon-bedrock-agentcore
  • anthropic-claude
  • aws-codebuild
  • boto3
  • claude-sonnet-4.6
  • css3
  • fastapi
  • html5
  • javascript
  • json-schema
  • pydantic
  • python
  • strands-agents
  • univcorn
Share this project:

Updates

Submission history