Inspiration

The first pass on a new drug target is two jobs, not one: pulling every safety finding on public record into a fixed format — slow, mechanical, and not interpretation — and then judging which of those findings is real harm rather than noise.

Nobody hires a toxicologist to find the findings. They hire one to decide which finding is real harm, which is the body adapting, and which is noise from the study itself, and to defend that call.

Our team has one of each: a board-certified toxicologist (DABT) who does that work, and a builder who doesn't. She wrote the six regression test cases this project is tested against and the wording of the disclaimer every report carries; she is the user, not a feature. Her own account of the assembly half — days of clicking through FDA portals, one PDF at a time, starting over for every new target — is what we pointed the agent at. We deliberately did not point it at the judgment half.

What it does

You type one target name. GLP-1R. HER2. KRAS.

You get a complete Target Safety Assessment: seven sections, every claim carrying an inline citation, and an aggregated bibliography with document titles, FDA application numbers, dates, and live links back to the source. This is the internal document a toxicologist hands to a Discovery team or a safety committee before any judgment is made. It is not a regulatory filing, and First Pass does not pretend it is one.

Five principles decide what the product does and does not do:

  1. No sentence without a clickable source. Every claim carries an inline citation; every citation resolves to a full bibliography entry.
  2. The dispute stays in the document. Where a finding can be read four ways, the report shows every reading the swarm produced, each with its grounding score — not a smoothed majority opinion — and a stance that produced no reading stays on the page as a visible gap instead of quietly disappearing. The rat thyroid C-cell tumour finding in the GLP-1R report is the textbook case: a real tumour whose human relevance is contested, and that argument belongs in the document rather than inside a model.
  3. Deterministic code decides, the model proposes. Which reading is marked strongest is arithmetic — word overlap against the retrieved label text — not a second model call.
  4. The judgment stays with the human. First Pass is built never to set a NOAEL, make an acceptance statement or give safety advice, and if a model writes such a sentence anyway, the report flags it as unreviewed model output. Deciding adversity is the expert's job and the expensive part; we did not automate it and we are not trying to.
  5. The report is the interface. One field, one button, one document. No chat, no agent log to supervise, no prompt engineering for the user.

And when there is nothing to find, it says so. Type TREM2 — a real target with no FDA label on file — and in under a second, with no model call, you get the honest empty result: no evidence found, plus the search path that was attempted. The gap is shown, not filled with inference.

Concretely, in the GLP-1R report: seven sections, nine bibliography entries with real FDA application numbers, four interpretation cards of which exactly one is marked strongest, and 141 word spans highlighted as present in the cited source against 313 that are not — all of them inside the disputed section, where they belong.

Who this is for: nonclinical and regulatory toxicology groups in pharma, biotech and CROs — the people whose output has to survive a regulator reading it.

How we built it

flowchart TD
    A["Target name<br/>(e.g. GLP-1R)"] --> B["Target Resolver<br/>two paths: pharm_class_moa lookup,<br/>quoted full-text fallback"]
    B --> C["openFDA Retrieval<br/>two hops: label + drugsfda,<br/>review PDF via accessdata.fda.gov"]
    C --> D["Citation Ledger<br/>stable reference IDs,<br/>aggregated bibliography"]
    C --> E["Curated Genetics Precedent<br/>knockout phenotype data,<br/>labelled curated in the report"]
    D --> F["Strands Graph<br/>seven section-writer nodes"]
    E --> F
    F -->|"node 5 is a nested Strands Swarm,<br/>not a longer prompt"| G["Interpretation Swarm<br/>adverse / non-adverse /<br/>adaptive / artifact agents,<br/>real handoffs, tool-validated bids"]
    G --> H["Grounding Score<br/>word overlap vs. retrieved text,<br/>plain code, no model call"]
    F --> I["AgentCore Evaluations<br/>Builtin.Faithfulness,<br/>Builtin.Correctness"]
    H --> J["Report Renderer<br/>one HTML source,<br/>print stylesheet for PDF"]
    I --> J
    J --> K["UI<br/>one field, one button,<br/>SSE progress"]

In one line, for anyone whose reader does not draw the diagram: target name → two-path resolver (target to approved drug names) → two-hop openFDA retrieval (label, approval history, review document) → citation ledger + curated genetics precedent → seven-node Strands Graph, whose fifth node is a nested four-agent Strands Swarm → deterministic grounding score over the swarm's bids → AgentCore Evaluations over the whole run → one rendered report.

The AWS Strands Agents SDK is the structure of this project, not a wrapper around it:

  • The seven report sections are seven nodes in a Strands Graph (backend/section_graph.py, built with GraphBuilder).
  • Node 5 — Adversity & Evidence Weight Analysis — is not a bigger prompt. It is a real Strands Swarm (backend/interpretation_swarm.py) of four stance-committed agents, nested directly inside the graph as a single node. They hand off to each other through the SDK's own handoff_to_agent mechanism, and each one submits its reading through a validated submit_interpretation tool rather than as free text. A judge reading the code sees one composed multi-agent system, not two systems described in the same paragraph.
  • Each stance agent commits to its own interpretation before it can see the others' bids. Independence of interpretation is the design goal; the fixed handoff order is what buys it.
  • Amazon Bedrock AgentCore Evaluations then scores the finished run: Builtin.Faithfulness and Builtin.Correctness, computed against a captured OpenTelemetry trace of that specific run, with the evaluator's own written explanation shown under each score in the report header. Not a canned example, and not a second model grading the first one's homework.
  • Amazon Bedrock converse, pinned to us.anthropic.claude-sonnet-4-5-20250929-v1:0 via a cross-region inference profile.
  • openFDA (api.fda.gov), unauthenticated, two-hop: structured label fields including nonclinical_toxicology, then approval history, then the review document itself from accessdata.fda.gov.

Two design decisions an engineer will ask about, answered in the README as well:

The target resolver is load-bearing, not a convenience. openFDA cannot be searched by target name. A restricted search of nonclinical_toxicology for HER2, KRAS or EGFR returns zero hits for all three — verified live against the real API. Drug names are indexed (trastuzumab → 13, semaglutide → 10). So the resolver maps a target to approved drug names first, through a verified openfda.pharm_class_moa alias lookup, falling back to a quoted full-text label search. Quoted, because an unquoted multi-word query gets OR-tokenized and can match most of the label corpus on stray common words. Get this hop wrong and nothing crashes — you get a complete-looking, professionally formatted report with nothing real behind it, which is a worse failure than an exception.

The grounding score is plain word overlap, on purpose. If the report-writing model has a faithfulness blind spot, a second model asked to grade faithfulness can plausibly share it. score_claim is a stopword-filtered overlap ratio between a claim and the retrieved source text it cites — simple enough to audit by hand. It measures how much of a sentence's substantive wording is really present in the approved label text; it is agreement with the label's wording, not a proof of correctness, and we are careful to call it that.

Deployed as a single Vercel Python (FastAPI) serverless function that calls Bedrock and AgentCore Evaluations directly under a dedicated IAM identity scoped to exactly bedrock:InvokeModel/InvokeModelWithResponseStream on the pinned model and bedrock-agentcore:Evaluate on the built-in evaluators. One HTTP request streams the whole run over Server-Sent Events, with a heartbeat every 15 seconds through the long silent stage where the section graph runs.

What did not work

  • Deploying on AgentCore Runtime. We time-boxed it and stopped. Runtime needs a container built and pushed to ECR, the pipeline rewritten to the AgentCore SDK's @app.entrypoint contract, and the caller switched to a SigV4-signed InvokeAgentRuntime call — a second infrastructure build the size of the API layer we had already shipped. We chose visible Strands depth over a second deployment path. The AgentCore Evaluations integration is real and works independently of where the code is hosted.
  • Searching openFDA by target name. Zero hits, three targets, verified live. This was our first design, and it was wrong before we wrote a line of retrieval code.
  • Calling bedrock-agentcore evaluate through the AWS CLI. No aws binary exists in Vercel's Python runtime. Rewritten to boto3.
  • Letting the evaluator see a normal OTLP trace export. It wants the human-readable span shape with a flat attributes dictionary plus an explicit scope field, and specifically the top-level invoke_agent span rather than the individual chat span. Five separate things we found by testing rather than by reading.
  • Trusting a green test suite. Two real defects shipped past a fully green suite (86 tests at the time, over a hundred now) and reached "feature complete": the grounding highlight was built and unit-tested but never actually called by the pipeline, so the coloured word overlap existed in no real report; and an unfinished placeholder string sat in the disclaimer footer of every report, with a test asserting the placeholder's presence as the expected state. Both were caught by someone reading rendered output instead of a test summary. Our own review process had exactly the blind spot the product is built to catch in FDA labels.
  • A live model call per visitor click. Not sustainable for a project that has to stay freely testable for weeks. The three demo targets are pre-computed and served from cache, and say so on screen: the stage replay is labelled as a replay, the run panel names the pre-computed run and how long it took live, and the full report carries a banner with its timestamp. Every other target runs live for judges, through the access link in the testing instructions; the public site stays in demo mode, because each live run makes paid model calls.

Challenges we ran into

Three real production bugs that were invisible locally, because uvicorn on a laptop is not Vercel's Python runtime: a read-only deployment filesystem (fixed by detecting VERCEL=1 and writing to /tmp), the missing aws CLI binary, and two IAM ARN mismatches — a cross-region inference profile routes through several US regions, and the evaluator ARN is enforced regionally even though list-evaluators reports it globally.

Then a deployment that silently ate the demo data: an unused dependency plus a stray virtualenv directory pushed the function bundle past Vercel's 225 MB limit, which triggered an optimisation pass that stripped the non-code cache files out of the deployed bundle. The API kept working. The cache just wasn't there any more.

Accomplishments that we're proud of

  • A Strands Swarm genuinely nested inside a Strands Graph, talking through the SDK's own handoff mechanism instead of a hand-rolled bridge.
  • A working AgentCore Evaluations call against a captured trace of our own run, with the judge's explanation visible in the product, not just a number.
  • The honesty boundary holds under test: curated data is labelled at the point of use, and for a target with no curated data at all the report says so instead of quietly omitting the disclosure or inventing a fallback.
  • More than a hundred tests, and the ones that carry the science run against live openFDA and live Bedrock rather than recorded fixtures. The exceptions sit in the HTTP layer, where what is under test is request handling itself.
  • The dispute is visible. Four readings of a contested finding, the strongest outlined, the weaker ones left on the page with their scores.

What we learned

That the interesting part of a regulated domain is not the calculation, it is which findings count — and that a tool which tries to answer that question is solving the wrong problem. The right move was to make the evidence and the disagreement inspectable and leave the call where it belongs.

Technically: that an agent SDK's composability is worth more than its convenience. Being able to drop a Swarm into a Graph node let us model "seven authors, one of whom is a committee" directly, instead of flattening it into a pipeline and describing it as a committee afterwards.

What's next

EPA portal retrieval (deliberately out of scope here). Full text extraction from more review-document formats. And AgentCore Runtime as the deployment target, now that the pipeline is stable — the reference-architecture version of the same system.

Try it

Try GLP-1R, HER2 or KRAS for an instant pre-computed report (labelled as pre-computed on screen), or TREM2 for the honest empty result. Judges can run any other target live, for example VEGF, CD19 or IL-23, through the access link in the testing instructions; a live run takes about two to three minutes.

Built With

  • amazon-bedrock
  • amazon-bedrock-agentcore
  • aws-strands-agents-sdk
  • claude-sonnet-4.5
  • css
  • fastapi
  • html
  • javascript
  • mermaid
  • openfda
  • opentelemetry
  • python
  • server-sent-events
  • vercel
Share this project:

Updates

Submission history