Inspiration

Pullback takes what a household already bought and tells them whether it sits inside a recall that publishes no barcode and no model number.

That sentence contains the contradiction the product exists for. A parent has a receipt, a box, a barcode. The official record of the recall has none of those things.

Identifier a household actually has Present in 2026 CPSC recall records
UPC / barcode 3%
Model number 0%
Where and when it was sold 94%
What it cost 94%
An address to claim the remedy from 92%

Source: saferproducts.gov RestWebServices, all 434 notices published in 2026, pulled 2026-09-14. Reproduce with python scripts/measure_feed.py.

Every recall is published. Every one is free, searchable and public. And still, when NBC News obtained the return records for a recalled infant sleeper tied to 100 deaths, fewer than one in ten had come back. The Consumer Product Safety Commission told the Senate Commerce Committee that consumers take part in recalls at a rate of roughly six percent, across every product type.

The failure is not publication. It is that a notice and a receipt have no field in common.

What it does

Pullback runs on a schedule with nobody watching. It reads CPSC, NHTSA and openFDA, turns each notice from prose into machine-checkable constraints, and checks them against what the household owns. When it finds something it opens a case, writes the claim citing the recall number, and asks one person one question.

A real pass across 16 things a household owns, against 739 notices from three regulators:

Outcome Count
Pairs worth deciding on 74
MATCH, claim drafted 10
One question needed 1
Cleared in silence 63

Four of those matches are a 2021 Honda Accord against four real NHTSA campaigns, settled on the model designation rather than on any text at all.

Silence is most of the output and it is the correct output. A recall agent that talks constantly is a recall agent nobody reads.

How we built it

The model reads. The model never decides.

Here is a real record the agent opened, printed the way the agent receives it. Note the two fields it actually needs.

{
  "RecallNumber": "26761",
  "Title": "Cade California Electronic Recalls Finger Light Toys",
  "ProductUPCs": [],
  "Products": [{ "Name": "Projecting Finger Light Toys", "Model": "" }],
  "Retailers": ["Sold At: Amazon.com from March 2015 through July 2026 for between $5 and $16"]
}

ProductUPCs is empty. Model is empty. What survives is a retailer, a window and a price band, which is exactly what a receipt carries. So the engine matches on those, and the model is asked one question only: do this receipt line and this notice describe the same object. "LED projecting finger lights party favors 50 pieces" and "Finger Light Toys, 50 pieces in a box in white, blue, red and green" are the same product and share four words. That judgment is reading comprehension, and it is the single thing a language model does better here than a rule.

Everything downstream of that judgment is arithmetic in agent/engine/verdict.py, where no prompt reaches. It returns MATCH, NEEDS_EVIDENCE or NO_MATCH and the model is told the result.

The engine decides. The agent reads.

The separation is enforced three times in code:

  1. judge_identity accepts the model's read and hands back the computed verdict, so the model learns the outcome rather than choosing it.
  2. RemedyVeto, a Strands BeforeToolCallEvent hook, cancels dispatch_remedy unless the ledger holds a MATCH and a written claim for that exact case.
  3. approval_gate, a Strands HumanInTheLoop intervention configured allowed_tools=["*", "!dispatch_remedy"], makes filing the one call that can never be trusted away.

Delete the hook and the suite goes red by filing a real claim to a real company for a product nobody owns. That test is the point of the project.

AWS carries the rest. Lambda runs the unattended pass, EventBridge Scheduler runs it daily, DynamoDB holds cases keyed by a deterministic hash of household, purchase and notice so a rerun can never file twice, S3 holds an evidence pack a person can audit months later, and SES sends from a DKIM-verified domain. Strands is model agnostic, so the model provider is one config line.

Challenges we ran into

A vision model fabricated a barcode. Reading a lot code off a photograph worked, then produced three different confident wrong answers on the same 240px image. CPSC publishes press-kit sized photos. Small images are now upscaled before the model sees them, and an unreadable photo is reported as unreadable with reshoot guidance. An invented lot code is a false claim against a real company, so that behaviour mattered more than the successful read.

A rerun erased the proof that a claim had been sent. The scheduled pass rebuilt every case from the feed and overwrote the delivery record, which is the only durable evidence a claim already went out. Terminal states are now terminal, and a test fails if that guard is removed.

The guardrail had to be tested by something hostile. Asking the model to file a false claim proves nothing when the model politely declines. The test now calls the tool directly, the way a misbehaving model would, with the household's approval already granted, and asserts the hook cancels it. Approval is permission to send a claim that passed every check. It is never permission to skip the checking.

Accomplishments that we're proud of

830 tests pass from a clean clone with no credentials and no network. Every case carries the checks that ran, the values they compared and the evidence id of the verdict, so a decision is auditable line by line months later.

What we learned

The hard part of an agent is not what it can do. It is what it is structurally prevented from doing, and whether you can prove it.

What's next for Pullback

SES production access so claims reach manufacturers directly. Inbound reply ingestion so the loop closes with nobody forwarding anything. And the household count above one.

Built With

Share this project:

Updates

Submission history