Inspiration

Recalls exist because products injure people. The remedies are free. Most of them stay unclaimed. The notices go to everyone, and nobody matches them to the products in a home. Warranties expire with claims never filed. We wanted an agent that does this work at night. It speaks only when a decision is necessary.

What it does

Guardian is a background agent for one household. It builds an inventory from forwarded receipts and order emails. Each night, it pulls the public recall feeds from CPSC, NHTSA, and openFDA. It compares the new recalls with the inventory. It watches the warranty windows.

When a recall applies to a product that the household owns, Guardian sends one message with one decision. After the approval, it writes the remedy request, attaches the receipt, and sends it. Then it schedules a follow-up.

The best output of the agent is silence. The primary metric is how rarely it must speak to you.

Who it is for. The person who manages a household: parents of young children, car owners, and anyone with appliances.

How we built it

  • Strands Agents runs every graph. We built gren, a small graph runtime on the Strands Agents SDK. A YAML file declares the nodes: agents, code, verifiers, and a human gate. gren compiles the file into a Strands Graph. It derives the edges from the data references. It enforces a cost budget, retries, timeouts, and a quorum on fan-outs.
  • Code first, model second. The nightly sweep is a graph of 13 nodes. Code does the deterministic work: the feed fetch, the normalization, and match stages 1 to 3 (UPC, VIN, model number, lexical similarity, sold window). Only the ambiguous pairs go to a model. Claude adjudicates them with structured output, through the Strands model provider: the Anthropic API on the live site today, Amazon Bedrock the hour the account is authorized.
  • A verifier with authority. A triage agent applies the severity rules and the notification budget. An adversarial verifier examines the plan. It can reject the plan and send the triage agent back for one repair round.
  • A durable human gate. The gate is a Strands interrupt. The run pauses on disk. The process can exit. The answer of the household resumes the same graph. Completed nodes replay from the checkpoint without new model calls.
  • One send, never twice. A side-effect node sends the remedy one time. A frozen constraint makes the node unreachable without an approval.
  • A dashboard that shows the work. The React dashboard has the household view and the Agent flow tab. The Agent flow tab is a full trace workbench: the run graph, a timeline, the event log, metrics, the gate, the YAML spec, and a node inspector with fork. It installs as an app (PWA) and keeps working offline for reads.
  • AWS. One EC2 instance behind Caddy serves the live dashboard with automatic HTTPS. S3 keeps the state. SES sends the remedy email. The Bedrock provider and the AgentCore deployment are built and tested; both wait on the account (see the challenges).

Challenges we ran into

  • Our AWS account is new to Bedrock. Every model call answered "Operation not allowed" until AWS Support authorizes the account, and the case was still open at submission. In gren the provider is a configuration choice, not a code path, so we set one environment variable and the live server ran the same graphs on the Anthropic API through the Strands AnthropicModel. The Bedrock path stays in the repo; the switch back is the same one-line change. The full record is in docs/PROBLEMS_EXPERIENCED.md.
  • The AgentCore quotas of the account are 0. Only AWS Support can raise them. We built the EC2 path, so the demo has a public URL today.
  • A verifier that cannot see the item facts rejects correct plans. We gave it the confirmed matches. The repair loop became useful.
  • A late answer from the household must not block an approved one. An approval now releases the gate at once. A later answer forks the run from the gate.

Accomplishments that we are proud of

  • A real first sweep on live feeds. The sweep pulled 594 CPSC recalls, 3 NHTSA campaigns, and 1,200 openFDA reports. The code compared 7,143 pairs. The model adjudicated 3 ambiguous pairs. The run stopped at the gate after 235 seconds and $0.41. One approval produced a refund request to the real contact address of the recall.
  • A quiet night makes zero model calls. The graph skips every model node when nothing matches.
  • 50 tests run the graphs end to end on a mock provider: the gate, snoozes, late approvals, a verifier rejection with repair, and a quiet night.
  • The demo video is a real recording of the live site: a sweep on live feeds, the pause at the gate, the approval, and the remedy, with no staged data.
  • The Agent flow tab shows every run as the graph executed it. Nothing on it is illustrative.

What we learned

  • Put the deterministic work in code. Give the model only the ambiguous pairs. The cost stays at cents, and the result stays inspectable.
  • A verifier with rejection authority is worth more than a second prompt.
  • A gate must be durable. A household answers hours later, or never.
  • An agent that acts at night must show its work in the morning.

What is next

  • Bedrock as the live provider, the hour the account is authorized. It is one environment variable.
  • Email delivery through SES and SMS through SNS, when the account sandbox opens.
  • AgentCore Runtime, Memory, and Gateway, when the quotas allow it.
  • Photo intake of receipts, grocery advisories, and class-action settlements.

Built With

Share this project:

Updates

Submission history