Inspiration

A recall notice identifies a product, but someone still needs to find the right boxes, inspect their labels, and record what happened next.

I built Pantry Recall Response Agent around one practical question: Which shelf needs a person next, and what evidence or action is still missing?

The goal is to help community food pantry volunteers turn a recall notice into specific, traceable tasks.

What it does

Pantry Recall Response Agent investigates pantry inventory, identifies missing label information, and recommends the next checks with supporting source evidence. When a volunteer adds evidence or confirms an action, the agent revisits outstanding work.

The public demo gives each visitor a separate fictional pantry using a historical recall. One example follows four boxes with an incomplete printed code:

  1. The agent flags the boxes for label inspection.
  2. A volunteer supplies the missing code, which matches the recall.
  3. The identification task finishes, while the task to place the boxes on hold remains open.
  4. A simulated confirmation for two boxes leaves two outstanding.
  5. Confirming the remaining boxes completes the hold task and triggers another briefing.

Identifying affected stock and confirming that someone physically handled it are separate steps. The agent can investigate and recommend; people record physical actions.

How I built it

I built the application in Python using Strands Agents SDK and Amazon Nova Pro on Amazon Bedrock.

The agent uses seven tools to inspect recall information, compare inventory, review case histories, and submit a briefing. Programmed comparison rules check product scope and preserve missing information as unknown. The application validates task states and quantities, then attaches exact stored source evidence to the agent’s recommendations.

SQLite stores evidence versions, event history, tasks, and human confirmation receipts. Accepted label updates and hold confirmations trigger follow-up investigations after the initial agent start.

The application runs on Amazon EC2 behind Amazon CloudFront, with encrypted Amazon EBS storage, a private Amazon S3 release bucket, and an IAM instance role for Bedrock access. The interface uses HTML, CSS, and JavaScript.

Challenges I ran into

Getting the agent to finish its investigation

In the first live evaluation of a second recall, the agent skipped all ten required scope comparisons. I preserved the failed run and added a bounded continuation that identifies missing tool work and lets the agent complete it. Automated tests exercise that recovery path.

Keeping citations tied to real evidence

Another run completed its comparisons but repeatedly invented citation identifiers and exhausted its call allowance. I changed the interface so the agent selects existing task IDs and explains its recommendations, while the application attaches the exact stored evidence.

The previously failing session then completed without resetting the pantry.

Tracking partial completion

Inspecting a label does not mean the stock has been isolated. Confirming two boxes does not complete a four-box task. Keeping evidence, quantities, and human receipts separate made those distinctions explicit and auditable.

Accomplishments I'm proud of

  • A public AWS demo with a real agent, visible tool activity, and case history.
  • Follow-up briefings triggered by new evidence and human confirmations.
  • 165 passing automated tests, including pantry workflow and benchmark tests.
  • 18 of 18 frozen synthetic recall scenarios passing across two reviewed recall projections.
  • A recorded four-stage public Nova Pro walkthrough with eight inventory comparisons per stage and no agent-generated physical-action confirmations.

These checks demonstrate specific prototype behaviors; they do not establish a general success rate or measured time savings for volunteers.

What I learned

The most useful role for the model is deciding what to investigate next and explaining why. Exact product identifiers, current quantities, and completion records need checks enforced by the application.

I also learned to treat failed evaluations as design evidence. Skipped comparisons changed the execution controls. Invented citations changed how recommendations connect to sources. Preserving those failures makes the reasoning behind each improvement inspectable.

What's next

Next steps include independent review of the recall scenarios, evaluation on a previously unused recall, and usability testing with pantry volunteers. Broader use would require additional reviewed product adapters, recall ingestion, and operational monitoring.

The current project is a prototype using historical notices and synthetic inventory. It has not been tested in a pantry pilot. Recommendations remain advisory, and physical actions require human confirmation.

Built With

Share this project:

Updates

Submission history