-
-
She feeds twelve kids and files paperwork for every bite. Tally does the counting.
-
52 percent of licensed homes are gone, and the average home leaves $436.92 a month unclaimed.
-
Four steps, and none of them ask her to type. Say who is here, photograph the plate, fix it, approve.
-
Every meal beside the photograph it was judged from, with each food, its confidence, and the verdict.
-
The snack that will not be paid, the one change that would fix it, and everything Tally said out loud.
-
Nine children, their allergies, ages and subsidy status, checked before anything is written down.
-
The allergy question is red and jumps the queue. Safety never waits for the interruption budget.
-
Notes home for every family, written in the language that family actually reads.
-
The month's claim, $161.98, computed inside Amazon Bedrock AgentCore Code Interpreter.
-
Every step the agents took, in order, including the questions they decided were not worth asking.
-
The same month from the state reviewer's desk, with the rulebook version that judged each meal.
-
Every meal, day by day, with its photograph and rule version, so a claim holds up a year later.
-
Photograph any plate of your own. The same vision model and rulebook say whether it would be paid.
-
The same day in dark mode. Both themes are designed, not inverted.
-
One photograph in, the day answered in seconds, and the paperwork closed while the food is still warm.
Inspiration
The paperwork happens at nine at night, and that is why it goes wrong.
A licensed family child care provider looks after four to twelve children in her own home. She is a small business owner. To be paid for the meals she serves, she must log every meal for every child by component under the federal food program, and since 2024 she must document actual daily attendance for every subsidised child. She does all of it after the last child goes home, from memory, hours after the plates were cleared.
The sector has been saying this for years, and the numbers are stark:
- Licensed family child care homes fell 52 percent between 2005 and 2017, more than 97,000 homes. Providers serving subsidised children fell 51 percent. (HHS Office of Child Care)
- Almost 80 percent of state agencies name burdensome paperwork as the top barrier to food program participation. (American Journal of Public Health, 2023)
- For home providers, meal documentation "frequently becomes evening responsibilities". (PMC, 2023)
- The National CACFP Association published a report on exactly this burden on 19 August 2026, three weeks before this hackathon closed. (Stretched Thin)
Here is what it costs her. One lunch a day that fails the log, for six children, over 22 serving days, is 436.92 dollars gone in a month at this year's tier 1 rate of 3.31 dollars per lunch (USDA payment rates, July 2026 to June 2027). Not a fine. Money she earned, cooked for, and does not receive, because the record was written at nine at night instead of at noon.
So the goal was never to make the paperwork faster. It was to move it to the one moment when it can still be fixed: while the food is on the table.
What it does
1. Say who is here. "Maya's here, Leo's here, Ava's mom said she's sick, Mateo's coming at eight." From that one sentence: attendance, the subsidy check, and the staff to child ratio for the rest of the day, including the part of the afternoon that has not happened yet.
2. Photograph the plate. The agent names each food and maps it to a food program component. Then deterministic rules decide whether the meal is reimbursable, for every age group at the table, because a plate that qualifies for a four year old may not qualify for a one year old.
3. Fix it at the table. "Add milk and this breakfast qualifies." Not a rejection next month, when nothing can be done. This is the entire product in one sentence.
4. Allergies are checked before anything is written. In the demo, this catches peanut butter in a bowl of porridge, on a real photograph, and that catch was not scripted. I put a real picture of porridge in front of it and it found a real problem.
5. The evening closes itself. A Strands Graph runs the ledger, drafts a note home for each child in that family's language, checks the compliance clock for drills and training and subsidy days, and then puts at most two questions to her. That budget is enforced in code, because her attention is the thing this product exists to protect.
6. The month is already filed. Every meal carries the photograph it was judged from and the rulebook version that judged it, which is what makes a claim defensible to a state reviewer a year later.
And you can point it at your own lunch. There is a page that takes a photograph of any plate, from your camera or your files, and returns the components it found with the confidence behind each one, whether it would be paid, and the smallest change that would fix it. Nothing is stored: the photograph is read, the answer is returned, the file is deleted.
How we built it
The day and the evening are different problems, so they have different shapes.
During the day the provider is standing there with a plate in her hand, so the work is
request driven and shallow, and only the step that genuinely needs a model uses one. The
API calls one tool per event. log_plate runs the Plate agent, a Strands agent on Bedrock
vision at temperature 0, and everything after the reading is deterministic. take_roll
uses no model at all. Latency is the product, so there is no orchestrating agent standing
between her and the answer.
In the evening the work is a fixed pipeline whose steps feed each other, so it is a
Strands Graph: Ledger, then Parent Notes, then Compliance, then the Digest, in that
order, every time.
This is deliberately a different Strands shape from Turnout, my other submission, which uses a conditional Graph and the Agent to Agent protocol. Two problems, two shapes. I did not want to reach for the same pattern twice and call it architecture.
The model does the one thing only a model can do. Bedrock vision at temperature 0 reads what is on the plate, so one photograph gives one record. Everything touching money or licensing is plain Python against versioned JSON: whether the meal is reimbursable, the claim arithmetic, the daily maximum, the ratio limits, and the decision whether to interrupt her. A rule a prompt can talk its way around is not a rule, and a year from now an auditor has to get the same answer.
The sponsor should not have to trust any of that. The organisation that questions a claim is the sponsoring agency, and what reviews a claim on their side is increasingly an agent rather than a person with a browser. So the same rulebook is served over the Model Context Protocol, and a separate Strands agent consumes it as a sponsor's reviewer. That reviewer holds no Tally code. Its entire toolset is discovered at runtime over stdio, it is read only by construction, and every answer carries the rule version it came from. A reviewer that disagreed with the provider's software over the same facts is the failure that design exists to prevent.
On AWS: Amazon Bedrock (Claude Sonnet 4.6 for vision, Haiku 4.5 for the fast paths), Bedrock AgentCore Code Interpreter, which computes the month's claim, and AgentCore Memory, which holds the answers she has already given. Both are live on the deployment and the screen says which path answered. It runs on AWS App Runner, live, with no login.
Challenges we ran into
Knowing which half of the problem to take. Vision language models recognise food well and estimate portions and grams poorly. The food program happens to reimburse on components, not grams. So Tally does the half these models are good at and refuses the other half out loud, rather than quietly overreaching into nutrition numbers it cannot stand behind.
Not guessing. Below a confidence threshold, an item is left out of the record and one question is asked instead. Guessing a component into compliance would create a false claim, which is far worse for a provider than being asked whether the cup is milk or juice. Seven of the thirteen demo photographs prompt a clarifying question, and that is the gate working, not a failure.
Memory keyed on the wrong thing. The two questions a night only mean anything if both are new. The subsidy reconciliation question is what breaks without memory: it fires when a subsidised child is present on a day their authorisation does not cover, and a family whose Tuesday is not on the paperwork has an unlisted Tuesday every week. Keyed on question id, she answers the same thing every Tuesday forever. So an answer is now stored against its topic, the child, and the weekday, which is the thing that actually recurs, and it ages out after 45 days because a family's schedule does change and an answer from three months ago is not evidence about this week. A failed lookup asks her anyway: failing safe here means a repeated question, not a subsidy day she is owed and never claims.
Measuring it on photographs I did not choose. On 13 curated demo photographs it gets 100 percent component recall over 2 trials each, with 0 spurious components in 26 readings. That number is flattering and I knew it. So I ran it against 45 photographs it had never seen, the first five files in each of ten Wikimedia Commons categories, and it got 35. The interesting part is the 10 it got wrong: 9 of them credited nothing at all rather than inventing a component. When it does not know, it stays quiet, which is the failure mode a provider can survive.
Accomplishments that we're proud of
- The allergy catch nobody planned. Real photograph, real hazard, found at the table before anything was written.
- A sponsor can verify the rulebook without trusting the software, through their own agent, over an open protocol, with a version stamp on every answer.
- When it is wrong, it is silent rather than confident. 9 times out of 10 on photographs it had never seen.
- It refuses 13 of 13 adversarial cases and wrongly refuses 0 of 6 legitimate ones. The second number is the one that matters. Refusing everything would score perfectly on the first and make the product worthless, since the whole point is getting her paid for the meals she did serve.
- Two questions a night, enforced in code, not requested in a prompt.
- It works at 200 and 400 percent zoom, in both themes, with a keyboard, because the person using it is holding a phone in a kitchen.
What we learned
Take the half of the problem the model is actually good at, and say plainly that you are not doing the other half. Scope is a feature.
Silence beats a confident guess when the output is a federal claim.
Memory should be keyed on whatever recurs in the real world, not on whatever the software happens to have an id for. That one change is the difference between a memory feature and a memory demo.
Say what is live and what is designed. AgentCore Code Interpreter and Memory are running here. Gateway and Runtime are designed and labelled as designed. Precision is more persuasive than claiming everything.
What's next for Tally
- Real speech. The roll call is typed in this build. The parser and the spoken responses are real, and Amazon Transcribe sits in front of the same function with Amazon Polly behind it.
- Infant meal patterns. Infants follow a separate pattern by age in months, so today Tally records their meals and explicitly declines to judge them rather than applying the wrong rule. That is the honest stop, not the finished one.
- More states. Only Texas is in the ratio table. An unknown state raises rather than guessing, because a wrong ratio limit is worse than no answer. Adding a state is a data change, not a code change.
- AgentCore Gateway and Runtime, moving the rulebook and claim filing behind managed tools with credentials out of agent code.
- A pilot with one sponsoring agency using the MCP reviewer against real claims.
Rosa and every child are fictional. The photographs are real, openly licensed pictures of real food from Wikimedia Commons, credited in the repository, because testing food recognition against drawings would prove nothing.
Built With
- agentcore-code-interpreter
- agentcore-memory
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-ecr
- aws-app-runner
- claude
- computer-vision
- docker
- fastapi
- javascript
- mcp
- model-context-protocol
- opentelemetry
- pydantic
- python
- strands-agents

Log in or sign up for Devpost to join the conversation.