Inspiration

A household is holding a notice from the state. It says their monthly SNAP benefit is $210. They ask a benefits navigator the only question that matters: is this number right?

Today the honest answer is that nobody checks. The notice gives one figure and does not show the budget behind it, and there is no independent way to re-derive it. So the agency's number is the only number there is.

USDA's own quality-control review put the FY2024 SNAP payment error rate at 10.93% — 9.26% overpayment, 1.67% underpayment. An underpayment means a household eats less than the law entitles it to, every month, and nobody in the room can tell.

What it does

Second Budget re-derives the allotment from the household's own circumstances and, when it disagrees with the notice, localises the disagreement.

It does not claim to know which stage the agency got wrong — a single final figure cannot support that. It answers the question a single figure can support:

For this notice to be correct, what would have to be true about this household that the household says is not true?

On the worked example that comes out as:

The notice says $210. The engine derives $473.
Difference: $263 a month in the household's favour.

For the notice to be correct, one of these would have to be true:
  * unearned income is $584.41 higher than stated  -- 7 CFR 273.9(b)
  * earned income is $728.71 higher than stated    -- 7 CFR 273.9(b)

Shelter costs and the medical deduction are absent from that list, and their absence is the point: driven all the way to zero they move the allotment from $473 to $313, nowhere near $210. They cannot explain the gap, so they are dropped rather than printed as an impossible number. Every line that remains is verified by feeding the value it names back through the budget and requiring the agency's figure to come out.

The output is a fair-hearing request in which every monetary figure records where it came from and every regulation is quoted verbatim.

How I built it

The architecture follows one rule: the model elicits facts and drafts prose. It never computes.

  • The engine is a MultiAgentBase peer node in the graph, not a @tool. A tool is something a model chooses to call, chooses arguments for, and can decline. The SNAP budget is a statute compiled to arithmetic; the graph simply runs it and the result is not negotiable.
  • Elicitation is a bounded cyclic GraphBuilder graph. The required-fact set is conditional — learning that a household member is elderly or disabled adds a medical-expense requirement that did not exist a moment earlier — so the number of rounds is not knowable in advance. The loop ends when a pure Python predicate over a fact ledger returns true, not when the model announces that it feels finished.
  • NumbersGate is an InterventionHandler on before_tool_call that denies any tool payload carrying a currency token absent from a frozen engine certificate. Set membership: decidable, and unit-tested in both directions.
  • CitationGate denies any quotation that is not a byte-identical span of the section it cites. An agency representative at a hearing looks the citation up and reads it back; a quotation smoothed into better English is not a small error, it is a reason to dismiss the filing.
  • A BeforeToolsEvent interrupt confirms a whole batch of proposed facts in one gate, with anything the model inferred rather than heard marked for scrutiny. Rejecting one re-opens the fact frontier and sends the graph round again.
  • The 7 CFR 273 index is a custom MemoryStore that fails closed at three seams, because MemoryManager.search catches a store's exception and returns an empty list — so a dead index and a regulation that says nothing are otherwise indistinguishable.

Runs on Amazon Bedrock (amazon.nova-pro-v1:0, us-east-1) with the Strands Agents SDK 1.53.

Is any of it true?

The engine is validated by exact replay against the USDA SNAP Quality Control public-use microdata for FY2024 — 44,891 real, de-identified households carrying both the inputs to the budget and the benefit that resulted. It runs from raw components and never reads the file's net-income column; net income is derived. That is what makes it a test of the budget rather than of one subtraction.

stage agreement
excess shelter deduction 42,386 / 42,388 = 99.995%
total deductions 42,386 / 42,388 = 99.995%
net income 42,383 / 42,388 = 99.988%
allotment 42,385 / 42,388 = 99.993%

Integer equality, no tolerance. The three remaining households are singletons, not a pattern.

Then the question the product exists for. Comparing an independent recomputation against the benefit actually received, and against the federal reviewer's own error findings, over 43,299 households:

reviewer found no error reviewer found an error
engine agrees 23,596 204
engine disagrees 2,585 16,914

Agreement implies no finding in 99.14% of cases; disagreement implies a finding in 86.74%. Agreement is measured — causation is not claimed, and both error cells are published rather than summarised away.

Challenges I ran into

A naive implementation of the allotment formula scores 91.287%. The misses are not noise. They are shaped, and the shape names the rule that is missing:

  1. −23 on 3,259 records — the minimum benefit, which applies only to household sizes 1–2 and applies even when the computed allotment is zero, not merely when it is small.
  2. The rounding sits before the subtraction — 7 CFR 273.10(e)(2)(ii)(C) rounds the household's 30% share up and then subtracts.
  3. The earned income deduction truncates. floor matches 99.955%; rounding to nearest matches 90.059%, missing low by exactly one dollar on 4,209 households.
  4. Adjusted income is rounded before it is halved, not after. Halving first scores 74.502%.
  5. Python's built-in round() is banker's rounding. Using it for the half-of-adjusted-income step scores 87.487% instead of 100.000% — 5,304 households wrong by a dollar, with no exception and nothing in the output to suggest it.
  6. −54 on 38 records — a homeless household's flat shelter deduction replaces the excess-shelter stage rather than adding to it. ceil(0.30 × 180) = 54, exactly. That is how it was found.

Worth stating plainly: the statute does not round consistently. The earned income deduction truncates down, lowering the deduction; the household's share rounds up, lowering the benefit. Both defaults fall the same way for the household.

On the SDK side, four behaviours were measured rather than assumed, and each changed the build:

  • Edge conditions dispatch on the parameter name invocation_state. Spell it ctx=None and there is no error and no warning — the condition gets its default forever. Measured on this exact graph: a 3-iteration convergent loop became a 12-node runaway that only the execution cap stopped.
  • A limit breach does not raise. status becomes FAILED, but failed_nodes is an int and is 0, and every node still reports completed. if result.failed_nodes: misses it entirely.
  • BeforeToolsEvent derives the interrupt id from the interrupt name alone, and an id that already carries a response is returned rather than raised. A fixed name gave all three elicitation rounds the same id — so one approval would silently have covered the next batch. The name now carries a content hash of the proposed facts.
  • A Guide at after_model_call does not survive a session restore, while a Deny at before_tool_call does. For a case a navigator reopens days later, guidance-based control is not weaker — it is absent. Every hard gate here sits on before_tool_call.

What I'm proud of

148 tests pass in under two seconds with no AWS credentials in the environment. They are driven by a real Strands Model provider that replays a script instead of calling a backend — not a mock: it emits the same StreamEvent sequence Bedrock does, so the agent loop cannot tell the difference. Graph topology, every intervention denial, every interrupt round and the exact model-call count are all assertable offline.

And what the project refuses. The regulatory constants are transcribed from the published schedule and then proven, per household, against the microdata's own columns. That check found something: all 861 Illinois households carry a standard deduction exactly $7 below the published contiguous figure, at every size band. 861 of 861 is a rule, not noise — and no published source for the variation could be found.

So Illinois is dropped from coverage rather than having the value copied out of the microdata. A table fitted to the data it is checked against has stopped being evidence. Alaska is refused for a different reason: it runs three separate benefit schedules and the public-use file records only the state, so falling back to the contiguous figure would produce an answer that is plausible, wrong, and indistinguishable from a right one — which is the exact failure this project exists to attack.

What I learned

Before wiring a single agent, find a public dataset that contains both the inputs and the answer, and make the deterministic core reproduce it exactly. Then insist the residual histogram be empty, not small. A flat mismatch rate tells you nothing; a residual that clusters on one value is a specific rule missing from your code, and it names itself.

The LLM part of an agent cannot be unit-tested. So make the part that can be as large as possible, and prove it against reality.

What's next

Coverage beyond FY2024 and beyond the contiguous benefit schedule; state-level utility allowances; and a second reviewer for the elicitation step, which is the one part of the system that tests cannot vouch for.

Honest limitations

  • Fact extraction is not verifiable by tests. Only the engine is. Every fact is human-confirmed before it enters a budget and inferred facts are a distinct provenance class, but that is a mitigation, not a proof.
  • Coverage is FY2024, the contiguous benefit schedule minus Illinois. Everything outside it is refused with the missing constant named.
  • Citations are section-level rather than paragraph-level. eCFR publishes no paragraph paths, and two rounds of heuristics still left 187 duplicate citations across 2,418 paragraphs — roughly 8% of paths wrong. A wrong citation in a legal filing is worse than a coarse one. The quotation itself is exact, and the quotation is what an advocate argues from.
  • Layer B measures agreement, not causation. A reviewer's finding can rest on facts the file does not expose, and a disagreement can equally be the engine's own limitation.
  • Reproducing quality-control records proves the engine implements the rules as QC applied them — not that any given state agency's production system does.
  • Not legal advice. The packet is a draft for a navigator or advocate to review, correct and file.

Built With

  • amazon-bedrock
  • amazon-nova
  • pytest
  • python
  • rich
  • sqlite
  • strands-agents
Share this project:

Updates

Submission history