Inspiration

Every year, $60 billion in federal benefits go unclaimed in the United States — not because the programs don't exist, but because no one connects families to them. Eight separate programs, eight different websites, eight sets of eligibility rules. A working parent juggling two jobs doesn't have time to figure out whether they qualify for SNAP or LIHEAP or WIC, let alone navigate the applications. We built AidRadar because the problem isn't generosity — it's access.


What it does

AidRadar is an AI-powered benefit finder for low-to-moderate income families across all 50 US states. It runs a four-agent pipeline that takes a household from zero to a complete, personalized benefits report in under two minutes.

The Intake Agent collects household details through a natural conversation — no Social Security number, no documents. A state dropdown eliminates the first friction point; a three-tier fuzzy matching system handles corrections and misspellings conversationally. The Eligibility Agent runs the profile through PolicyEngine, the same open-source microsimulation engine used by governments and academic researchers, checking eight federal programs simultaneously: SNAP, Medicaid, WIC, TANF, SSI, Lifeline, LIHEAP, and Free School Meals. The Recommendation Agent generates a plain-English report with estimated monthly dollar amounts, required documents, and direct state-specific application links — plus cliff effect warnings when a raise could cost more in lost benefits than it gains in income. The Monitor Agent runs on a schedule via AWS EventBridge and re-checks every saved profile when federal poverty guidelines update, notifying only the families whose eligibility actually changed.

The results screen includes a What If Simulator that lets families explore how a change in income or household size shifts their benefits — using the same real PolicyEngine calculation, instantly, without re-running the full pipeline. A Data Confidence Indicator shows High / Medium / Low confidence based on data completeness — transparently explaining which fields were estimated or missing and how that affects precision, so families can trust or improve their results


How we built it

AidRadar is built on Amazon Bedrock with the Strands Agents SDK. Each of the four agents is a dedicated Strands agent with its own system prompt, tool set, and guardrails. The pipeline is hardened: a malformed profile cannot reach PolicyEngine. Every agent call is wrapped in a timeout guard. The Recommendation Agent calls an estimate_cliff_effect tool to detect benefit cliffs. The Monitor Agent uses two tools — get_profile_history and check_policy_change — to distinguish a policy rule change from a personal situation change, producing fundamentally different notifications.

Profiles persist to DynamoDB as native maps (queryable fields, not JSON strings) with a 90-day TTL. The architecture is stateless at the agent level and stateful at the data layer. The frontend is Streamlit — chat intake with state dropdown, results dashboard, What If simulator, confidence indicator, and Monitor Agent demo in a single app.

The eligibility engine is PolicyEngine, an open-source microsimulation library. It runs as a direct Python call — deterministic, no LLM involved — making results authoritative and never parsed from free text. State-specific LIHEAP thresholds and program names are maintained for all 50 states.

We validated AidRadar against three diverse household profiles — a single mother in Texas, an elderly couple in Florida, and a disabled veteran in California — cross-checking results against USDA SNAP tables, HHS FPL guidelines, and program eligibility rules. All three profiles produced correct results across all eight programs.

We have 117 unit tests covering intake validation, eligibility edge cases, cliff effect detection, monitor diffing, and pipeline orchestration — plus 16 integration tests hitting the real PolicyEngine and DynamoDB, and an eval suite across 10 household profiles.


Challenges we ran into

PolicyEngine integration was the biggest technical challenge. The library expects a precise schema — adults with per-person income, children with exact ages, boolean flags wired to specific field names like is_ssi_disabled (not just is_disabled). Getting the intake agent's natural language output to reliably produce a profile that PolicyEngine could consume without errors required a dedicated guardrails layer with normalization, bounds checking, PII scrubbing, and a series of edge case guards — elderly headcount injection, citizenship normalization for DACA/refugee/TPS statuses, child age clamping, and an adults=0 safety net.

50-state expansion required more than adding state codes. Each program needed verified state-specific application URLs, state program names (Colorado calls LIHEAP "LEAP"; New York calls it "HEAP"), and state-specific income thresholds. The guardrails layer, intake prompt, and all program JSON datasets were updated consistently across all 50 states.

Strands Agents SDK was brand new to me. Understanding when to give an agent a tool versus calling a function directly from the pipeline runner — and why that separation matters for reliability — took real iteration. The key insight: PolicyEngine results should never be parsed from LLM text. The eligibility checker runs outside the agent; the agent only interprets results.

Amazon Bedrock AgentCore was the intended deployment target — it's the right architectural choice for a production multi-agent system and would have strengthened the infrastructure story significantly. I raised an AWS service quota increase request 3 weeks before the hackathon to lift the agent creation limit in my account, but it hadn't been approved by submission time. The pipeline is architected to migrate to AgentCore with minimal changes — the agent separation and tool boundaries are already clean.


Accomplishments that we're proud of

  • Real math, not keyword matching. AidRadar uses PolicyEngine's actual microsimulation — the same engine governments use — not a lookup table or LLM guess. Results are defensible and accurate, validated against ground truth across three household types.
  • All 50 states, verified. State-specific application portals, program names, and LIHEAP thresholds for every US state — not a generic federal fallback.
  • Transparent confidence scoring. AidRadar tells you not just what you qualify for, but how confident it is and exactly why — income approximation, missing citizenship status, defaulted child ages. Transparency builds trust.
  • Cliff effect detection. Most benefit finders tell you what you qualify for. AidRadar tells you when a raise could cost you more than it's worth — and by how much.
  • Monitor Agent that actually differentiates. It distinguishes whether an eligibility change was caused by a federal rule update or the user's own situation changing. Those two things warrant completely different responses, and the agent produces them.
  • Production-quality architecture in 48 hours. Guardrails, timeout protection, DynamoDB with queryable native maps, an eval suite, and integration tests — not just a working demo.

What we learned

  • Separation of concerns between agents and tools matters enormously. Giving PolicyEngine results to an agent in its prompt rather than as a tool call means the math is always authoritative. The agent reasons about results; it doesn't produce them.
  • Guardrails are not optional. The gap between "the LLM said the right thing" and "PolicyEngine got valid input" is where most errors live. A dedicated validation layer with typed errors and user-facing recovery messages is what makes the pipeline reliable rather than fragile.
  • Microsimulation is surprisingly accessible. PolicyEngine's Python API is well-documented and fast. The hard part isn't calling it — it's knowing exactly what schema it expects and building the intake pipeline to reliably produce it.

What's next for AidRadar

AidRadar already covers all 50 states and eight federal programs. The next stage adds a mobile-first React frontend, real SMS and email notifications via Amazon SNS, user accounts, and coverage across fifteen or more programs. Long-term: multilingual support (Spanish, Mandarin, Vietnamese), pre-filled application forms, and an open API layer that government portals and nonprofits can plug into directly.

$60 billion goes unclaimed every year. AidRadar finds it.

Built With

Share this project:

Updates

Submission history