Inspiration
Money leaks in slow motion. A Netflix hike you never noticed, a gym you stopped using in April, a "free trial" that quietly started billing $12.99/month, a late fee nobody disputed. Each one is small — together it's hundreds of dollars a year and hours of attention: statements to review, calls to make, the same emails written again and again.
The hackathon brief described an agent that "runs quietly in the background and only pings you when there's a real decision to make." We realized money admin is the perfect domain for that: budgeting apps show you the problem, but the fix is always the same handful of judgment-heavy emails — cancel this, dispute that, negotiate the rate back. That's agent work. The human is only needed for the decision.
What it does
SubSentry is a background agent that guards your recurring payments end to end:
- Ingests bank and card statement exports (CSV/PDF), dedupes by hash, and normalizes merchants ("POS DEBIT NETFLIX.COM 866-579-7172 CA" → Netflix)
- Detects with deterministic math: recurring series, price hikes, zombie subscriptions (via usage exports), trial conversions, late fees, upcoming bills
- Triage and act with a Strands multi-agent graph: an analyst separates MATTERS from NOISE, an advisor verifies against the ledger and drafts ready-to-send emails (cancel / negotiate / fee-waiver / dispute), a steward files all-clear reports on quiet cycles
- Asks you only for real decisions: every outbound email pauses at Strands'
HumanInTheLoopintervention and lands in the dashboard Approval Inbox — one click to approve or reject. Approved emails default to.emldrafts, so nothing ever sends unless you opt into SMTP - Shows it all in a live dashboard: money-saved stats, findings feed, subscriptions, and an "Ask SubSentry" chat behind the same gate
On the seeded demo, one cycle found five issues and projected $414.54/year in savings plus a $25 late-fee recovery — from statements the user never had to open.

How we built it
- Python 3.13 +
strands-agents1.54:Agentwith typed@tools,GraphBuilderwith conditional edges (analyst → advisor only when something matters; analyst → steward when quiet),FileSessionManagerfor auditable transcripts and chat memory - Human-in-the-loop as a framework feature: the advisor's only outbound tool sits behind Strands' vended
HumanInTheLoopintervention. Itsaskcallback writes the approval to SQLite and parks the agent run on an asyncio future; the FastAPI dashboard resolves it across threads. The gate lives inside the agent loop — there is no code path where the agent acts unilaterally - Provider-agnostic models: Amazon Bedrock (Claude) by default, local Ollama fallback with one config switch, so the demo runs anywhere
- A deterministic core: recurrence detection, anomaly flags, and money math are pure, unit-tested Python — the LLM never does arithmetic
- A dependency-free single-page dashboard (FastAPI + vanilla JS), a Typer CLI (
demo,watch,run,ingest), 20 tests on the detection engine, ruff, and GitHub Actions CI
Full architecture: docs/architecture.md · one-command demo: subsentry demo --provider ollama
Challenges we ran into
- Graph constraints: Strands told us "Session persistence is not supported for Graph agents yet" — we restructured so graph nodes stay stateless per cycle and the persistent-session chat agent carries memory instead.
- Flaky local tool calling: our Ollama fallback (Llama 3.1 8B) would sometimes narrate tool calls as text instead of invoking them, and parallel tool calls raced the interrupt state.
SequentialToolExecutorfixed the race; digging through the raw tool results revealed the real killer — the small model sometimes dropped a requiredrationaleargument, so validation failed before the tool body ever ran. We made display-only args optional and folded bookkeeping into the action tool itself (one call per finding). Bedrock Claude handles the full orchestration natively — which is why it's the default. - Wrong late-fee attribution: the engine first attached a DriveSafe late fee to the wrong subscription. The fix: attach fees to the nearest preceding charge, not the first match.
- Cross-thread approvals: resuming a paused agent from a web request required wiring SQLite-backed approval records to asyncio futures resolved via
call_soon_threadsafe— the agent waits for a human for as long as it takes. - Background-agent semantics: unattended runs need to be safe to fail. Unhandled findings are automatically requeued for the next cycle, and action keys + note-based stats make retries idempotent — the agent never double-sends or double-counts.
Accomplishments that we're proud of
- The safety story is enforced by the framework, not by convention: the agent physically cannot contact anyone without a human approval captured in the database
- A complete, honest money trail: every finding links to exact amounts and dates, and savings stats are projected from real transaction data
- The entire demo is reproducible offline in one command on fully synthetic statements — no personal data, no API keys needed
- The deterministic core is genuinely tested: 20 unit tests covering cadence fitting, hike/trial/fee/zombie detection, ingestion dedupe, and merchant normalization
What we learned
- Let math do math and LLMs do judgment. Asking the model to find subscriptions from raw transactions was slow and hallucination-prone; moving detection into deterministic code made the agent faster, testable, and honest.
- Small models fail in instructive ways. Every 8B-model failure (narration, dropped arguments, parallel-call races) pushed us toward better tool ergonomics — which made the Bedrock path better too.
- Strands' intervention system is production-grade human-in-the-loop. Bridging
askto a real UI gave us "autonomous, but only with your approval" without fighting the SDK. - Background agents are a reliability problem, not a prompt problem. Requeue loops, idempotent actions, and dedupe matter as much as the model.
What's next for SubSentry: your bills, handled. Your attention, respected.
- Email ingestion so statements arrive and process themselves (no manual exports)
- Bank-aggregator connections (Plaid-style) as an alternative to statement files
- Deploying to Amazon Bedrock AgentCore Runtime with
S3SessionManagerfor multi-instance state - Confirmation tracking: when a merchant replies "you're cancelled," mark the subscription resolved automatically
- Household mode: one guardian for the whole family's accounts
Built With
- amazon-bedrock
- amazon-web-services
- claude
- fastapi
- github-actions
- human-in-the-loop
- javascript
- multi-agent
- ollama
- pydantic
- python
- sqlite
- strands-agents
- typer
Log in or sign up for Devpost to join the conversation.