Inspiration
Every month I do the same ritual: open my inbox, find the electricity bill, check the amount looks normal, log in, pay. Then broadband. Then the gym I haven't been to since June. Then wonder why my streaming bill went up. None of these is hard — they're just constant, and the expensive part is what slips through: the price hike I didn't notice, the trial that started charging, the "URGENT overdue balance" email that was actually a scam.
The Agents for Humans brief said it exactly: build something that runs in the background and only surfaces when there's a real decision. Bills are the purest version of that problem. Ninety percent of them need no decision at all. The other ten percent need a human — fast.
What it does
Ledger runs on a daily schedule and reads the inbox for bills, renewals, trial notices and payment demands. For each one it looks up the payee's history and runs a policy check. Routine bills — known payee, amount within ±10% of the trailing average, no duplicate — get paid silently. Anything else becomes a single, short question delivered to your phone:
AirFiber sent an updated invoice at ₹1,499 — 50% above your usual ₹999. Pay / Hold / Dispute?
It catches price hikes, duplicate invoices, subscriptions you haven't used in 60+ days, free trials about to convert, and urgent payment demands from senders it's never seen (probable phishing). Then it writes a five-line digest and goes back to sleep.
How we built it
- Strands Agents SDK for the agent loop. Every capability is a plain Python function with a
@tooldecorator — the docstring and type hints become the schema, so the whole tool layer took minutes, not hours. - Claude on Amazon Bedrock as the model, doing the reading, extraction and judgement about what to ask.
- A deterministic policy tool that holds all payment authority. This was the most important design decision: the LLM can call
schedule_paymentonly afterevaluate_policyreturnsauto. The model never gets to decide, on its own, that money should move. - Amazon SNS for delivering decisions, EventBridge Scheduler for the daily trigger, and an Amazon Bedrock AgentCore Runtime entrypoint for deployment. Strands' built-in OpenTelemetry tracing sends the agent's reasoning trail to CloudWatch.
- A seeded mock inbox and payee history so the demo is deterministic, with the Gmail API as the drop-in replacement for the real thing.
@tool
def evaluate_policy(sender_domain: str, amount: float, due_date: str) -> dict:
"""Deterministic policy check. Returns verdict 'auto', 'ask' or 'block'.
The agent MUST call this before paying anything."""
Challenges we ran into
- Duplicate reminders. The first version happily paid the electricity bill twice when a reminder email arrived. The fix was a duplicate window in the policy engine — a good reminder that "the model will figure it out" is not a payment strategy.
- Deciding what deserves a ping. Too many questions and the agent is just a noisy inbox; too few and it's dangerous. The ±10% variance rule and the 60-day unused threshold came from asking "would I want to be interrupted for this?"
- Time. This was built in a single sitting, so every hour went to the thing the judges will actually see: a clean end-to-end run, not a polished UI.
What we learned
- Strands makes the agent the easy part. The real work is in the tools and the policy — deciding what the agent is allowed to do, not what it can do.
- Separating judgement (LLM) from authority (deterministic code) made the whole system easier to trust and much easier to demo.
- The phishing case was a surprise: the agent flagged an urgent demand from an unknown sender as suspicious before we'd written a rule for it.
What's next for Ledger
Gmail OAuth ingestion, a real payments rail, a reply webhook so "Pay" on your phone resumes the paused action, AgentCore Memory for per-user policy that learns from your answers, and a monthly "here's what I saved you" digest.
AirFiber sent an updated invoice at ₹1,499 — 50% above your usual ₹999. Pay / Hold / Dispute?
It catches price hikes, duplicate invoices, subscriptions you haven't used in 60+ days, free trials about to convert, and urgent payment demands from senders it's never seen (probable phishing). Then it writes a five-line digest and goes back to sleep.
How we built it
- Strands Agents SDK for the agent loop. Every capability is a plain Python function with a
@tooldecorator — the docstring and type hints become the schema, so the whole tool layer took minutes, not hours. - Claude on Amazon Bedrock as the model, doing the reading, extraction and judgement about what to ask.
- A deterministic policy tool that holds all payment authority. This was the most important design decision: the LLM can call
schedule_paymentonly afterevaluate_policyreturnsauto. The model never gets to decide, on its own, that money should move. - Amazon SNS for delivering decisions, EventBridge Scheduler for the daily trigger, and an Amazon Bedrock AgentCore Runtime entrypoint for deployment. Strands' built-in OpenTelemetry tracing sends the agent's reasoning trail to CloudWatch.
- A seeded mock inbox and payee history so the demo is deterministic, with the Gmail API as the drop-in replacement for the real thing.
@tool
def evaluate_policy(sender_domain: str, amount: float, due_date: str) -> dict:
"""Deterministic policy check. Returns verdict 'auto', 'ask' or 'block'.
The agent MUST call this before paying anything."""
Challenges we ran into
- Duplicate reminders. The first version happily paid the electricity bill twice when a reminder email arrived. The fix was a duplicate window in the policy engine — a good reminder that "the model will figure it out" is not a payment strategy.
- Deciding what deserves a ping. Too many questions and the agent is just a noisy inbox; too few and it's dangerous. The ±10% variance rule and the 60-day unused threshold came from asking "would I want to be interrupted for this?"
- Time. This was built in a single sitting, so every hour went to the thing the judges will actually see: a clean end-to-end run, not a polished UI.
What we learned
- Strands makes the agent the easy part. The real work is in the tools and the policy — deciding what the agent is allowed to do, not what it can do.
- Separating judgement (LLM) from authority (deterministic code) made the whole system easier to trust and much easier to demo.
- The phishing case was a surprise: the agent flagged an urgent demand from an unknown sender as suspicious before we'd written a rule for it.
What's next for Ledger
Gmail OAuth ingestion, a real payments rail, a reply webhook so "Pay" on your phone resumes the paused action, AgentCore Memory for per-user policy that learns from your answers, and a monthly "here's what I saved you" digest.
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-cloudwatch
- amazon-eventbridge
- amazon-sns
- boto3
- claude
- gmail-api
- opentelemetry
- python
- strands-agents-sdk
- streamlit
Log in or sign up for Devpost to join the conversation.