Inspiration

Every month I do the same ritual: open my inbox, find the electricity bill, check the amount looks normal, log in, pay. Then broadband. Then the gym I haven't been to since June. Then wonder why my streaming bill went up. None of these is hard — they're just constant, and the expensive part is what slips through: the price hike I didn't notice, the trial that started charging, the "URGENT overdue balance" email that was actually a scam.

The Agents for Humans brief said it exactly: build something that runs in the background and only surfaces when there's a real decision. Bills are the purest version of that problem. Ninety percent of them need no decision at all. The other ten percent need a human — fast.

What it does

Ledger runs on a daily schedule and reads the inbox for bills, renewals, trial notices and payment demands. For each one it looks up the payee's history and runs a policy check. Routine bills — known payee, amount within ±10% of the trailing average, no duplicate — get paid silently. Anything else becomes a single, short question delivered to your phone:

AirFiber sent an updated invoice at ₹1,499 — 50% above your usual ₹999. Pay / Hold / Dispute?

It catches price hikes, duplicate invoices, subscriptions you haven't used in 60+ days, free trials about to convert, and urgent payment demands from senders it's never seen (probable phishing). Then it writes a five-line digest and goes back to sleep.

How we built it

  • Strands Agents SDK for the agent loop. Every capability is a plain Python function with a @tool decorator — the docstring and type hints become the schema, so the whole tool layer took minutes, not hours.
  • Claude on Amazon Bedrock as the model, doing the reading, extraction and judgement about what to ask.
  • A deterministic policy tool that holds all payment authority. This was the most important design decision: the LLM can call schedule_payment only after evaluate_policy returns auto. The model never gets to decide, on its own, that money should move.
  • Amazon SNS for delivering decisions, EventBridge Scheduler for the daily trigger, and an Amazon Bedrock AgentCore Runtime entrypoint for deployment. Strands' built-in OpenTelemetry tracing sends the agent's reasoning trail to CloudWatch.
  • A seeded mock inbox and payee history so the demo is deterministic, with the Gmail API as the drop-in replacement for the real thing.
@tool
def evaluate_policy(sender_domain: str, amount: float, due_date: str) -> dict:
    """Deterministic policy check. Returns verdict 'auto', 'ask' or 'block'.
    The agent MUST call this before paying anything."""

Challenges we ran into

  • Duplicate reminders. The first version happily paid the electricity bill twice when a reminder email arrived. The fix was a duplicate window in the policy engine — a good reminder that "the model will figure it out" is not a payment strategy.
  • Deciding what deserves a ping. Too many questions and the agent is just a noisy inbox; too few and it's dangerous. The ±10% variance rule and the 60-day unused threshold came from asking "would I want to be interrupted for this?"
  • Time. This was built in a single sitting, so every hour went to the thing the judges will actually see: a clean end-to-end run, not a polished UI.

What we learned

  • Strands makes the agent the easy part. The real work is in the tools and the policy — deciding what the agent is allowed to do, not what it can do.
  • Separating judgement (LLM) from authority (deterministic code) made the whole system easier to trust and much easier to demo.
  • The phishing case was a surprise: the agent flagged an urgent demand from an unknown sender as suspicious before we'd written a rule for it.

What's next for Ledger

Gmail OAuth ingestion, a real payments rail, a reply webhook so "Pay" on your phone resumes the paused action, AgentCore Memory for per-user policy that learns from your answers, and a monthly "here's what I saved you" digest.

AirFiber sent an updated invoice at ₹1,499 — 50% above your usual ₹999. Pay / Hold / Dispute?

It catches price hikes, duplicate invoices, subscriptions you haven't used in 60+ days, free trials about to convert, and urgent payment demands from senders it's never seen (probable phishing). Then it writes a five-line digest and goes back to sleep.

How we built it

  • Strands Agents SDK for the agent loop. Every capability is a plain Python function with a @tool decorator — the docstring and type hints become the schema, so the whole tool layer took minutes, not hours.
  • Claude on Amazon Bedrock as the model, doing the reading, extraction and judgement about what to ask.
  • A deterministic policy tool that holds all payment authority. This was the most important design decision: the LLM can call schedule_payment only after evaluate_policy returns auto. The model never gets to decide, on its own, that money should move.
  • Amazon SNS for delivering decisions, EventBridge Scheduler for the daily trigger, and an Amazon Bedrock AgentCore Runtime entrypoint for deployment. Strands' built-in OpenTelemetry tracing sends the agent's reasoning trail to CloudWatch.
  • A seeded mock inbox and payee history so the demo is deterministic, with the Gmail API as the drop-in replacement for the real thing.
@tool
def evaluate_policy(sender_domain: str, amount: float, due_date: str) -> dict:
    """Deterministic policy check. Returns verdict 'auto', 'ask' or 'block'.
    The agent MUST call this before paying anything."""

Challenges we ran into

  • Duplicate reminders. The first version happily paid the electricity bill twice when a reminder email arrived. The fix was a duplicate window in the policy engine — a good reminder that "the model will figure it out" is not a payment strategy.
  • Deciding what deserves a ping. Too many questions and the agent is just a noisy inbox; too few and it's dangerous. The ±10% variance rule and the 60-day unused threshold came from asking "would I want to be interrupted for this?"
  • Time. This was built in a single sitting, so every hour went to the thing the judges will actually see: a clean end-to-end run, not a polished UI.

What we learned

  • Strands makes the agent the easy part. The real work is in the tools and the policy — deciding what the agent is allowed to do, not what it can do.
  • Separating judgement (LLM) from authority (deterministic code) made the whole system easier to trust and much easier to demo.
  • The phishing case was a surprise: the agent flagged an urgent demand from an unknown sender as suspicious before we'd written a rule for it.

What's next for Ledger

Gmail OAuth ingestion, a real payments rail, a reply webhook so "Pay" on your phone resumes the paused action, AgentCore Memory for per-user policy that learns from your answers, and a monthly "here's what I saved you" digest.

Built With

Share this project:

Updates

Submission history