Inspiration

I run a one-person business. I assumed my admin time went to doing things. It doesn't — it goes to deciding the same things again. Which project does this coffee receipt belong to. Is this inside the retainer or extra. Do I chase this invoice today. Twenty seconds each, thirty a day, plus the context switching. That is an hour gone on judgments I had already made once. I call it re-decision cost. Automation tools don't touch it, because they try to make the decision for you. I wanted the opposite: something that remembers the decision I made and stops asking.

What it does

Standing Order watches the routine admin stream of a one-person business — transactions, invoices, subscriptions, calendar — and handles each item using the standing orders you already gave it. A standing order is a plain sentence you approved once: "Food and drink under ₩50,000 with a client meeting on the calendar goes to entertainment expense on that project."

Three things on screen, none of them a chat box:

  • Interruption budget. The agent gets three interruptions a day. What it can't decide competes for those slots on a deterministic priority score; the rest waits for the daily digest.
  • Standing orders book. Every rule in plain language, with version, effective date, and the queue item it came from.
  • Receipt log. Every item handled without asking, with the rule ID that authorized it and the tool calls that ran.

The metric is not how much it automated. It is how often it interrupted you — and that number goes down. Over the 60-day replay, daily interrupt demand falls from 4 to 1 (first-10-day average 2.7 → last-10-day average 0.9), with 27 approved standing orders covering 75.3% of decision-bearing events by day 60. The replay itself makes zero model calls and completes in 7.7 seconds.

How we built it

Strands Agents is the whole runtime, but the interesting part is what the model is not allowed to do.

Matching is a pure Python policy engine. Rules are Pydantic specs compiled into predicates — no eval, no model call. A test blocks the socket layer and runs the full 60-day replay to prove the decision path never touches the network. The model appears in exactly two places: proposing a generalized rule from one human answer via structured_output, and writing the one-line explanation on a receipt.

The invariant is enforced with Strands hooks. Before any tool call executes, a BeforeToolCallEvent hook checks the tool name against the scope of the standing order that authorized this item. Out of scope means the call does not run — not "the prompt asks it not to." The demo shows the blocked-call counter going up, and the test asserts the database is untouched after a blocked call.

The pipeline is a Strands multi-agent Graph — ingest, match, act, escalate, learn — with most nodes containing no LLM at all, and a SessionManager so the background loop survives a restart. The rule compiler also runs as a deployed Bedrock AgentCore runtime (ARM64 container implementing the /invocations contract directly).

Challenges we ran into

The rule compiler generalized too far: one answer about one café became "all food is entertainment." The fix was not a better prompt. Every candidate rule is now replayed against the previous 60 days and shows how many past items it would have decided; anything above a threshold is flagged before you can approve it. Rules also never apply retroactively — they carry an effective date.

Deployment on a brand-new AWS account was its own jungle: App Runner stopped accepting new customers in April 2026 (our fallback assumption silently died), Lambda Function URLs returned 403 even to signed requests (a fresh-account restriction — ConcurrentExecutions capped at 10 was the tell), so the live demo runs behind an API Gateway HTTP API instead. The AgentCore starter toolkit needed CodeBuild permissions we didn't have, so we implemented the runtime contract directly — ARM64 image, /invocations, /ping — and had it READY in about 40 minutes. Mid-build, Bedrock revoked Anthropic model access pending a use-case form; it turns out aws bedrock put-use-case-for-model-access submits it from the CLI.

The dumbest one: hook event class names differ between Strands versions, and we lost time trusting documentation instead of printing dir(strands.hooks) on the installed package.

What's next

Bank and accounting connectors so the input stream is real instead of seeded. Multi-user with per-user rule books. A conflict-resolution screen for when two standing orders both match.

Not in scope (today): no bank or accounting API integration — input is seeded data, CSV, or webhook. No tax advice. Single user only. It never moves money. The 60-day ledger is generated by simulation, not a real production log.

Built With

  • amazon-bedrock
  • api-gateway
  • aws-lambda
  • bedrock-agentcore
  • claude-sonnet
  • docker
  • fastapi
  • nextjs
  • python
  • sqlite
  • strands-agents
Share this project:

Updates

Submission history