Inspiration
One of us paid £456 for a gym cancelled in March. Not because the money didn't matter — because cancelling was a four-minute job, and four-minute jobs are exactly the ones that never get done.
The average household runs about twelve subscriptions and eight recurring bills. A price goes up. A trial converts. A charge lands twice. An appointment needs moving. Individually minor; together, hours a month you never find.
The usual answer is an app. But an app just moves the work — now you have to remember to open the thing that reminds you.
And the usual answer from agents is worse. Most agents ask you about everything, which is just a more talkative to-do list. An agent that needs supervising has not saved you the work; it has changed the format of the work.
So we built the opposite. Quiet Hours earns the right to stop asking.
What it does
Quiet Hours runs on a schedule. There is no app to open and no chat box — the web app is a decision inbox, and on a good day it is empty.
- Ingest — new email, card transactions and calendar events since the last run.
- Triage — a Strands multi-agent graph classifies and correlates them into findings: a price increase, a converting trial, a duplicate service, an unused subscription, a double charge.
- Specialists act — a bill analyst compares against twelve months of history, a negotiator drafts cancellations and disputes, a scheduler handles calendar work.
- The policy gate — every proposed action passes a deterministic gate before it runs. Reversible, low-risk work executes silently and is logged with the reasoning that produced it. Anything that spends money, sends a message or cancels a service raises a decision card and the run suspends.
- You decide — one tap, hours or days later. The agent resumes the same session exactly where it stopped and finishes the job.
The part that makes it different: autonomy is earned. Answering "always" creates a policy — and the card shows you the rule in plain English before you grant it: "Always cancel subscription for FitLife up to GBP 38.00." The policy engine then handles that class of action silently on every future run.
Over four simulated weeks the interrupt rate falls from four decisions in week one to one in week four — 75% fewer interruptions — while the number of actions handled without asking rises from three a week to six. In the run you can open right now, 19 of 26 actions were handled silently and £411.90 was saved.
That trend line is the product. Every policy is listed, counted and revocable in one click, and some things — disputing a charge, closing an account — no policy can ever auto-approve. That is enforced in code, not requested in a prompt.
How we built it
The agent is the Strands Agents SDK doing real work, not a wrapper around one prompt:
| Strands feature | What it does here |
|---|---|
GraphBuilder |
The triage → specialists → brief pipeline, six agent nodes |
HookProvider on BeforeToolCallEvent |
The policy gate. Every tool call is governed — including ones the model improvises — so safety is structural |
event.interrupt() |
Raises the decision card and suspends the run |
Durable SessionManager |
The suspended run survives process exit and resumes days later where it stopped |
structured_output_model |
Typed Finding and ProposedAction objects, not JSON scraped out of prose |
invocation_state |
Household context reaches tools and hooks without entering the model's context window |
Every tool carries a RiskTier — SILENT, NOTIFY, CONFIRM, NEVER_AUTO — and the floor comes from the action kind, not from the tool's own say-so, so a new tool cannot quietly grant itself permission. The policy engine that applies learned rules is deliberately ordinary code with no model call in it: it cannot be argued out of a rule, and it cannot downgrade NEVER_AUTO.
Around it: FastAPI + Pydantic v2 for the API, Next.js 15 and TypeScript for the decision inbox, DynamoDB for decisions and runs, S3 for Strands session state, SES for the weekly digest. The wire format is snake_case end to end with generated TypeScript types, so the contract between the three of us could not drift.
Three people, three lanes, three weeks, each with an AI coding assistant, sharing a frozen /contracts directory and a per-lane AGENTS.md. 307 tests run in CI — 214 agent, 50 integrations, 43 API — plus a guardrail job that fails the build on a cross-lane edit.
Money is integer minor units everywhere. Never floats.
Challenges we ran into
We could not get Amazon Bedrock working, and we are not going to pretend otherwise. Every model on our AWS account — Claude and Amazon's own Nova alike, from the SDK and from the console playground — returns ValidationException: Operation not allowed. Our support case has been open and unassigned since 31 August. We eventually found two causes: new AWS accounts default to a Free plan that excludes Bedrock cross-region inference, and the account itself is still stuck in AWS's verification hold, which blocks provisioning account-wide. We upgraded the plan; the hold outlasted the deadline.
So the demo runs in simulation mode: the Strands graph, the hook dispatch, the interrupt, the session suspend-and-resume and the tool execution are all genuine SDK machinery — only the model provider is swapped for a scripted one. Nothing is deployed to AgentCore. The Bedrock and AgentCore code paths are written and tested, and have never been run against the live service. That is an honest cap on our technical score and we would rather state it than have it found.
What that forced us to get right. Because we could not lean on a model to make the demo impressive, everything the product claims had to be true structurally. The policy gate is deterministic code. The autonomy curve is counted from the audit trail of four real runs, not authored. The hosted demo re-implements the API in the browser so a judge needs no credentials at all.
Three other things bit us. A Strands minor release moved the half-finished tool call from interrupt_state["context"]["tool_use_message"] to interrupt_state["pending_tool_execution"]["assistant_message"], and our interrupt probe went red in CI while every laptop stayed green — the probe now accepts both spellings. Our CI was reporting green for weeks behind a pytest … || true; removing that mask immediately caught a real dependency break. And we found and removed three overclaims from our own docs, including a "deployed on AgentCore" line for something that is not deployed.
Accomplishments that we're proud of
- The interrupt genuinely survives process exit. The agent stops mid-tool-call, the process ends, and days later a tap in the browser resumes the same session and completes the action. It is tested on all three answer paths.
- The autonomy curve is measured, not asserted. Both lines are counted from the audit trail.
- Safety that is structural rather than prompted. A hook on every tool call, a risk floor keyed to the action, and a tier no policy can downgrade.
git clone && make demoworks with zero credentials, and the hosted demo loads for a stranger with no login.- Nothing in this submission is oversold. The architecture diagram draws deployed components as dashed boxes. That discipline cost us points and we kept it anyway.
What we learned
The interrupt is a product decision before it is a technical one. Strands gives you event.interrupt(); what it does not give you is the judgement about when an agent has earned the right not to use it. Getting that wrong in either direction ruins the product — ask about everything and you have built a to-do list; ask about nothing and nobody will connect their bank.
Determinism is a feature you can show a user. The most persuasive thing in the whole demo is that the rules page is boring: plain English, counted, revocable, and applied by code that cannot be talked round.
Model access is infrastructure, and infrastructure fails early. We treated "get Bedrock working" as a task for week three. It was the one thing we could not solve by working harder.
Freeze the contracts. Three people and three AI assistants in one repo for three weeks, and the shared /contracts directory froze on 24 August. Not one merge conflict came from a payload shape.
What's next for Quiet Hours
- Get it onto Bedrock and AgentCore Runtime. The code path exists and is tested; it needs an account that will serve a model. This is the first thing we do.
- Real providers behind the integration interfaces — Gmail, an open-banking feed, CalDAV — behind the same risk tiers, still drafts-only and scheduled-payments-only.
- Policy suggestions from observed patterns, always proposed and never self-granted: "You have approved this four times — want a rule?"
- Shared households, where a decision can be routed to whichever person actually owns that bill.
- Widen the trust model beyond money: renewals, warranties, council and school admin — the same earned-autonomy loop.
Built With
- amazon-bedrock
- amazon-ses
- amazon-web-services
- bedrock-agentcore
- dynamodb
- fastapi
- github
- github-actions
- next.js
- pydantic
- python
- react
- strands-agents
- tailwindcss
- typescript
Log in or sign up for Devpost to join the conversation.