Inspiration
Spoke & Sprocket is a two-person mobile bike-repair shop. Between jobs, its request inbox fills from three channels at once — a web form, a shared inbox, and SMS — with bookings, reschedules, cancellations, questions, refunds, and the occasional damage claim, all mixed together. Sorting them by hand costs an hour a day that nobody has, and the cost of answering the wrong one badly is a customer, not a support ticket.
What it does
Inbox to Done is a Strands Agents CLI that clears that inbox. For each request it classifies the message, drafts a reply, proposes and books an appointment slot where one is needed, and files the routine ones itself. Out of a seeded 14-request inbox it files ten and hands back four.
The four it hands back are the point. A refund at or above the shop's $25.00 auto-limit, any dispute or damage claim at any amount, anything classified ambiguous, anything below a 0.70 confidence floor, and anything no policy rule covers are never filed by the agent — and not merely discouraged in a prompt. The check runs inside invoke(action, args, actor), the single function every state change in the codebase goes through, so the agent's own file_routine call comes back refused, with a reason and an instruction to escalate. The audit log records the refusal alongside the successes.
Each refusal becomes a one-line decision request the owner clears from the CLI — one question, two to four concrete options, and the agent's recommendation. The owner's answer travels back through that same invoke, tagged actor="owner", and the agent does the follow-through work itself. One log, two actors, one code path.
Nothing is ever sent: drafted replies sit in out/outbox/ until the owner releases them, and no email, SMS, or payment provider is wired up.
How we built it
A Strands agent with ten @tool functions runs over the seeded inbox. Four are read-only. The other six are thin wrappers around one function, core.invoke(action, args, actor), which evaluates the shop's policy, writes to the store only if policy allows it, and appends an audit entry either way. The owner CLI calls that same function with actor="owner" — there is no second write path, and no argument the model can pass makes it the owner.
Built with the Strands Agents SDK, Amazon Bedrock (the default model provider; the exact model id is deliberately left unpinned and resolved against the live catalog at the cheapest Claude tier, so cost stays a code guarantee rather than documentation), Python 3.12 with uv, and a 33-test pytest suite that runs entirely against a scripted offline provider — no AWS credential needed to verify the plumbing works. Ollama is wired in as an optional local, no-spend provider.
Challenges we ran into
Keeping the actor honest was the real design problem. actor="agent" is passed only from the tool wrappers the model calls; actor="owner" is passed only from the CLI it cannot reach — there's no argument the model can hand invoke that makes it the owner, which is what makes the audit log's actor column a real distinction rather than a label the agent could forge.
The architecture diagram also documents two places where the build honestly departs from the original spec, rather than papering over them: the tool split is drawn 4 read-only / 6 mutating (not "every tool into invoke"), because four tools never touch the store at all and drawing it otherwise would put a false claim on the artifact judges look at hardest; and no AgentCore deployment box appears, because that path was deliberately not pursued in favor of shipping the local CLI first.
What's next
The README's Known Limitations section is unchanged from what a reviewer would find today: no persistence across separate CLI invocations by design (demo chains a full run in one process on purpose), and the offline demo replays a deterministic scripted model — evidence the plumbing works, not that a live model behaves identically.
Built With
- amazon-bedrock
- mermaid
- ollama
- pytest
- python
- ruff
- strands-agents-sdk
- uv
Log in or sign up for Devpost to join the conversation.