posted an update —

Build log — receipts, not claims

Built in a two-day sprint. Every phase below is timestamped in the repo — commits, PR review threads, and CI runs are the receipts.

The spike. Proved the load-bearing pieces with the Strands SDK: a real agent stopping at a typed interrupt, session persistence across processes, seeded resume, and artifacts validating against open HACP schemas. Four design rulings came out of a governed debate between two different AI models (blind first passes, severity-ordered findings, one human ruling each motion).

The spine. Deterministic artifact pipeline: task packet → review findings → typed stop → human decision → consumption receipt → dry-run effect → agent report. Tampered artifacts fail validation by design. Commit-or-pivot gate on Bedrock passed at ~$0.05/run against a $5 ceiling.

The console. Three-state decision console over HTTP: nothing preselected, rationale required, the server is the source of truth. Survived a reviewer loop that caught (and we fixed) real correctness bugs — including non-approval branches initially executing the approval path. That loop is why the demo's core promise holds for all three choices.

The live proof + hardening. A real Bedrock model ran the full loop end to end: one interrupt, resume on the recorded decision, competing replay rejected. The consumption store got domain-separated digests, a multi-process contention proof (now in CI), fail-closed reservations, and power-loss durability. Eleven review rounds total — every finding public in the PR threads.

CI runs the whole verification stack on every push: https://github.com/joefeser/who-decides/actions

Log in or sign up for Devpost to join the conversation.