-
-
The decision card: nothing preselected, a written rationale required. Nothing executes until the human decides.
-
After approval: the decision claimed exactly once, exact dry-run payload, all artifacts schema-valid, duplicate replay REJECTED.
-
The loop: bounded job → prepare + verify → typed stop for the human → execute only the approved branch, claimed exactly once.
-
Live Bedrock run: one interrupt, resume on the recorded decision, competing replay rejected — nothing staged.
Inspiration
We spend our days in a human-gated engineering workflow: AI agents draft, review, and implement — but a human decides what actually ships. After months of that, one pattern kept standing out: the interesting moment in any agentic system isn't what the agent did, it's the transfer of judgment. When does the machine decide, and when does the human?
What it does
The agent works a bounded job in the background — a dependency security update with tests green and one unresolved tradeoff. When it reaches a decision that belongs to a human, it emits a typed decision request and halts. You get the question, the evidence, and the tradeoff in plain terms; you choose and write your rationale. The agent resumes from its persisted session, executes only the branch you approved, produces a receipt proving your decision was consumed exactly once, and shows the exact payload — dry-run, no external mutation. A duplicate replay attempt is rejected on screen.
Most agent demos optimize for autonomy — let it run, let it act, review the aftermath. We wanted the inversion: an agent that works in the background and surfaces only when there's a real decision, then stops and waits.
So we built the entry the same way it works. Its four core design decisions were made in a governed debate between two different AI agents — one GLM, one GPT — posting blind first passes (no peeking at each other's positions), severity-ordered findings, and typed stop reasons, with one human ruling each motion. 27 receipted messages, four human rulings, zero agent-held authority. The project's design record is itself the demo of its thesis: who decides is the only question that matters.
What we learned
- Blind review works. On one motion both agents independently converged on the same recommendation and the same top-severity finding. On another they genuinely diverged — and the exchange resolved it on a structural argument, not taste, with one agent conceding on the record.
- Typed stops prevent real waste. During setup, a prompt landed in the wrong agent session; the agent refused with
WRONG_TOOL_OR_MODEinstead of improvising. Wrong work that never happens is the cheapest work. - Approval theater is a design failure. If the agent surfaces a decision with all-green checks, "approve" is a rubber stamp. The demo deliberately seeds an honest tradeoff the agent cannot own.
- Claiming ≠ doing. An atomic consumption receipt proves one decision authorized one resume — it does not prove the effect happened exactly once. Saying that plainly in the spec was more valuable than over-promising.
How we built it
Built on the Strands Agents TypeScript SDK, who-decides runs a dependency-security update as a bounded task packet: inspect, update, test, build, gather evidence. When it reaches a judgment call — a security patch whose compatibility check is genuinely unresolved — it terminates with a typed decision request validated against open HACP JSON Schemas. No process hovers, no effect occurs.
The human approves or rejects in a console with three core states (running → decision required → resumed) plus typed blocked outcomes for failures, with wording that tells the truth: "Decision required — waiting for you", never "paused," because nothing is paused — the run ended. A new invocation resumes seeded with the prior session plus the recorded decision, and an immutable consumption receipt atomically binds that decision to exactly one successor invocation — two concurrent resumes can never both succeed.
Every effect defaults to dry-run: the receipt carries the exact payload it would execute, with an explicit "no external mutation performed." The model layer is a thin provider abstraction — Amazon Bedrock by default, any OpenAI-compatible endpoint as an escape hatch — so the demo is reproducible either way.
The agent is deployed, not just local: packaged as a Node 22 ARM64 container on AWS AgentCore Runtime (us-east-1), authenticated by a scoped machine-principal service token, and invoked by the console through the AWS SDK with no silent fallback. A run advances only when the deployed runtime confirms the result — including proof that the executed effect matches the approved choice. A strict live gate exercises the full A → decision → B cycle against the deployed runtime through the production wiring: 6/6 passing, with duplicate-replay, choice-conflict, and rejection discipline on typed dispatches.
Challenges we faced
- Resume semantics. Mid-run pause depended on unproven continuation primitives, so we committed to terminal-request + seeded resume — then discovered the SDK ships human-in-the-loop interrupt and checkpoint primitives mid-spike, and re-evaluated honestly against our own bar: does it survive process restart, restore exact state, and consume decisions exactly once?
- Exactly-once consumption. An idempotency key deduplicates requests; it doesn't stop one decision authorizing two competing resumes. The fix — a separate immutable receipt with an atomic claim — took a real design debate to get right.
- Closed schemas. HACP's schemas reject unknown properties, so extensions can't just bolt fields on. The demo publishes a separately versioned extension profile, and our engine gates execution on those records — a bare decision artifact alone authorizes nothing.
- The crowded room. Agent PR-reviewers are a commodity. We chose dependency-release judgment instead, where the decision gate is the product, not the output format.
Built With
- ai-agents
- amazon-bedrock
- amazon-web-services
- aws-agentcore
- ffmpeg
- gemini
- glm
- human-in-the-loop
- json-schema
- next-js
- node.js
- react
- sqlite
- strands-agents-sdk
- tailwindcss
- typescript
- z-code
Log in or sign up for Devpost to join the conversation.