Inspiration

Put AI agents on real work and output goes up sharply. So does a failure mode people rarely produce: an agent reports "done," confidently, having made something plausible and wrong. A slugify that forgot to lowercase. A launch email missing its unsubscribe line. A refund policy that never names the window. Every project tracker records that as success, because "done" is a status a worker sets. That was defensible when workers were people who feel embarrassment. It is not defensible now. We wanted the opposite default: an agent should not be able to mark its own work done.

What it does

You describe what done means, an agent produces the work, and Conductor delivers it only after a deterministic check passes. The work can be code, a document, a policy, or a report: anything where "done" is a specific, checkable claim.

  • Code: point it at a real GitHub repo and it opens a draft PR only after the check passes, with the check run isolated in AWS CodeBuild so a stranger's check never touches the host.
  • Documents, no code: describe the bar in plain words ("must include an unsubscribe line and name the product"). Conductor compiles that into a real check you approve, an agent drafts, and you get the verified document in the app with no repo and no PR.
  • A manager layer a person actually uses: work ranked by value, a throughput forecast, blockers, team health, and a stakeholder update generated from real state. It surfaces only the one decision that genuinely needs a human.

How we built it

Conductor is a Strands Agents multi-agent system on Amazon Bedrock: a Planner, Worker, Recovery, Compressor, and Orchestrator that propose, and deterministic code (a verification runner, policy gate, attention and trust ledgers) that disposes. No agent stage can reach "done."

The signature piece is the proof hook: a Strands HookProvider subscribes to BeforeToolCallEvent, and when a worker agent calls its own claim_done, the hook runs the commitment's real check on the spot and, on failure, sets cancel_tool to veto the call and hand the agent the reason to keep working. The SDK itself makes a dishonest "done" impossible.

State is event-sourced and replayable, so trust earned and lost survives a restart. It is deployed on Amazon Bedrock AgentCore Runtime and AWS App Runner, with a provider fallback and a backup key so a live run survives a model throttle, and full OIDC/JWKS auth so it never handles a password.

Challenges we ran into

  • A relative-import bug in the proof hook crashed every real-agent run and only surfaced in live testing, not in unit tests. A reminder that the demo has to be run for real, not just tested.
  • A connected-repo check first ran on the host. That was a remote code execution risk, so we moved every untrusted check into an isolated AWS CodeBuild sandbox.
  • Free model keys throttle fast, which is exactly why the provider fallback and backup key exist.
  • Making it usable by non-engineers: the guarantee originally only reached people who could write a shell check, so we built "checks without code" and an in-app delivery surface.

Accomplishments that we're proud of

  • An agent genuinely cannot mark its own work done, enforced inside the Strands loop, not just around it.
  • The same guarantee works for code, content, and compliance, each proven live with a real artifact.
  • Untrusted checks run sandboxed in AWS CodeBuild, so running a stranger's check is safe.
  • Event-sourced trust that survives restarts, a real manager layer, and 225 passing tests plus agent evals in CI.
  • An honest product: it proves the work passes your check, and says plainly what it does not do.

What we learned

The hard part of agents isn't making them produce; it's trusting what they produced. The bottleneck moves from doing the work to checking it, and the scarce resource becomes human attention. The honest framing matters most: Conductor proves the work passes your check, never that the check is the right check. So plans are graded for how checkable they are, and the human owns every check and every priority. That honesty is the product, not a caveat.

What's next for Conductor

  • More delivery surfaces beyond a repo: an approval inbox and connectors to Google Docs, Notion, and Sheets so content and compliance teams work where they already are.
  • A richer library of plain-words checks per domain (required clauses, link-resolves, schema-validates, metric-in-range).
  • Live metric sources for the outcome tier, so "done only when the metric holds" runs against real analytics.
  • Deeper enterprise foundations: RBAC, audit log, SSO hardening, and regional deployment for lower latency worldwide.

Built With

Share this project:

Updates

Submission history