Inspiration

Every team ships on vibes. A tired reviewer squints at a green CI badge, approves a deploy, and moves on. The badge says build passed — but nobody can answer: what exactly was verified, against what evidence, for which commit, and who approved it? When a release breaks, that answer can't be reconstructed, because the review happened in someone's head.

As a solo maintainer, I'm the worst case: the person approving the deploy is the same person who wrote the code at 2am. So I built evidenza — an agent that does the verification work autonomously, but cannot ship without a human at the one moment that matters.

(Disclosure: the release-verification concept builds on ideas from my own prior open-source project, autokeren. evidenza is a fresh, lean codebase written from scratch during the hackathon on the Strands Agents SDK — no code was copied.)

What it does

Given a release task, evidenza:

  1. Captures the exact git commit being released (git_sha)
  2. Runs the release candidate's real test suite and commands (shell)
  3. Inspects manifests, configs, source (read_file)
  4. Checks live endpoints when applicable (verify_url)
  5. Records every acceptance criterion with its raw evidence in a proof artifact (JSON)
  6. Computes an evidence-led verdict — SHIP / BLOCKED / NEEDS_HUMAN_REVIEW
  7. Pauses for a human before deploy: a Strands BeforeToolCallEvent hook raises an interrupt that blocks the deploy tool until a human explicitly approves
  8. Replays — the proof artifact renders a visual Release Card with zero API keys

Two safety properties are enforced in code, not by prompt:

  • A SHIP verdict requires every criterion to be passed, computed from recorded evidence — never from the model's own assertion.
  • approve() rejects any non-SHIP verdict, and the deploy hook raises an interrupt for unapproved SHIP proofs. No auto-ship, ever.

How I built it

  • Strands Agents SDK is the core: the Agent loop, @tool-decorated tools, the conversation manager, the hooks registry, and the interrupt system that powers human approval.
  • 7 lean tools — read_file, write_file, shell, git_sha, git_commit, verify_url, deploy — each a thin, typed Strands tool.
  • Approval hook — a BeforeToolCallEvent callback that gates deploy behind human approval via InterruptException. The entire safety property is ~80 lines of hook code.
  • Proof system — a plan → record → report → approve state machine serializing proofs to JSON, plus a replay renderer for the Release Card.
  • Model-provider agnostic runtime — a pluggable Strands Model implementation, so the agent runs on any Anthropic-Messages-compatible backend (local proxy for dev/demo, Anthropic, Amazon Bedrock).
  • 28 deterministic tests covering both safety properties with zero network dependency.

Challenges I ran into

  • Streaming crashes: the Anthropic SDK's streaming accumulator choked on tool_use index ordering from my proxy backends. Fix: a custom Strands Model provider that issues a single non-streaming request and synthesizes Strands StreamEvents from the response.
  • Models hallucinating tool arguments: Strands wraps tool inputSchema under a json key; backends saw an empty schema and the model invented arguments. Fix: unwrap the schema before sending.
  • Tools never called: some backends wouldn't call tools without an explicit tool_choice. Fix: default to tool_choice: auto whenever tools are present.
  • Judging without credentials: I didn't want judges to need any API key or live model to evaluate the system. That constraint became the best feature: the proof artifact is replayable offline — evidenza --proof-replay examples/demo/proof-run.json renders the full evidence chain with nothing but Python.

Accomplishments I'm proud of

  • The human-approval gate is a deterministic, tested mechanism (InterruptException in a hook), not a polite request to the model.
  • 28/28 tests pass, covering verdict logic, proof lifecycle, replay, and the approval hook — no network needed.
  • The demo runs end-to-end reliably on video: real pytest output becomes evidence, deploy gets blocked, human approves, deploy proceeds.

What I learned

  • Strands' hooks + interrupt design maps exactly to a human-approval gate — this is the non-obvious SDK usage the best agent architectures are made of.
  • "Evidence-led" beats "model says it's fine": forcing the verdict to be computed from recorded criteria — instead of trusting model narration — surfaced every assumption I'd otherwise have shipped on.
  • Safety properties that matter should live in code with tests, not in prompts.

What's next for evidenza

  • Amazon Bedrock AgentCore deployment for a hosted, live-verifiable demo
  • CI integration (GitHub Action that posts the Release Card on every release PR)
  • Richer verifiers: security scans, performance budgets, changelog consistency

Built With

Share this project:

Updates

Submission history