Inspiration
Every team ships on vibes. A tired reviewer squints at a green CI badge, approves a deploy, and moves on. The badge says build passed — but nobody can answer: what exactly was verified, against what evidence, for which commit, and who approved it? When a release breaks, that answer can't be reconstructed, because the review happened in someone's head.
As a solo maintainer, I'm the worst case: the person approving the deploy is the same person who wrote the code at 2am. So I built evidenza — an agent that does the verification work autonomously, but cannot ship without a human at the one moment that matters.
(Disclosure: the release-verification concept builds on ideas from my own prior open-source project, autokeren. evidenza is a fresh, lean codebase written from scratch during the hackathon on the Strands Agents SDK — no code was copied.)
What it does
Given a release task, evidenza:
- Captures the exact git commit being released (
git_sha) - Runs the release candidate's real test suite and commands (
shell) - Inspects manifests, configs, source (
read_file) - Checks live endpoints when applicable (
verify_url) - Records every acceptance criterion with its raw evidence in a proof artifact (JSON)
- Computes an evidence-led verdict —
SHIP / BLOCKED / NEEDS_HUMAN_REVIEW - Pauses for a human before deploy: a Strands
BeforeToolCallEventhook raises an interrupt that blocks thedeploytool until a human explicitly approves - Replays — the proof artifact renders a visual Release Card with zero API keys
Two safety properties are enforced in code, not by prompt:
- A
SHIPverdict requires every criterion to bepassed, computed from recorded evidence — never from the model's own assertion. approve()rejects any non-SHIP verdict, and the deploy hook raises an interrupt for unapproved SHIP proofs. No auto-ship, ever.
How I built it
- Strands Agents SDK is the core: the
Agentloop,@tool-decorated tools, the conversation manager, the hooks registry, and the interrupt system that powers human approval. - 7 lean tools —
read_file,write_file,shell,git_sha,git_commit,verify_url,deploy— each a thin, typed Strands tool. - Approval hook — a
BeforeToolCallEventcallback that gatesdeploybehind human approval viaInterruptException. The entire safety property is ~80 lines of hook code. - Proof system — a
plan → record → report → approvestate machine serializing proofs to JSON, plus a replay renderer for the Release Card. - Model-provider agnostic runtime — a pluggable Strands
Modelimplementation, so the agent runs on any Anthropic-Messages-compatible backend (local proxy for dev/demo, Anthropic, Amazon Bedrock). - 28 deterministic tests covering both safety properties with zero network dependency.
Challenges I ran into
- Streaming crashes: the Anthropic SDK's streaming accumulator choked on tool_use index ordering from my proxy backends. Fix: a custom Strands
Modelprovider that issues a single non-streaming request and synthesizes StrandsStreamEvents from the response. - Models hallucinating tool arguments: Strands wraps tool
inputSchemaunder ajsonkey; backends saw an empty schema and the model invented arguments. Fix: unwrap the schema before sending. - Tools never called: some backends wouldn't call tools without an explicit
tool_choice. Fix: default totool_choice: autowhenever tools are present. - Judging without credentials: I didn't want judges to need any API key or live model to evaluate the system. That constraint became the best feature: the proof artifact is replayable offline —
evidenza --proof-replay examples/demo/proof-run.jsonrenders the full evidence chain with nothing but Python.
Accomplishments I'm proud of
- The human-approval gate is a deterministic, tested mechanism (
InterruptExceptionin a hook), not a polite request to the model. - 28/28 tests pass, covering verdict logic, proof lifecycle, replay, and the approval hook — no network needed.
- The demo runs end-to-end reliably on video: real pytest output becomes evidence, deploy gets blocked, human approves, deploy proceeds.
What I learned
- Strands' hooks + interrupt design maps exactly to a human-approval gate — this is the non-obvious SDK usage the best agent architectures are made of.
- "Evidence-led" beats "model says it's fine": forcing the verdict to be computed from recorded criteria — instead of trusting model narration — surfaced every assumption I'd otherwise have shipped on.
- Safety properties that matter should live in code with tests, not in prompts.
What's next for evidenza
- Amazon Bedrock AgentCore deployment for a hosted, live-verifiable demo
- CI integration (GitHub Action that posts the Release Card on every release PR)
- Richer verifiers: security scans, performance budgets, changelog consistency
Built With
- amazon-web-services
- httpx
- pytest
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.