Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned## Inspiration
Agents now write most of our code and read almost none of it — and when you ask what an agent did, you get its answer: a commit message, a summary, a green checkmark. The Quirq challenge framed the problem exactly: agents run for hours, consume tokens, modify files — and it's hard to see what they actually did or whether they stayed inside their boundaries. We wanted to push that one step further than observability: not just display what the agent did, but prove it, in a record the agent itself cannot quietly edit.
What it does
Custody is a chain of custody for machine-written code. One agent (Claude Opus 5) hardens a repository — but only under a scope contract it must declare before touching anything. A second, adversarial auditor never reads the agent's explanation: it re-derives what happened from the diff, the project's own test results, and a fresh deterministic survey, runs seven cheat detectors (deleted tests, silenced linters, loosened quality gates, scope escapes, under-declared spend...), and rules PROVEN / NOT_OBSERVED / INSUFFICIENT_EVIDENCE / REJECTED. Only PROVEN survives as a commit. Everything lands in an append-only, hash-chained ledger — edit it, truncate it, or rewrite a verdict, and custody verify says so offline with exit code 2.
The distinction the whole system refuses to collapse: "we checked and found nothing" is never the same as "we could not check."
How we built it
- Python 3.9–3.13, standard library only — zero runtime dependencies, because an auditing tool that drags in a supply chain undermines its own premise. The Claude API client is ~30 lines of
urllib; analysis isast; hashing ishashlib; the dashboard ishttp.server. - Deterministic where it must be, agentic where it helps. Detection is fully deterministic so any accusation is reproducible. The model does exactly two jobs: proposing fixes, and adjudicating the one ambiguous state (a PROVEN ruling carrying advisory suspicion) — as a one-way ratchet: the model may add caution to the audit; it may not remove any.
- Measured, not asserted: 256 tests, a ground-truth eval (8/8 injected cheats caught, 0/6 clean fixtures flagged — false positives fail the eval too), and an adversarial trial of seven scripted cheats that runs on every CI build.
- Proven on real code: run against Bottle (15 years in production, 381 tests) — 1,457 findings surveyed, 3 fixed and PROVEN against Bottle's own suite, 42 honestly declined.
- Contributed back: a Custody adapter for
quirq-ai/xo-space(zero core edits, 12/12 tests) that makes the ledger a session store the Space can read — the only agent whose activity feed is tamper-evident.
Challenges we ran into
The best bugs were the tool's own principles violated inside the tool:
- Finding IDs were line-keyed — inserting one blank line above a finding retired its ID, and an unfixed finding adjudicated as PROVEN. IDs are now content-anchored, and clearing requires two independent signals.
shell=1bypassed the flagship rule (it only matched the literalTrue). One character.- A red baseline convicted the innocent: hardening Bottle on Windows hit 9 pre-existing test failures, and every honest attempt was rejected as "the suite failed after the change." Custody now refuses a red baseline the way it refuses a dirty tree.
- The auto-revert once destroyed the ledger it existed to protect — and reported "ledger intact: True," because an empty chain verifies clean. A destroyed audit trail presenting as a clean result, inside the tool built to catch exactly that.
- Custody flagged our own test fixtures for containing key-shaped strings (true positive — we now assemble them at runtime), and flagged our own reviewer code for excessive complexity mid-development. Being audited by your own tool is humbling and extremely effective.
What we learned
Trust has to be structural, not narrative. Every place we let a check fail silently — an unpriced model, an oversized file skipped, a .env never scanned — "not looked" became indistinguishable from "looked and found nothing." The fix was never to trust harder; it was to make the gap itself a recorded finding.
What's next
TypeScript/JavaScript rule packs (the architecture is language-agnostic; today's rules are Python), signed ledgers, and landing the xo-space adapter upstream so every Quirq workspace can display tamper-evident agent activity.
Log in or sign up for Devpost to join the conversation.