Inspiration

2025's Word of the Year, per Merriam-Webster, was "AI slop." By 2026 it had stopped being a joke. curl shut down its bug bounty program after AI-generated submissions crowded out real vulnerability reports. Ghostty started permanently banning contributors who submitted bad AI-generated code. A developer submitted a 13,000-line AI-generated pull request to the OCaml compiler, admitting he'd written zero lines of it himself, and maintainers rejected it purely because the review burden was unsustainable. GitHub itself opened a public discussion about possibly disabling pull requests entirely to cope.

Every tool built in response to this answers one question: who or what wrote this code. Attribution tools exist. What none of them answer is the question that actually determines whether a maintainer should trust a contribution: does the person committing it understand it? A developer can accept a generated function, never read it, commit it under their own name, and pass every attribution check that exists today.

That gap is Vouchcode.

What it does

Vouchcode is a local-first CLI that tracks AI versus human authorship at the commit level and requires the committing developer to demonstrate they understand AI-generated code before it's sealed into a tamper-evident, cryptographically signed ledger.

  • Detects AI-generated code via direct tool signals (a Claude Code adapter) or a confidence-capped stylometric fallback, never asserted as certain
  • Segments diffs by abstract syntax tree, not by line, so a renamed function is never mistaken for a rewritten one
  • Asks targeted comprehension questions derived from the actual control flow of AI-generated code, and scores answers deterministically, no external model involved
  • Seals every commit's outcome into a hash-chained, Ed25519-signed ledger, so tampering is mathematically detectable, not just inconvenient
  • Generates a signed, portable JSON and PDF report anyone can verify offline, no account, no install
  • Gates a CI pipeline: vouchcode gate fails a build if AI-attributed code lacks a passing comprehension record
  • Publishes an honest, self-generated README badge that says "comprehension not evaluated" rather than rounding up when nothing was evaluated

Everything runs entirely on the developer's machine. Vouchcode makes zero calls to any external AI or LLM API in its own runtime.

How we built it

Six layers: Capture (git hooks plus direct tool signals), Segment (AST-based diffing), Verify (the comprehension engine), Seal (the cryptographic ledger), Report (signed JSON and PDF), and Integrate (the CI gate and badge).

Every phase was built against a hard adversarial exit criterion, not just a happy-path test. We didn't just test that a renamed function is recognized as a rename, we tested a function renamed and internally edited, to force the segmentation logic to correctly separate the two effects rather than collapsing them into one signal. We didn't just test that a correct comprehension answer scores well, we tested a keyword-stuffed fake answer against it in the same run, to prove the scorer discriminates on reasoning rather than vocabulary overlap. We didn't just test that a forged, re-signed ledger fails verification, we proved it passes internal verification cleanly (a signature only proves a document wasn't altered, not who signed it), and that only an independently published key fingerprint catches the forgery.

Then we ran Vouchcode on itself. Every commit in this repository's history has been analyzed by Vouchcode's own ledger.

Challenges we ran into

The hardest bug wasn't a feature gap, it was a hash-chain implementation that compared each entry against its stored predecessor hash instead of its recomputed one. A tampered entry's damage stayed confined to itself instead of propagating, which would have made the entire tamper-evidence claim silently false while every demo still looked correct. It only surfaced because the exit criterion for that phase specifically demanded mutating a field and tracing exactly which entry broke, not just checking that something failed.

A close second: literal \x08 control-code bytes from a heredoc-mangled \b regex, silently disabling a comprehension check with no visible trace in the source file. Found only because the test suite was itself scanned for control characters as a matter of discipline, not because anyone was looking for that specific bug.

Accomplishments that we're proud of

Running Vouchcode on its own 61-commit development history and getting an honest answer back: 43.4 percent AI-attributed, reported at a mean confidence of 0.245, explicitly flagged as a probable undercount because the stylometric baseline is drawn from the very code it's scoring against. A tool that verifies AI code honesty, being honest about its own.

What we learned

The gap in this space isn't detection, it's accountability. The research literature already says so: a peer-reviewed study of over a thousand developer discussions on AI-generated code concluded that tool developers should shift focus from generation to verification, specifically through uncertainty indicators and provenance information. We built toward that conclusion without having read it first.

What's next for Vouchcode

Multi-language support beyond Python via tree-sitter, a hosted badge verification option alongside today's honest self-generated one, a GitHub PR bot that posts the gate result directly as a comment, and passphrase-protected key rotation.

Built With

  • ast
  • ci-cd
  • claude-code
  • cli
  • cryptography
  • ed25519
  • git
  • github-actions
  • gitpython
  • json
  • mypy
  • pdf
  • pytest
  • python
  • reportlab
  • rich
  • ruff
  • svg
  • typer
Share this project:

Updates

Submission history