CodeProof
Analyze the code. Prove the report.
Code analysis reports are easy to share — and surprisingly hard to trust once they leave the machine that produced them. A score can be edited. Findings can be removed. Configuration can change. An AI explanation can be rewritten. The final report may still look perfectly legitimate.
CodeProof turns a static-analysis result into portable, independently verifiable evidence.
The problem
Developer teams increasingly use automated and AI-assisted analysis in reviews, audits, CI, vendor assessments, and security workflows. But the recipient of a report often has no independent way to answer:
- Does this report correspond to the repository snapshot it claims?
- Were findings, scores, or configuration changed afterward?
- Was an AI explanation modified?
- Can I verify the result without trusting the original machine or calling the AI provider again?
CodeProof makes those questions mechanically checkable.
What CodeProof does
Repository snapshot
↓
Deterministic CodePulse analysis
↓
Findings + scores + effective config
↓
Canonical JSON + SHA-256
↓
codeproof/1 evidence envelope
↓
Hash-chained ledger
↓
Self-contained HTML report + offline verifier
A normal scan produces sealed JSON evidence, a hash-chained ledger, and a self-contained HTML report. Verification recomputes the integrity chain independently and returns one of three explicit states:
VALID · INVALID · NOT_CHECKED
No ambiguous green checkmark is used where verification did not actually occur.
30-second demo
pip install -e .
codeproof demo
The demo:
- creates a synthetic repository;
- runs deterministic analysis;
- seals the resulting evidence;
- verifies it as VALID;
- modifies a sealed field;
- verifies the modified evidence as INVALID;
- restores the original record.
The same flow is visible in the browser report: Verify Evidence → Simulate Tampering → INVALID → Restore Original → VERIFIED.
Why AI is useful here — and why it is not trusted blindly
CodeProof deliberately keeps AI downstream of deterministic evidence.
The analyzer produces findings and scores without AI. Optional AI can then explain, prioritize, and summarize those deterministic findings, but it cannot create findings, change scores, or affect verification. Its output is sealed separately inside the evidence envelope.
That means CodeProof gets the usability benefit of AI while keeping the trust boundary explicit:
AI interprets the evidence. It does not become the evidence authority.
Verification works offline and never requires another AI call.
What is bound into the evidence
CodeProof seals:
- repository snapshot identity;
- commit and dirty-state information;
- effective analyzer configuration;
- deterministic findings and scores;
- analyzer identity and version;
- optional AI output hash;
- report payload identity;
- evidence-chain state.
A changed finding, score, config value, snapshot field, AI output, or ledger entry produces a verifiable mismatch instead of silently becoming the new truth.
Practical value
CodeProof is aimed at situations where the report travels farther than the analyzer:
- code-review handoffs;
- CI artifacts;
- security and quality audits;
- vendor or contractor assessments;
- compliance evidence;
- long-lived engineering records;
- AI-assisted developer workflows.
The recipient does not need to trust a screenshot, a PDF, or the person who generated the report. They can retain the evidence and verify it independently later.
Technical architecture
The system separates five concerns:
- Snapshot — identify the repository state being analyzed.
- Analysis — deterministic AST-based analysis across six dimensions.
- Evidence — canonical serialization and SHA-256 integrity sealing.
- Ledger — append-only hash chaining for sequential records.
- AI interpretation — optional, downstream, separately hashed.
The analyzer dependency is pinned to an exact revision, and the repository includes explicit architecture, threat-model, claims, provenance, demo-script, and testing documentation.
Honest verification boundary
CodeProof does not claim that SHA-256 proves authorship, machine identity, timestamp authenticity, vulnerability absence, or semantic correctness of every finding.
It proves something narrower and useful: whether the retained evidence remains internally consistent with the sealed analysis record.
For complete-rewrite or tail-truncation scenarios, CodeProof supports externally retained expected hashes / ledger heads rather than pretending a self-contained chain can prove history it never witnessed.
Why this matters for AI-assisted software
As AI-generated interpretations become more common, provenance becomes part of product quality. CodeProof demonstrates a pattern where deterministic software evidence and AI assistance can coexist without collapsing into the same trust domain.
The product is intentionally simple to try:
codeproof scan . --no-ai --output out/
codeproof verify out/codeproof.evidence.json
Then open the generated HTML report and verify the same evidence directly in the browser.
Built for real review, not just a demo
- deterministic core analysis;
- optional AI layer;
- portable sealed evidence;
- offline verification;
- tamper demonstration;
- self-contained browser report;
- hash-chained ledger;
- explicit threat model and limitations;
- reproducible CLI workflow;
- MIT licensed.
CodeProof does not ask you to trust the report. It gives you something you can verify.
Built With
- codepulse
- github-actions
- pytest
- python
- sha-256
Log in or sign up for Devpost to join the conversation.