Inspiration

Smart contract audits cost $20,000–$150,000 and take 1–4 weeks. Most new DeFi protocols can't afford that, so they either ship unaudited or skip security entirely — and $2B+ has been lost to exploits that better tooling could have caught before deployment.

We kept seeing the same pattern in contest reports on Code4rena and Sherlock: critical bugs that were simple in hindsight but buried in thousands of lines of Solidity. Static analyzers like Slither catch maybe 15-20% of what a real audit finds. Raw LLMs hallucinate findings with no way to verify them. Nobody had actually connected AI reasoning to real, executable proof.

We wanted to build the layer that doesn't exist yet: something that sits between "free but shallow" and "$50K but slow" — cheap enough that any solo developer runs it before every deploy, and rigorous enough that its findings are backed by actual code execution, not AI opinion.

What it does

ChainAudit AI is an 8-layer automated security pipeline for Solidity smart contracts:

  1. Static analysis — Slither with auto-detected Foundry remappings
  2. AI logic review — a two-stage LLM pipeline (cheap model scouts every file, stronger model validates against actual code)
  3. Cross-contract & economic attack analysis — flash loans, sandwich attacks, oracle manipulation
  4. Foundry fuzzing engine — the core innovation: converts AI suspicions into structured JSON specs, routes them through protocol-specific adapters (lending/AMM/vault/staking), and generates real Forge test harnesses that actually execute
  5. Historical exploit matching — cross-references a 50K-entry database of real audit findings
  6. Confidence scoring, AI-generated fixes, and a full HTML report with exploit reproduction code

The result: for every finding, you get either a forge-verified counterexample with exact gas cost and call trace, or an honest "not reproduced" — never a fake pass.

How we built it

The architecture went through several dead ends before it worked. Early on we tried letting the LLM write Foundry test code directly — it hallucinated syntax constantly and almost never compiled. We rebuilt around a strict principle: the AI only ever outputs structured JSON specs. Python owns the schema. Templates generate the Solidity. This one architectural decision is what made the whole engine reliable.

We built protocol adapters that detect contract type (lending, AMM, ERC4626 vault, staking) and route to appropriate mock dependencies — including realistic mocks for Uniswap position managers, Permit2, and Chainlink oracles, so fuzzing actually exercises real business logic instead of bouncing off empty mock returns.

We benchmarked against real, unmodified Code4rena contest repositories — Revert Lend, DYAD, and Kinetiq — comparing our output against the actual published audit findings, not synthetic test cases.

Challenges we ran into

Two bugs took us days each to isolate:

The proxy wall. Modern protocols are almost universally built with OpenZeppelin's upgradeable pattern. Our fuzzing engine could detect these proxy contracts correctly — but had no way to actually deploy and test them, so it correctly refused to fake results and just marked them untestable. On one benchmark, every single core contract in the protocol was upgradeable, meaning our engine could test 0% of the actual logic. We built a two-step deployment harness (deploy implementation → deploy ERC1967Proxy → call initialize through the proxy) that took our testable coverage on that repo from 0% to 46% in a single fix.

Accomplishments that we're proud of

Our fuzzing engine independently rediscovered and forge-proved a real critical vulnerability in Revert Lend's V3Utils.execute() — a missing caller-validation bug that lets the contract's operator approval persist after use — with exact reproduction evidence:

Built With

  • any
  • bug-free
  • doesn't
  • exposing
  • let-me-generate-the-architecture-diagram-(offered-earlier)-?-a-clean
  • pipeline
  • risk
  • that
  • visual
Share this project:

Updates

Submission history