Inspiration

I hunt smart contracts for bug bounties. Every new program starts the same way: a scope list of 30 to 50 contract addresses, and thirty minutes of clerical work before any interesting analysis happens. Is this address actually a contract? What does name() return? Does the supply look sane? Does the proxy point where the docs say? The same five reads, forty-seven times, in Etherscan tabs.

The boring part is where triage mistakes hide: wrong address pasted, stale docs trusted, copy-paste assumptions. It is also the part a machine should do. Judgment stays with me; the clerical loop should run without me and surface a written note only when there is something to read. That is what a professional agent looks like in my day job.

What it does

ChainPilot is an autonomous EVM triage agent. Give it a prompt like "Triage 0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2" and it runs the loop end to end against live Ethereum mainnet:

  1. Checks the address is really a contract via eth_get_code (3,124 bytes of WETH here)
  2. Plans and executes raw eth_call reads (name, totalSupply) over permissionless public RPC with automatic fallback
  3. Decodes the ABI itself, keeping the raw hex beside every decoded value
  4. Reads storage slots directly (eth_get_storage_at, EIP-1967 proxy targets)
  5. Writes a triage_note.md artifact, then re-reads it from disk to verify it landed
  6. Surfaces a summary. You make the call.

No wallets, no chain API keys, no read/write to any chain service. Blast radius: zero. The web console streams the real Strands event loop over Server-Sent Events: tool calls, tool results, deltas. Nothing is simulated; when a call fails the card says "no data".

How we built it

The whole product is a tool-use loop, so I did not want to own a control loop. Strands Agents SDK gives me the event loop (planning, tool dispatch, retries, streaming) in about ten lines:

from strands import Agent
from strands.models.gemini import GeminiModel

agent = Agent(
    model=GeminiModel(...),
    tools=[eth_call, eth_get_code, eth_get_storage_at, read_file, write_file],
    system_prompt=SYSTEM_PROMPT,
)

Five @tool functions, plain Python over requests, no web3 dependency. The custom part is the EVM tool layer and the trust discipline: reads cap at 4 KB, a publicnode-to-llamarpc fallback chain, and a system prompt that bans fabricated chain data. Every value the agent reports must trace to a tool result, and the raw hex travels with the decoding so a human can audit the claim in seconds.

The model is a config value: the demo runs Gemini, swapping to Amazon Bedrock is a one-line change to model=, tools and prompt untouched. The frontend is FastAPI serving one HTML file: a three-column triage console (target and guardrails, live pipeline, rendered dossier).

Challenges we ran into

  • An LLM describing on-chain data is a liability. One invented address and the whole artifact is worthless. The fix was structural, not prompt-only: every fact is a tool result, hex is preserved next to every decode, and the note is verified by reading it back from disk.
  • Fresh-clone demo silently ran in mock mode. load_dotenv() resolves relative to the working directory, so a judge's bash demo.sh would have shown a deterministic stub while my machine showed the live run. The mock path exits 0 and looks plausible. Caught it only by cloning the published repo and running it exactly as a judge would.
  • localhost worked in curl and failed in the browser. The server bound IPv4-only; browsers resolving localhost to ::1 got connection refused. Fixed with an explicit dual-stack socket.
  • Streaming event shapes. Strands' stream_async emits toolResult events without the tool name, so the console correlates toolUseId against the preceding toolCall to label each pipeline card correctly.

Accomplishments that we're proud of

  • bash demo.sh on a fresh clone: installs, runs the agent, prints the trace plus the written artifact. Keyless mock mode exits 0 for CI; any free AI Studio key gets the real mainnet loop. The agent writes its own proof, and that is the claim judges are invited to test.
  • A console that shows the actual event loop, not a mock. Tool cards with raw JSON-RPC payloads, decoded values beside their hex, disk readback verification badge.
  • The whole thing is verifiable-by-construction in an area where that is the product: read-only by design, permissionless RPC, no keys, no simulated traces.

What we learned

  • Tool discipline matters more than prompt cleverness. Size caps, fallback chains, and structured output rules are what make agent output auditable.
  • Reproducibility is a feature, not cleanup. The dotenv bug is exactly the class of failure that kills submissions at judging time, not build time.
  • Strands kept its promise: the SDK owns the loop, I own the tools and the workflow. Model portability (Gemini to Bedrock in one line) turned out to be real, not marketing.

What's next for ChainPilot

  • Cross-check chain reads against an independent security oracle so verdicts come from two agreeing sources, disagreements surfaced, not averaged.
  • Batch mode: feed the scope CSV of a new program, wake up to a triage folder.
  • Spending authority is a later design question, not a demo shortcut. Today the agent is read-only by design, which is the kind of trust professionals need first.

Repo: https://github.com/cryptoflops/chainpilot (MIT). Demo video: https://youtu.be/kLO_UwODNeE

Share this project:

Updates

Submission history