Inspiration
I hunt smart contracts for bug bounties. Every new program starts the same way: a scope list of 30 to 50 contract addresses, and thirty minutes of clerical work before any interesting analysis happens. Is this address actually a contract? What does name() return? Does the supply look sane? Does the proxy point where the docs say? The same five reads, forty-seven times, in Etherscan tabs.
The boring part is where triage mistakes hide: wrong address pasted, stale docs trusted, copy-paste assumptions. It is also the part a machine should do. Judgment stays with me; the clerical loop should run without me and surface a written note only when there is something to read. That is what a professional agent looks like in my day job.
What it does
ChainPilot is an autonomous EVM triage agent. Give it a prompt like "Triage 0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2" and it runs the loop end to end against live Ethereum mainnet:
- Checks the address is really a contract via
eth_get_code(3,124 bytes of WETH here) - Plans and executes raw
eth_callreads (name,totalSupply) over permissionless public RPC with automatic fallback - Decodes the ABI itself, keeping the raw hex beside every decoded value
- Reads storage slots directly (
eth_get_storage_at, EIP-1967 proxy targets) - Writes a
triage_note.mdartifact, then re-reads it from disk to verify it landed - Surfaces a summary. You make the call.
No wallets, no chain API keys, no read/write to any chain service. Blast radius: zero. The web console streams the real Strands event loop over Server-Sent Events: tool calls, tool results, deltas. Nothing is simulated; when a call fails the card says "no data".
How we built it
The whole product is a tool-use loop, so I did not want to own a control loop. Strands Agents SDK gives me the event loop (planning, tool dispatch, retries, streaming) in about ten lines:
from strands import Agent
from strands.models.gemini import GeminiModel
agent = Agent(
model=GeminiModel(...),
tools=[eth_call, eth_get_code, eth_get_storage_at, read_file, write_file],
system_prompt=SYSTEM_PROMPT,
)
Five @tool functions, plain Python over requests, no web3 dependency. The custom part is the EVM tool layer and the trust discipline: reads cap at 4 KB, a publicnode-to-llamarpc fallback chain, and a system prompt that bans fabricated chain data. Every value the agent reports must trace to a tool result, and the raw hex travels with the decoding so a human can audit the claim in seconds.
The model is a config value: the demo runs Gemini, swapping to Amazon Bedrock is a one-line change to model=, tools and prompt untouched. The frontend is FastAPI serving one HTML file: a three-column triage console (target and guardrails, live pipeline, rendered dossier).
Challenges we ran into
- An LLM describing on-chain data is a liability. One invented address and the whole artifact is worthless. The fix was structural, not prompt-only: every fact is a tool result, hex is preserved next to every decode, and the note is verified by reading it back from disk.
- Fresh-clone demo silently ran in mock mode.
load_dotenv()resolves relative to the working directory, so a judge'sbash demo.shwould have shown a deterministic stub while my machine showed the live run. The mock path exits 0 and looks plausible. Caught it only by cloning the published repo and running it exactly as a judge would. - localhost worked in curl and failed in the browser. The server bound IPv4-only; browsers resolving
localhostto::1got connection refused. Fixed with an explicit dual-stack socket. - Streaming event shapes. Strands'
stream_asyncemitstoolResultevents without the tool name, so the console correlatestoolUseIdagainst the precedingtoolCallto label each pipeline card correctly.
Accomplishments that we're proud of
bash demo.shon a fresh clone: installs, runs the agent, prints the trace plus the written artifact. Keyless mock mode exits 0 for CI; any free AI Studio key gets the real mainnet loop. The agent writes its own proof, and that is the claim judges are invited to test.- A console that shows the actual event loop, not a mock. Tool cards with raw JSON-RPC payloads, decoded values beside their hex, disk readback verification badge.
- The whole thing is verifiable-by-construction in an area where that is the product: read-only by design, permissionless RPC, no keys, no simulated traces.
What we learned
- Tool discipline matters more than prompt cleverness. Size caps, fallback chains, and structured output rules are what make agent output auditable.
- Reproducibility is a feature, not cleanup. The dotenv bug is exactly the class of failure that kills submissions at judging time, not build time.
- Strands kept its promise: the SDK owns the loop, I own the tools and the workflow. Model portability (Gemini to Bedrock in one line) turned out to be real, not marketing.
What's next for ChainPilot
- Cross-check chain reads against an independent security oracle so verdicts come from two agreeing sources, disagreements surfaced, not averaged.
- Batch mode: feed the scope CSV of a new program, wake up to a triage folder.
- Spending authority is a later design question, not a demo shortcut. Today the agent is read-only by design, which is the kind of trust professionals need first.
Repo: https://github.com/cryptoflops/chainpilot (MIT). Demo video: https://youtu.be/kLO_UwODNeE
Built With
- agent
- ai
- amazon-web-services
- blockchain
- css
- gemini
- html5
- python
- sse
- strands
Log in or sign up for Devpost to join the conversation.