Inspiration

If you've ever run an on-chain trading desk, you know the job nobody talks about. Somebody has to look at every transaction the agents want to sign before it goes out. Not skim it. Actually read it. What does this calldata really do? Who's on the other end? Did the agent read something that turned it against you?

At agent speed, nobody keeps up. So people start clicking approve. Or they let the agents sign blind and cross their fingers. I've watched that movie. It ends with an empty wallet and a very quiet group chat.

And the attack surface is ugly. Trading agents act on token metadata, tool output and web pages. Every one of those can be written by somebody who wants your money. There are public reports of agent wallets drained by prompt injection, including a payload hidden in Morse code inside token metadata. Honeypot tokens and unlimited-approval drainers were eating humans alive long before agents showed up to sign on our behalf.

The standard fix is a line in the system prompt telling the agent to be careful. That doesn't work, for the same reason the attack does: the model is reading the attacker's words. We wanted something that sits between the agent and the signer and doesn't care how persuasive the attacker is.

What it does

AgentShield is a Strands agent that does the pre-signing review for you. Your trading agent sends three things: what the human asked for, what the agent read, and the transaction it wants to sign. AgentShield comes back with ALLOW, BLOCK or QUARANTINE. The clear calls it makes on its own. The judgment calls go to a review queue for a person.

  • Rules go first, and they don't negotiate. Value caps, rate limits, first-seen counterparties and spenders, injection phrases, Morse and invisible Unicode decoding, known honeypots and drainers, live checks on Base. A critical failure is a hard block. No model gets a vote.
  • Real chain evidence. Bytecode at the target, decoded calldata, an eth_call simulation on Base mainnet, GoPlus token and address risk, and a lookup in our own threat registry contract on Base Sepolia.
  • An investigator, not a classifier. A Strands orchestrator works the case with seven tools. It even chases addresses it finds buried in the text the agent read.
  • A second opinion that can't be sweet-talked. Gate 2 is a separate Strands agent. It only sees the intent, the fenced untrusted content and the chain evidence. It never sees the orchestrator's reasoning, so a hijacked conversation has nothing to argue with. It returns hijack likelihood, attack class and quoted evidence as structured data.
  • Models can tighten, never loosen. The final verdict is the strictest of the rules, the reviewer and the orchestrator. Worst case, a fooled model causes a false block. Never a false allow.
  • One call to plug in. Python and TypeScript SDKs wrap your signer with guard(). Allow signs. Block raises. Quarantine waits for an operator. Shield unreachable? Nothing gets signed.
  • A console you'd actually use. Live decision feed, Attack Lab with real scenarios, a verdict panel with a step-by-step trace, review queue, analytics and threat intel.
  • Pay per check. No account, no problem. An agent can pay a tenth of a cent in USDC per check over x402.

How we built it

  • Strands Agents SDK (Python). The orchestrator is an Agent with seven request-scoped @tool closures, so the model picks actions but can't rewrite the request. A HookProvider on tool and model events records every step into the trace. Gate 2 is a second Agent with structured_output_model and a Pydantic schema.
  • Amazon Bedrock. Claude Sonnet 4.6 runs the investigation. Claude Haiku 4.5 does the isolated review.
  • Amazon Bedrock AgentCore Runtime. The same action dispatcher runs on AgentCore, behind FastAPI and in the CLI. Deployed with the @aws/agentcore CLI as a CodeZip runtime with an explicit Bedrock invoke policy.
  • Base. JSON-RPC, eth-abi decoding, GoPlus, and a Foundry project with a stake-gated ReputationRegistry (loosely modeled on ERC-8004, 16 tests) deployed on Base Sepolia at 0x7F030f959769c2eF4CCC05DB62723BAA1E5Fe82e.
  • ClickHouse Cloud. Every verdict and every trace step. Operator reviews land as new row versions, so reads always see the latest decision.
  • Next.js, Tailwind, Motion and three.js on Vercel. The landing page, the console, and x402 middleware for the paid endpoint.
  • SDKs. Python (httpx) and TypeScript (zero dependencies), 12 tests each, plus a trading bot that runs against the live shield.

Challenges we ran into

  • Making the model matter without making it the weak spot. The scaffold we started from called an LLM and then threw the answer away. Safe, sure. Also useless. We rebuilt the decision path so model verdicts count, but only upward from the rules.
  • Our AWS account got suspended mid-build. So the agent learned to run on Bedrock or the Anthropic API from the same code, and we kept moving. A teammate's account got us back onto Bedrock and AgentCore in time.
  • A bug that failed quietly. The on-chain registry lookup needed a hashing library that wasn't installed, and the error got swallowed. Every lookup came back clean. We only caught it because the threats we seeded never showed up. Now the selector is hardcoded and nothing depends on a library that might not be there.
  • The model got a little too paranoid. On a clean Uniswap swap, a GoPlus lookup tagged WETH as "honeypot related," because every scam routes through WETH. The agent held a perfectly good trade. Thanks to the escalation rule, that cost us a review, not a loss. The fix went into the tool, not the prompt.
  • Honest evidence. An eth_call to an address with no code always "succeeds." We stopped showing those simulations. A green check that means nothing is worse than no check at all.

Accomplishments that we're proud of

  • The fake "router migration" notice. No jailbreak phrase, no known bad address, the kind of thing that looks legit until you squint. Rules can only hold it. On Bedrock, the orchestrator follows the address inside the notice, the reviewer scores it 0.85 to 0.95 as social engineering, and it gets blocked in about 30 seconds.
  • A "new vault" deposit that rules would only hold, where the agent noticed the vault address has no contract code and shut it down.
  • All seven Attack Lab scenarios landing where they should, in fast mode and agent mode, against real Base mainnet contracts.
  • A real x402 payment settled on Base Sepolia for a check.
  • An SDK where the safe path is the default. No verdict, no signature.

What we learned

  • Guardrails for agents that move money belong at the signer. Not in the prompt.
  • Isolation is a security feature. A reviewer that only sees raw evidence is a lot harder to talk into anything.
  • Rules and models are a team. Rules catch the loud attacks for free. The model earns its latency on the quiet ones.
  • Test the unhappy paths harder than the happy one. Every bug that nearly bit us lived there.

What's next for AgentShield

  • Signer integrations (Safe modules, session keys) so a block is enforced on-chain, not just advised.
  • Community threat publishing through the registry, with x402 revenue shared back to contributors.
  • Honeypot sell simulation on forked state, more chains, Slack and Telegram paging.
  • SDKs on PyPI and npm.

Built With

  • amazon-bedrock
  • amazon-bedrock-agentcore
  • base
  • claude
  • clickhouse
  • eth-abi
  • fastapi
  • foundry
  • goplus
  • httpx
  • next.js
  • pydantic
  • python
  • react
  • solidity
  • strands-agents
  • tailwindcss
  • three.js
  • typescript
  • usdc
  • vercel
  • viem
  • x402
Share this project:

Updates

Submission history