Inspiration
Firmware security analysis matters, but very few people ever get to watch it happen. It is slow and quiet, and it takes place inside a terminal. At the same time, everyone is asking how capable AI agents really are on technical work. We wanted to answer that question live, in front of a crowd, in a way that is fun to watch. So we gave several AI agents the same piece of firmware and the same rules, let them race, and let the audience bet on the result.
What it does
Vault Heist is a live betting game. Three AI agents, powered by Gemini, OpenAI and Claude Haiku, race to find a planted vulnerability in IoT firmware.
- Players join from their phones and each place one bet on an agent. The odds update live as the pot fills.
- The host locks betting and the race begins. All three agents search the same read-only firmware filesystem at once.
- A big screen shows each agent's reasoning live, progress bars for each milestone, the current odds, and an announcer calling the action.
- The first agent to correctly name the vulnerability (the file, function or hardcoded string) wins. The vault on screen cracks open, and the pot is split among everyone who backed the winner.
The system only identifies vulnerabilities. Agents never write exploits or recover live secrets. The target is OWASP IoTGoat, firmware published specifically for this kind of authorized security analysis.
How we built it
Architecture. The whole system runs on one event bus, which is the single source of truth. The agents, the judge, the betting engine and the game state machine all publish events to it. A WebSocket layer sends every event to the big-screen dashboard and to each player's phone. The announcer and the recorder listen on the same bus.
Backend. Node.js with pure ES modules and no build step. A state machine moves each round through a fixed order:
$$ \texttt{LOBBY} \rightarrow \texttt{BETTING_OPEN} \rightarrow \texttt{BETS_LOCKED} \rightarrow \texttt{RACING} \rightarrow \texttt{SETTLED} $$
Agent sandbox. We extracted the IoTGoat root filesystem (about 1,000 files) and checked our answer key against it. Agents can use only four read-only tools: list_dir, read_file, grep and strings. They get no shell access.
Betting. Payouts are pari-mutuel. If $P$ is the total pot, $W$ is the total bet on the winning agent, and $b_i$ is player $i$'s bet on that agent, then player $i$ receives
$$ \text{payout}_i = b_i \cdot \frac{P}{W} $$
Before betting locks, the live odds shown for agent $k$ are $P / W_k$, where $W_k$ is the total bet on agent $k$.
Frontend. A static, vanilla JavaScript dashboard and a phone player screen. A pure reducer turns the event stream into what appears on screen, which makes the display logic easy to unit test.
Announcer. We generated 17 voice lines ahead of time with ElevenLabs, so no text-to-speech calls happen during a show.
Record and replay. Any race can be recorded and played back through the same event bus, so a replay looks exactly like a live run. If the Wi-Fi or an API fails on stage, we still have a complete demo.
Challenges we ran into
- Making different model APIs behave the same way. Gemini, OpenAI and Anthropic each handle multi-turn tool calls differently. Gemini gave us the most trouble: we had to fix how tool-call signatures carried across turns and update to a newer model ID before it raced reliably.
- Keeping the race fair. Agents could submit guesses over and over until one landed. We capped each agent at three wrong submissions and started showing rejected submissions on the dashboard.
- Controlling cost. Real model calls cost money, and a crowd can start races quickly. We added a hard spend cap of $2 per race, a cooldown between races, and a host token that controls who can start a race.
- Announcer spam. Early versions repeated lines and kept talking after the race was over. We added per-round deduplication and made the announcer go quiet after a win.
- Demo-day risk. A live AI race over conference Wi-Fi can fail in many ways. That is why we built recording, replay, and mock agents that run without any API keys.
Accomplishments that we're proud of
- The full experience works end to end on separate devices: joining, betting, a live race, a cracking vault, payouts and a narrated announcer.
- Every event goes through one event bus, which made replay nearly free to build and kept the system easy to reason about.
- The safety boundary holds. Agents use read-only tools on purpose-built training firmware, and the task is identification only.
- A test suite covers the bus, the state machine, the sandbox (run against the real firmware), the judge, the betting math, recording and replay, the announcer and the dashboard reducer.
What we learned
- Agent behavior depends heavily on the details of each provider's API. A small difference in how tool calls work can change whether an agent solves the task.
- Telling agents the exact goal of the round improved their focus noticeably.
- A live demo needs a backup plan built in from the start. Record and replay turned out to be one of the most valuable features we built.
- Making security work enjoyable to watch is mostly a presentation problem: clear milestones, visible reasoning and a good announcer.
What's next for Rowdy's Security Agents
- Add more firmware targets and more vulnerability rounds so the race doesn't stay the same.
- Add more agents to the crew and keep a leaderboard of results across many races.
- Publish benchmark-style statistics, such as solve rate, time to solve and cost per solve for each model.
- Let players bet on milestones within the race, not only on the final winner.
Built With
- ai-agents
- anthropic
- binwalk
- claude
- claude-haiku
- css
- cybersecurity
- elevenlabs
- firmware-analysis
- gemini
- google-genai
- gpt-4o-mini
- html
- javascript
- llm
- node.js
- openai
- openwrt
- owasp-iotgoat
- squashfs
- tool-calling
- websockets
- ws
Log in or sign up for Devpost to join the conversation.