Inspiration

AI agents are advancing fast, but the field is bottlenecked on data. Most web agents are trained on scraped pages and synthetic demonstrations. Almost none of it captures the how an agent behaves when it comes across something unexpected, such as a modal appearing over the button it was about to click, a control it was promised is suddenly disabled, or a decoy sitting exactly where the real one was. Failures like that can't be scraped, but rather has to be manufactured. This data is what AI agents need to get better.

So we built a way to create it, with a fun twist. Four agents race the same task on the same page at the same time, a master model actively sabotages them mid-run, and we log everything (every decision, every keystroke, every pass and failure, timestamped against the exact sabotage). To make each run more interesting, we wrapped it in a live prediction market: spectators bet on which agent wins, and the market's odds become a live read on each agent's recovery as it happens.

The race is the spectacle and the dataset is the purpose.

What it does

PolyBot runs four AI agents against the same task at the same time, each in its own Steel.dev cloud browser:

  • Session events for training data. We capture navigations, loads and failures from the browser's own side, which becomes training data for future web agents. Checkpoints and sabotage are event-driven: each sabotage fires at a verified checkpoint, so a fast agent and a slow one meet the same sabotage on the same screen at the same point in the task. That is what makes the four runs comparable instead of merely simultaneous.
  • Live sessions. Four real Chrome instances spin up per race, one per model, fully isolated from each other. We attach Playwright to each over CDP.
  • Session viewers. Every agent's browser is embedded live in the UI through Steel's read-only viewer, so you watch four models shop the same store simultaneously, in real time.
  • Session replay. After a race, Steel's recording lets you scrub back through exactly what each agent saw and clicked.

While the agents work, a master model chooses sabotage and we inject it over CDP into one agent's live page: a decoy button beside the real one, the real control disabled, a modal dropped over the checkout. Same page, same moment, different browser, so we can watch one model get fooled while another routes around it.

Spectators bet virtual credits on who wins. Odds move as agents clear checkpoints, get stuck, recover, or fail. When the race ends, PolyBot replays what each agent did and scores how it handled every hazard.

How we built it

We built the backend in TypeScript and used Steel to run four isolated browser sessions simultaneously. Each competing model receives a screenshot and information about the page, chooses a browser action, and continues until it completes the task or fails.

We used OpenRouter to connect the competing models and the master agent. The backend verifies progress through checkpoints, applies sabotages, records every action, and updates the prediction market.

The frontend was built with React and Vite. It displays the four browser sessions, agent activity, race progress, betting controls, changing odds, and post-race evaluations. We also built our own shopping website so every model receives a consistent and verifiable task.

Challenges we ran into

The hardest part was keeping everything synchronized. The four agents do not move at the same speed, so each sabotage must activate separately when an agent reaches the correct point.

We also dealt with models returning invalid actions, agents repeating the same steps, rate limits, disconnected browser sessions, and sabotages that were too difficult to recover from. We originally planned three sabotages, but the third caused every agent to get stuck, so we reduced each race to two.

Getting the live website working on phones was another challenge. We had to handle Cloudflare tunnels, QR-code access, responsive layouts, video playback, and browser autoplay restrictions.

Accomplishments that we're proud of

We got four different models racing in real browser sessions at the same time. Judges can watch the agents work, bet from their phones, and see the odds react to what is happening.

PolyBot also produces more than a final winner. It records what each agent observed, which actions it chose, where it struggled, and whether it recovered from sabotage. This creates useful evaluation data for comparing how models behave under pressure.

What we learned

We learned that an agent performing well on a normal webpage does not necessarily mean it will handle unexpected changes well. Something as simple as a fake button or obstructive modal can expose major differences in observation, planning, persistence, and recovery.

We also learned how much infrastructure is required to run browser agents reliably. Steel gave us isolated browser sessions and the visibility needed to watch and evaluate each agent. Building the prediction market taught us how to turn complicated agent behavior into something spectators can quickly understand.

What's next for PolyBot

We want to add more websites, tasks, and types of controlled sabotage. We also want users to submit their own agents and compare them against the same standardized challenges.

Eventually, PolyBot could become a live benchmark where models are ranked on more than task completion. We want to measure how quickly they finish, how often they fail, and how well they recover when a website does something unexpected.

Built With

  • ai-agents
  • cdp
  • cloudflare
  • computer-use
  • cors
  • dom-injection
  • fastify
  • hls
  • model-benchmarking
  • multi-agent-systems
  • node.js
  • openrouter
  • playwright
  • prediction-markets
  • react
  • sse
  • steel
  • typescript
  • vite
  • vitest
+ 6 more
Share this project:

Updates

Submission history