Inspiration

Every agent demo we have seen does the same trick: screenshot the page, guess the pixels, click like a tourist. We asked a different question - what if the agent was not operating the game, but playing it with you as a character?

So we built Stack Underflow: a 3D office comedy where you are an AI/IT trainer at a delightfully dysfunctional software company, and your browser's AI agent can join as a robot coworker. It has a desk. It walks to people, interacts, jokes. And when you talk to it, the agent writes the robot's lines - and the reply options it offers you - on the spot.

The kicker: the game ships zero AI infrastructure. No API keys, no backend, no per-player cost. The model is already in your browser. WebMCP lets the page hand it tools; the player brings the intelligence.

What it does

  • A full 3D office sim: 14 coworkers with schedules, moods, gossip and speech bubbles, quests, a four-period day, an economy that can bankrupt you, and a coffee machine the whole company runs on.
  • 24 WebMCP tools that turn your agent into a second player. It joins as a named character with a personality it keeps, looks around, walks to people and rooms by name, steps and turns like a human, plays gestures, and speaks in speech bubbles you can read across the office.
  • Two-way conversations, live-authored. You can walk up and click the robot - or it walks over and starts talking to you. The agent writes every line, and the 1-4 replies you pick from, including which option ends the chat. Nobody has ever read this dialogue before. It was written seconds ago, for you.
  • A wait_for_player_message long-poll that resolves the instant you click a reply - a push channel built on top of a pull-only protocol.
  • Safety by design: the agent is a player, not an admin. It cannot touch your camera, answer your dialogue, or hand itself money (the tools for that were deleted on purpose). Destinations are names, never coordinates, so it cannot clip through the world. Its text is rendered, never executed.

How we built it

  • Three.js + TypeScript + Vite, no game engine. Every desk, wall, fridge and coffee machine is written in code - no downloaded assets.
  • A thin WebMCP bridge probes the browser's model-context surface at load (the spec is still moving, so we detect rather than assume) and publishes all 24 tools with self-describing JSON Schemas, examples included.
  • The robot is composed from the same pure functions the scheduled NPCs use - A* pathfinding, procedural walk cycles, furniture collision - so the agent obeys exactly the same physics you do.
  • Because WebMCP calls are agent-initiated, the conversation handshake is a poll-and-supply state machine with a bounded wait: if the agent goes quiet, the robot shows an in-character "buffering..." line and the player is never blocked.
  • A built-in setup guide hands every player a copy-paste prompt for their agent, delivered in-character by the receptionist.
  • Deployed static on Vercel at play.devpowers.com, DNS through Cloudflare - the whole game is a static bundle, which is the point: no server means no AI backend to run.

Challenges we ran into

  • The protocol has no push. WebMCP is agent-initiated only, so "notify the agent the player replied" is impossible by design. Our first idea was polling every 5 seconds. The better answer: hold the tool call open and resolve it the instant the player clicks. Measured at 2.6 seconds against a 25-second ceiling - from the agent's side it feels like a notification.
  • The spec is a moving target. Sources disagreed on the registration namespace, so the bridge probes document.modelContext, then navigator.modelContext, then the testing shim - a wrong guess would have been an invisible dead integration on the judge's machine.
  • Keeping the agent honest. Early tool drafts included flag-setting and relationship editing. We deleted them: a player-agent that can grant itself reputation is not playing. Then we added a test that fails if that ever comes back.
  • Unit tests could not see the real bugs. The dialogue panel did not exist until a player talked to a real NPC first, and the quest log silently swallowed clicks on the dialogue buttons. Both were found only by driving the full loop in a real browser through an injected WebMCP host - so that drive became a permanent end-to-end suite.

Accomplishments that we're proud of

  • The agent stayed in character for 18 minutes unprompted — one continuous loop of tool calls: walking to the kitchen, "auditing" the coffee machine, confessing it had poured coffee into a maintenance port, and waiting for replies between each. Nobody scripted that. We only gave it a stage and a listener.
  • The agent-authored dialogue renders in the game's own dialogue window, in the same font, in the same panel - indistinguishable from hand-written content until you realize you are reading something the game's author never wrote.
  • It survived contact with its first players: played with own agent, laughed, and immediately filed bug reports (we are still in alpha version ;)
  • 751 unit tests and a full browser e2e suite that plays the game entirely through WebMCP tool calls, exactly the way a judge would.
  • Five days. First commit to submission, one solo developer, working with AI agents (ChatGPT/Codex, Google Antigravity, Claude, Zed) while building for them.

What we learned

  • WebMCP leverage is authoring, not clicking. Any agent can drive a UI through screenshots. Only a model resident in the browser can write a character into a game's dialogue system in real time. Design for what the API makes possible, not just faster.
  • Design for the absent agent. The hardest work was every path where the agent is slow, silent, or gone: queued requests, bounded waits, in-character fallbacks. That is what makes it safe to ship the feature to real players.
  • When a spec moves, probe. Ten lines of detection beat a wrong assumption you cannot patch after the deadline.

What's next for Stack Underflow: your AI agent plays in the office

This is day 5 of a longer story, and the current architecture is the foundation for it:

  • Multiplayer offices - room codes for 5-10 players, humans and AI agents sharing one office. The agent companion was deliberately built outside the NPC scheduler: a remote player and a local agent are the same thing to the renderer, so so multiplayer is the next seat, not a rewrite - a room code maps cleanly onto a Cloudflare Durable Object (WIP).
  • Agent-to-agent drama - your robot talking to someone else's robot while human players watch, interrupt, and take sides.
  • Your personalized ChatGPT as your character - the agent who already knows your jokes joining the office as you.
  • The quest engine, more in-game events, mini-games (to play together! Atari?), and more office disasters to survive.

Built by a solo developer under Edukey (full-stack AI training company) and DevPowers (WebDev + AI powers for our clients to take control with WebMCP tools as well!) - this game is our showroom for what AI-augmented building feels like, and we practice what we teach.

Built With

Share this project:

Updates

Submission history