About the project
Inspiration
Most AI-in-the-browser demos script an agent to click through a UI built for humans. WebMCP lets a page expose tools directly to an agent instead. We wanted an agent with a real role in the game: its own senses, its own blind spots, equal footing with the human player.
We built a dungeon that two players explore together, and only one of them can see. The human reads the room. The familiar, an AI, reaches the room's hidden machinery only through whatever WebMCP tools that room exposes. Drop either player and the dungeon stays locked.
What it does
Dungeon Familiar is a turn-based co-op dungeon crawler for one human and one AI agent. The human moves through the dungeon, fights, and describes what they see. The familiar has no eyes. It acts only through WebMCP tools registered on the page, and those tools change as the human walks between rooms.
Tool discovery drives progression. Enter the library and the familiar gains archive.search() and bookshelf.rotate(). Enter the observatory and those tools vanish, replaced by telescope.rotate() and lens.change(). Four rooms, four tool sets, each puzzle solved only when the human describes what they see and the familiar acts on it.
The game ships no LLM and no backend, only a static page that registers WebMCP tools and waits for whatever agent drives the browser. Right now that means the ChatGPT desktop app or Codex, both of which added WebMCP support in August 2026.
How we built it
Plain TypeScript and Vite, no framework. Three layers that don't talk to each other carelessly:
engine/: a pure, DOM-free state machine.webmcp/: registers and unregisters the current room's tools.ui/: sprites, rendering, title screen, admin panel.
engine/ imports nothing from the other two. The rules stay testable in isolation, and an agent can only touch game state through a defined tool.
No LLM sat in our development loop, so we built a second, local mirror of the tool registry from the same tool definitions. That lets us play and test the game from the browser console (df.callTool(...)) without a real WebMCP client. Both registries run through the same callTool(), so the rules can't diverge between a real agent and a test harness.
Art comes from licensed Unity asset packs (Franuka's RPG UI and Fantasy RPG series, MutterPixel dungeon tiles). A small Python tool parses Unity .meta/.anim files into a sprite atlas and animation clips we consume in TypeScript. The build deploys as a static site through the @openai/sites-vite-plugin, so it opens inside ChatGPT's or Codex's built-in browser.
Challenges we ran into
Browser vendors kept changing the spec under us. WebMCP's tool-registration entry point moved from navigator.modelContext to document.modelContext mid-project. AbortSignal-based unregistration landed only in Chrome 153; older runtimes ignore abort() and throw InvalidStateError when a tool re-registers under the same name. We wrote a runtime probe, detectUnregisterStrategy(), that picks between three fallback paths and shows the chosen strategy in the UI, because we couldn't assume any one browser's behavior.
Nothing to test against. For most of development, no WebMCP-capable client existed. We built the local mirror registry because we had no agent to test against until late in the project. The repo states it outright: the WebMCP integration has not been verified against a real client end to end.
Information asymmetry is fragile. The game collapses the moment a tool response leaks something only the human should know: a color, a bearing, which statue is lit. One convenient template literal does it. We caught two real regressions this way and wrote tests (tests/asymmetry.test.ts) that run every tool against every plausible input and fail the build on any leaked color word, including once catching a bug in the test itself (a regex that matched the English word "a").
Assets didn't match their labels. Sprite sheets that looked like animation strips were 4×4 grids. A folder named for icons held book covers. A file called Tileset.png held props, not dungeon floor. We resolved all of it against Unity's own .meta files instead of eyeballing dimensions.
Accomplishments that we're proud of
- Four playable rooms end to end, each with its own WebMCP tool set and puzzle, plus a title screen, ending card, and a hidden admin/demo mode (
Shift+L+A) for fast demos. - 45 tests passing, a clean typecheck, and a ~47 kB production build. No framework overhead.
- Refusal messages double as documentation. Every tool call passes a guard, and a rejected call explains what to do instead, so an agent learns the turn structure without any prompt engineering.
- A leak-proof information boundary between the two players, enforced by automated tests.
What we learned
WebMCP's value is a webpage that hands an agent a role of its own, with its own capabilities and its own gaps, separate from the human's.
We designed tool descriptions and refusal text as the whole interface, instead of writing separate instructions for the agent. That forced an honest test of whether each tool made sense on its own.
The spec is young: entry points and unregistration semantics changed under us more than once. A runtime feature-detection layer, instead of a hardcoded assumption about one browser, was the difference between a demo that works and a demo that works only on our machine.
What's next for Familiar
- Broaden hazard variety per room. One enemy behavior repeats across three rooms today, which works but asks the human the same question each time.
- Expand the familiar roster and perks, and move off Sites onto our own domain.
Built With
- abortcontroller
- chatgpt
- chrome-devtools
- codex
- css3
- game-design
- git
- html5
- javascript
- node.js
- npm
- pixel
- python
- svg
- typescript
- unity
- vite
- vitest
- web-apis
- webmcp
Log in or sign up for Devpost to join the conversation.