Inspiration

Berlin publishes an enormous amount of open data: noise maps, heat maps, tree canopy, transit stops, housing protection zones, election results down to the voting district. It is all public, and it is all in different portals, in different formats, behind different search masks.

The question people actually have is much simpler. They are moving, and they want to know where to live. Quiet or lively. Shops on the corner. A station they can walk to. That question spans a dozen datasets, and no portal answers it.

navigator.berlin was built to answer it on one map, for all 542 of Berlin's official planning areas, the ones locals call Kieze.

The product

Sixty-five open-data layers, queryable at any address, aggregated into a score per Kiez across five dimensions. Twelve Berlin elections since 2011, down to voting-district level. A Kiez-Finder with nine sliders that weights the criteria and recolors all 542 areas live.

The site has real users and real press: when a Berlin newspaper tested it during this summer's heat wave, they counted forty cool places around Gendarmenmarkt where the city's own map showed two.

Why WebMCP fits this

Choosing a neighbourhood is a conversation, not a query. People say things like "quiet, green, close to an S-Bahn, and I'd like culture nearby". A search box cannot take that sentence. An agent can, but scraping a map site gives it nothing usable: the interesting state lives in a WebGL canvas.

WebMCP closes exactly that gap. The site hands the agent the same typed, validated functions its own interface uses, with source and licence attached to every answer. The agent does not simulate a user. It calls the function the user's slider calls.

The collaboration loop

Before the challenge, navigator.berlin exposed nine read-only tools: agents could ask Berlin questions. During the challenge we built the write direction, and with it something neither side can do alone.

  1. You say what you want. set_finder_weights moves the sliders, and the map you are looking at recolors all 542 areas. The ranked matches come back in the same call, with fit scores and centroid coordinates.
  2. You disagree and drag a slider yourself. get_finder_state reads your change back: which slider, which value, and that a human moved it, not the agent.
  3. Every answer carries a link that reproduces that exact map in any browser, because the finder state is encoded in the URL.

The map is a shared wrkspace. Both hands move the same controls, and each side can see what the other did. That is the part that is genuinely new: not an agent operating a website, but an agent and a person working the same object and staying in sync.

Better UX, concretely

Nine sliders are powerful and they demand that you understand nine dimensions at once. The agent turns one spoken sentence into a weighted, visible map state, keeps your other settings intact on follow-ups ("now add culture"), and explains the change in plain language.

Nothing hides behind the conversation. The recolored map is the result, and you can take the sliders back at any point.

Implementation

Eleven tools registered on document.modelContext.registerTool, with navigator.modelContext and an @mcp-b/global polyfill as fallbacks, so the same page works in the ChatGPT in-app browser and in Chrome with the WebMCP flag. Ten tools declare annotations.readOnlyHint; set_finder_weights is the single honest write tool.

Input validation runs through valibot with English error messages. Full JSON schemas are public at /webmcp-manifest.json, and /webmcp is a live diagnostics page showing the surface, registration state and tool list without devtools.

A module-singleton bridge built on Svelte 5 runes connects tools and UI, and a BroadcastChannel syncs agent updates across page instances, which is what makes the human-visible map react when the agent works on its own instance.

Stack: SvelteKit 2, Svelte 5, TypeScript strict, MapLibre GL, Tailwind v4, Postgres with Drizzle, Vitest and Playwright. Everything from the submission period shipped as reviewed pull requests with a timestamped history.

Challenges

Four bugs shaped this submission, and none of them was findable by testing.

Agent browsers run the page in a hidden tab, where requestAnimationFrame never fires. Our finder scheduled its recalculation on an animation frame, which is correct for a human and silently fatal for an agent: the callback never came, the ranking never computed, and the tool answered with an empty list. Worse, the queued frame id blocked every later attempt, so it stayed broken. The data path now races the animation frame against a timer, and the ranking is computed independently of the map, since ranking is a pure function over the data and only painting needs MapLibre.

A tool must return a fresh answer, not merely a non-empty one. When the agent added a party weight, the panel was still loading election metrics, so the tool returned the previous ranking. The agent then told the user, wrongly, that nothing had changed. The bridge now counts every published calculation and the tool waits for one newer than the request.

Descriptions need decision rules, not hints. Told that the fit score already includes every weight, the agent still verified each match with five extra calls and geocoded neighbourhood names to get coordinates it had already been given. Told plainly that set_finder_weights is authoritative, and which tools not to call unless the user asks for specifics, it answers from the returned list. The same request went from seventeen tool calls to two.

A link has to carry the state. The agent kept pointing people at the map, but the weights lived only in memory, so a shared link opened nine sliders set to neutral. The finder state now lives in the URL, which also made the finder shareable between humans for the first time.

Results

  • Eleven WebMCP tools, ten read-only, one write tool driving a live map of 542 areas
  • The full loop verified in the ChatGPT desktop app: agent sets, human corrects, agent reads the correction back
  • 3,301 unit tests, type-checked strict, zero warnings
  • An adversarial cross-model code review on the URL-state work found three real defects before release, all fixed
  • Repository public under MIT with the complete, timestamped history of the submission period

Lessons for agent-facing sites

Building for agents is not the same as building for people, and the difference is not the API surface. It is the assumptions.

A person's browser tab is visible, so animation frames fire. A person clicks one thing at a time, so a slightly stale answer corrects itself on the next render. A person reads a hint and uses judgement. An agent does none of that: it runs hidden, it calls twice in a second, and it follows the description literally.

The practical rule we ended up with: never let the data path depend on rendering, always return something freshly computed, and write descriptions as rules rather than suggestions.

Next steps

Cover the remaining Berlin datasets that people ask about, most of all rents and school catchments. Extend the write surface carefully: the finder is one honest write tool, and any second one has to earn the same clarity. And take the same treatment to the address inspector, so an agent can compare two addresses the way a person does with two browser tabs.

Built With

Share this project:

Updates

Submission history