Signpost explores what happens when a user's compound request becomes a journey across independently operated WebMCP sites.

Inspiration

WebMCP gives individual websites a way to expose structured capabilities to browser agents. Those capabilities are exposed locally, while a user's compound objective can span several independently owned providers. I wanted to see whether a generic agent could compose that kind of journey without requiring any one provider, resolver, or platform to own it.

For the challenge, that meant separating the problem: providers declare what they can do, Signpost resolves where those capabilities are available, and the agent decides which capabilities it needs and in what sequence. Consequential providers retain control over whether their own mutations may execute.

The user can therefore begin with the outcome they want rather than a prearranged itinerary across specific sites: "order a coffee, book a basic haircut, and check this palette."

What it does

Signpost is a stateless capability resolver. An agent asks where a capability can be performed, and Signpost returns matching provider surfaces from declarations published by those providers.

Signpost never receives the user's full objective. It doesn't decompose the request, maintain a session, choose the next step, or own an itinerary. The agent remains responsible for deciding what capabilities it needs and what to do next.

For providers, this preserves a different boundary: they declare their own capabilities at their own origins rather than surrendering ownership of those capabilities to the system composing the journey.

For the challenge, I built three independently deployed WebMCP reference providers:

  • Deckhouse Coffee: order coffee
  • Chair & Comb: find and book an appointment
  • Hex Registry: check a hex value against a palette

In the live demo, ChatGPT receives one request spanning all three. It determines the capabilities it needs, resolves them through Signpost, navigates to each provider, discovers its WebMCP tools, and completes the journey.

For consequential actions, the agent can propose an action but cannot grant itself authority to execute it. Deckhouse Coffee and Chair & Comb display the exact material terms for user authorization at the provider. The pending invocation waits, then resumes and revalidates those terms immediately before execution. Only a valid, matching, single-use authorization permits the provider mutation.

Signpost also keeps evidence separate from agent narration. Participating surfaces record authorization decisions, provider calls, and provider results so the evidence view can show what actually happened without becoming part of the agent's planning state.

How I built it

Participating providers publish small capability declarations describing what can be done at their origin. Signpost fetches those declarations and exposes one WebMCP tool:

resolve_surface({ capability })

Matching is deliberately transparent: lexical retrieval with limited fuzzy matching, rather than another model in the loop. The result tells the agent where a capability is available, but never what it should do next.

Each reference provider is an independent WebMCP surface and knows nothing about the others. Consequential providers place authorization and execution control at their own mutation boundary.

The execution seam reuses @zioladev/[email protected], a small package I published before the challenge. Its rule is intentionally narrow: only allow permits execution; block and indeterminate do not. The exact-term authorization flow, await/resume integration, reference providers, trusted-activation guard, evidence capture, and multi-provider Signpost flow were built during the challenge.

I also included a deterministic drift test. It authorizes one set of material terms, presents different terms at execution, and verifies that the execution boundary returns BLOCK with zero provider calls.

Challenges and what I learned

The hardest challenge was not building too much into Signpost.

Once an agent begins moving across several sites, it is tempting to give the resolver a session, itinerary, cursor, or concept of "next." Those features make orchestration easier, but they also change what Signpost is. I moved those responsibilities back to the generic agent, so the resolver remained stateless and journey blind. Cross-origin execution made that boundary practical rather than theoretical. Each provider had to stand independently on its own origin, while the agent could return to Signpost whenever it needed to locate another capability.

The consequential flows evolved during the challenge as well. My first authorization design returned a "needs authorization" result and required the agent to retry after approval. I replaced it with await/resume: the original invocation now waits for authorization and then revalidates immediately before execution. That removed an unnecessary human-to-agent signaling round trip and narrowed the interval between authorization and use. That improvement made deliberate term drift difficult to demonstrate naturally in the live flow. Rather than weaken the runtime to manufacture a better demo, I moved that failure case into a deterministic test against the same execution boundary.

A much smaller bug reinforced why I wanted the resolver to stay inspectable. An early provider description used words such as "ordered" while explaining what the provider didn't do. The lexical resolver matched it to an order query anyway. It was doing exactly what its transparent scoring rules said it should do. The failure was easy to understand and correct without introducing an opaque ranking layer.

The larger lesson was that these responsibilities do not have to collapse into one agent system. A generic agent can choose the journey without becoming the capability registry, the source of execution authority, or the sole narrator of what happened.

What's next

Signpost is a reference implementation, not an attempt to prescribe how web-scale capability discovery must work. The broader question is what infrastructure an agent-operable web needs when agents begin moving across independently owned sites to complete a compound objective.

Signpost explores one possible separation: The agent owns why and sequence. Signpost answers where. Providers own what. Consequential providers control whether a mutation may execute. Evidence records what happened.

Built With

Share this project:

Updates

Submission history