Inspiration
Every engineering team I know wants an agent in the incident channel, and every one of them has the same objection: not the button. Reading dashboards, correlating a latency spike against a deploy log, drafting the status page, that's work a machine should do at 3am. Rolling back production is not.
The industry's current answer is a line in a system prompt: do not roll back without approval. That's a request, not a control. It's the security model of a sticky note on a server rack.
WebMCP suggested a different answer. If tools are registered from the page, and the page knows what state the incident is in, then permission stops being something you ask the model to respect and becomes something the tool surface simply doesn't offer.
What it does (→ why this is a strong fit for WebMCP)
SIGNAL is an incident response console that a human operator and their agent drive together. The active incident moves through four phases, and the agent's tool surface is recomputed on every transition.
In TRIAGE the agent has nine read-only tools: metrics, logs, deploys, topology, plus record_hypothesis to write its reasoning onto a timeline the operator can see. rollback_deploy is not in its list. Tell it to roll back and it doesn't refuse; it reports that it has no such capability, and calls request_escalation.
That paints a confirmation surface showing exactly which six capabilities would be granted and what the agent intends to do with them. A named human approves. toolchange fires. Six tools light up. Now it can act, each action still gated behind a confirmation showing the blast radius.
When the metrics recover, the agent calls declare_mitigated , which withdraws its own write access , and the surface contracts again to verification tools, then to postmortem tools.
This is a strong fit for WebMCP specifically because the permission boundary and the user interface are the same object. The operator isn't reading a description of what the agent can do; they're looking at the browser's own tool registry, rendered. getTools() is the source of truth for both the agent's abilities and the human's understanding of them. No server-side agent framework can offer that, because the human isn't there to see it.
How it creates a better user experience
Before: an operator with six dashboards open, correlating timestamps by eye under time pressure, while an agent in another window offers suggestions it has no way to verify and no way to execute.
With SIGNAL: the operator watches the agent investigate in a shared timeline, sees its hypothesis with the evidence attached, and makes one decision, do I trust this enough to unlock write access instead of twenty small ones. The agent handles correlation; the human handles judgement. Each is doing what it is good at, and the seam between them is a visible, auditable interface rather than a prompt.
The status page update is the sharpest example. The agent drafts; the operator gets that draft in a live text field and rewrites it; the tool resolves with their version, and the return value tells the agent it was edited and what actually shipped. That's not approval, it's co-authorship, and the agent's next update matches the operator's register because it can see what they changed.
What people and agents can do together that was difficult or impossible before
Three things, all specific to tools running in the page rather than on a server.
A tool can hand you a decision and wait. request_escalation's execute() renders a surface listing the exact capabilities to be granted and does not resolve until a human acts. A remote MCP server can return text asking a question. It cannot show you the blast radius, block, and continue with your answer, because it is not in the room.
A tool can come back edited. publish_status_update and draft_postmortem both return the human's revision, not a boolean. The agent learns from the diff.
The human can see the agent's power, live. The capability panel is fed by document.modelContext.getTools() on toolchange. When the phase changes, tools visibly light up or strike through. Nobody has been able to watch an agent's permissions change before, the registry was always somewhere else.
And the inverse: capability scoping by construction. Prompt-based restrictions fail under adversarial input; an unregistered tool has no attack surface. search_logs carries untrustedContentHint for the same reason , log lines are attacker-influenceable and should be treated as data.
How we implemented WebMCP
Imperative API throughout, 21 tools via document.modelContext.registerTool.
The core primitive is a useTool(def, enabled) hook that ties registration to React lifecycle through an AbortController. Phase is derived state in a Zustand store; each tool's enabled flag is a predicate over it. When the phase changes, React unmounts the effect, the controller aborts, and the capability genuinely leaves the registry. This is not a mock, the Tool Inspector extension shows the list shrinking.
toolchange plus getTools() drive the capability panel. AbortSignal is threaded from execute(input, { signal }) into every elicitation promise, so cancellation from either side tears down the dialog rather than stranding it. annotations.readOnlyHint marks the ten observational tools; untrustedContentHint marks log search.
Telemetry is a deterministic seeded simulation where metrics are a pure function of (service, region, time, mitigations applied), so rolling back the bad deploy actually recovers the curves. There's no second dataset. No backend, no network calls, nothing that can be down when a judge opens it.
Challenges
Getting the agent to reason about its own permissions rather than bumping into silence. Early on, told to roll back during triage, it would either hallucinate a tool or go quiet. The fix wasn't prompting, it was making get_incident_overview return the agent's current capability set and name request_escalation as the way to ask for more. Tools that describe the system's shape are as important as tools that act on it.
The second was tool selection between scale_service and rollback_deploy. Both plausibly address a saturated connection pool. It resolved by putting the reasoning in the description tha, t scaling relieves symptoms but won't fix a bad code path, rather than adding a rule.
Accomplishments
The moment the agent says "I don't have that capability in this phase, may I escalate?" and it's true, not a scripted refusal.
What we learned
Tool descriptions are the real interface. The behaviour we wanted came from sentences, not code. And the return value is a prompt: a tool that returns "OK" wastes the most valuable channel it has.
What's next
Real integrations behind the same phase gate, PagerDuty, Datadog, GitHub. The pattern generalises past incidents to any console with a destructive action: billing, admin panels, clinical systems. The interesting version is a policy engine that derives the tool surface from an org's existing RBAC, so the agent's capabilities are automatically a subset of the operator's.
Built With
- cloudflare-pages
- react
- tailwindcss
- typescript
- vite
- webmcp
- zustand
Log in or sign up for Devpost to join the conversation.