Inspiration

Incident response still forces humans to translate between dashboards, browser consoles, agents, runbooks, and change controls while an outage is unfolding. An agent can investigate quickly, but giving it broad action authority makes the human approval step vague or performative. We wanted the browser itself to expose a precise operational contract: enough authority for an agent to gather evidence and stage an exact response, but no authority to approve its own consequential action.

Runbook Zero is my answer. It treats WebMCP as a state-aware incident control plane and keeps the product interface as the shared source of truth for both the operator and the agent.

What it does

Runbook Zero is an installable incident product for the website an operator is actually using.

  • The Codex plugin begins read only, inspects bounded evidence from the selected site, inventories its WebMCP capabilities, and derives an issue specific component, dependency, user flow, telemetry, change, and evidence graph.
  • The generated Incident Pack opens in the Runbook Zero workbench, where real document.modelContext tools let an agent inspect services, trace paths, query signals, record a hypothesis, compare mitigations, and stage one exact action.
  • Every tool call updates and focuses the same topology, telemetry, evidence, mitigation, provenance, and capability interface the human sees.
  • The visible Capability Firewall shows active agent capabilities, unavailable consequential capabilities, application phase, and authority state.
  • apply_approved_mitigation is not registered before approval. The agent has no approval tool. A human must review the exact target, diff, risk, reversibility, assumptions, incident, and seed, then click the visible approval control.
  • Only after that click does apply appear, narrowed to the exact approved mitigation. It disappears immediately after use.
  • For a live target, Runbook Zero releases an origin-, tool-, and input bound receipt that Codex carries back to the target site. If no matching action exists, it produces an honest operator handoff instead of pretending automation happened.

The product also includes three polished deterministic packs—checkout pool regression (INC-042), payment queue backlog (INC-117), and catalog cache stampede (INC-203)—plus safe local JSON import. They all run through the same domain and WebMCP contracts. The canonical pack makes the judging flow reproducible the installed plugin proves Runbook Zero is not limited to a premade walkthrough.

How we built it

The workbench uses React 19, TypeScript, Vite, Zustand, a guarded state machine, and a validated Incident Pack v1 schema. Thirteen page defined WebMCP contracts are registered with document.modelContext.registerTool. The active subset is derived from incident phase, the current pack, approval state, and execution mode; stale registrations are cancelled with AbortController whenever that state changes.

The repository hosted Codex plugin defines a conservative Site Capture v2 contract. It treats page text, tool descriptions, and tool results as untrusted evidence, rejects browser secrets and invalid cross references, and deterministically builds a graph and provisional diagnosis from the evidence actually observed. Live action receipts bind the incident, seed, origin, target tool, and exact JSON input. Post action evidence must still satisfy the pack's recovery thresholds before an incident can resolve.

We validate the product with 17 Vitest files / 69 tests covering domain commands, state transitions, pack validation, dynamic registration, stale tools, the approval invariant, imported-pack failures, deterministic reset, and UI behavior. Four Playwright regression journeys cover reset-to-resolved, multi-pack behavior, keyboard and viewport usability, and live-site receipts. Real local, deployed in-app-browser, installed-plugin, and Chrome WebMCP evidence is recorded separately from the test-only browser harness. The public AGPL-3.0 app is deployed on ChatGPT Sites.

Challenges we ran into

The hardest challenge was making human authority a runtime property not a disabled button. Rejecting an early apply call was not enough, the consequential tool had to be completely absent from discovery until the exact staged object had visible human approval, and old handles had to become stale as soon as state changed.

The second challenge was cross origin honesty. Runbook Zero cannot claim that its page silently changed an unrelated website. We separated approval bound release from target site execution and verification, with an operator handoff fallback when the site exposes no applicable action.

The third challenge was escaping a scenario specific demo without sacrificing a deterministic judging path. Generalizing the schema, UI, tool inputs, graph layout, thresholds, and recovery engine.

Accomplishments that we're proud of

  • WebMCP is the product's control plane.
  • The same thirteen tool contracts operate against bundled, imported, and evidence derived live site incidents.
  • The agent and human share one synchronized operational interface with explicit Agent, Human, System, and Change provenance.
  • Pre-approval application is impossible through the registered WebMCP surface. Approval itself is human only.
  • The Capability Firewall makes the transition from locked apply to exact approved apply visually obvious.
  • New issues produce new evidence graphs instead of replaying a fixed demo.
  • Live actions use exact origin/tool/input receipts and an honest handoff fallback.
  • The canonical workflow resets deterministically and resolves through fixed recovery frames.
  • The public app, source, AGPL-3.0 license, installable Codex plugin, tests, and judging evidence are reproducible from the repository.

What we learned

WebMCP capability discovery can also be policy. The safest consequential tool is not one that merely promises to reject unsafe input. It is a tool that does not exist until the surrounding human-visible state makes its use valid.

We also learned that agent legibility matters as much as agent capability. When tool calls update the interface already in front of the operator, the human can see what the agent inspected, why it formed a hypothesis, what exact change it staged, and which authority boundary remains. Finally evidence topology is a far stronger bridge from demo to product than a library of hardcoded incident scripts.

What's next for Runbook Zero

Next we will validate more real sites, add capture adapters for richer browser and observability evidence, and expand the portable policy and receipt model. Production integrations will follow only where they can preserve the same boundary: read only observation, evidence backed diagnosis, exact staging, visible human approval, origin bound execution, independent verification, and auditable closure.

Built With

  • agentic-web
  • agpl-3.0
  • ai-agents
  • ai-safety
  • browser-automation
  • change-management
  • chatgpt
  • codex
  • developer-tools
  • devops
  • explainable-ai
  • human-in-the-loop
  • incident-management
  • incident-response
  • infrastructure
  • observability
  • open-source
  • react
  • site-reliability-engineering
  • sre
  • trust-and-safety
  • typescript
  • webmcp
  • workflow-automation
Share this project:

Updates

Submission history