Why this use case is a strong fit for WebMCP

MissionGraph lets a person collaborate with ChatGPT Work or Codex on the same live project page in the ChatGPT desktop app's built-in browser. The person uses the task graph on the page; through WebMCP, the AI agent can read that graph, inspect its event history, and take scoped actions without guessing from the DOM. When the person selects a task or relationship, the page exposes contextual tools for that exact selection.

This matters for work that continues between visits. The graph, approvals, and history remain available through the page, so the next agent turn can pick up from the current state instead of reconstructing context from a chat transcript.

How it creates a better user experience

A person can drag, connect, split, dispatch, and approve work on the canvas. ChatGPT Work or Codex can turn a goal into a dependency graph, add annotations, find ready or failed work, and explain what changed while the person was away. graph_digest returns the current graph state, while each tool result carries a cursor and a bounded changes_since delta of up to 50 recent events.

The AI agent can also work through an approval queue under a short policy written by the person and confirmed in the page. It cannot create that authority for itself. A confirmed policy becomes a session-bound grant with a limited lifetime and use count. Each use has a fresh nonce, and every policy-backed approval records the capability reference and nonce that authorized it.

What people and agents can do together that was difficult or impossible before

Running several agents in parallel is fast, but it is easy to lose track of who did what and why. Keeping everything in a narrated chat is easier to follow, but much slower. MissionGraph puts the work, the decisions, and the evidence in one shared graph.

Real Codex workers run in isolated Git worktrees and report lifecycle events, logs, and structured handoffs with short-lived credentials tied to one task. A Codex supervisor coordinates those workers and records each scheduling turn. The person works with ChatGPT Work or Codex in the built-in browser to inspect and supervise the same project from the same canvas.

How we implemented WebMCP

The main app registers 25 core tools and defines 5 contextual tools for the current selection and failure state. It uses the imperative API through document.modelContext.registerTool({ name, description, inputSchema, execute, annotations }) in app/src/webmcp/registry.ts.

At runtime, the page detects the registration features provided by the browser. With dynamic registration, only the contextual tools that currently apply are registered, each under an AbortController. A provideContext tier replaces the active catalog as state changes. The static fallback keeps all five contextual tools available and returns a clear not-applicable result when necessary. We verified the current dynamic path in the ChatGPT desktop app's built-in browser; compatibility with Chrome's WebMCP flag was tested during development.

Actions covered by the human-presence gate first open a visible confirmation showing the exact request, its SHA-256 binding, expiry, and use limit. A one-off action receives a single-use capability. A confirmed approval policy receives a bounded multi-use capability, with a new nonce for every use. The accepted event records the capability reference and nonce, never the bearer secret.

Inspiration

The idea came from coordinating long Codex sessions. Starting another worker was not the hard part. The hard part was returning later and answering basic questions: What changed? Which work is actually ready? What is waiting for my decision? Which action did I authorize?

That led us to use the page as the meeting point. On a later turn, ChatGPT Work or Codex can read the latest state, while the person can leave decisions and context in a form that will still be there. We started with software development because it exercises the hard cases: dependencies, parallel work, failures, handoffs, and approvals.

What it does

MissionGraph is a live task DAG shared by a person, an AI agent using WebMCP, and a worker fleet. The person arranges work and makes decisions on the canvas. Through WebMCP in the ChatGPT desktop app's built-in browser, ChatGPT Work or Codex can create a plan, inspect tasks and relationships, summarize recent activity, find ready or failed work, annotate the graph, and stage policies or consequential actions for visible confirmation.

When the bridge is enabled, a Codex supervisor receives structural events and coordinates workers in isolated worktrees. Workers report progress, handoffs, commits, and review requests back to the ledger. The bounded public judge fleet accepts confirmed dispatches for unchanged seeded tasks in first-in-first-out order. Tasks created or edited by a visitor remain supervision-only, and both the page and tool results say which mode will apply.

The active tool catalog follows the state of the page. Every result includes the current cursor and up to 50 recent changes_since entries. graph_digest also returns a full current-state summary, so ChatGPT Work or Codex can recover the current graph state even when older changes fall outside the 50-entry delta.

How we built it

The event-sourced server runs on Node 22 with Fastify and SQLite. Accepted events are appended to the ledger, a reducer folds them into current state, and updates reach the browser over WebSocket with SSE fallback. React Flow and elkjs render the canvas from that folded state.

The WebMCP registry supports AbortController-scoped tools, provideContext replacement, and a static fallback. A long-lived runtime waiter handles browsers that inject the WebMCP runtime after the page has already loaded.

A TypeScript bridge sends structural events to a Codex supervisor and executes its validated decisions. It can spawn, pause, resume, and terminate workers, and it can re-brief an idle worker thread. A live worker keeps its original brief until it exits.

Human confirmation is separate from any action proposed through WebMCP. Confirmation drafts are bound to the project, browser session, action or policy text, SHA-256 hash, expiry, and use budget. Capabilities are stored as hashes, and each accepted use is recorded with its nonce.

Challenges we ran into

Some problems were predictable: late WebMCP injection, differences between registration tiers, stale browser identities, and helping ChatGPT Work or Codex catch up on a later turn. Others only appeared with the real stack.

Three review rounds passed with mocks before the first real supervisor turn exposed a CLI flag regression. The ChatGPT desktop app's built-in browser rejected a pre-aborted registration probe that our mocks had accepted. An adversarial review found a prompt-injection path through model-supplied node IDs. On the production VM, the Codex sandbox failed because user namespaces were unavailable; a worker-mode boot probe was the only check that exposed it.

The bridge is deliberately at-least-once. Supervisor turns can be delivered again after a narrow crash window. Worker spawning is idempotent, while repeated pause or re-brief requests are bounded and tolerated rather than described as exactly-once.

Accomplishments that we're proud of

The current suites contain 87 server tests, 109 bridge tests, and 129 app tests, for 325 in total. A separate fleet-contract harness covers 19 scenarios. CI runs the package suites on pushes to main and track/**, and on pull requests.

The public seed is exported from a real run. It contains event history, handoffs, and approvals produced by real Codex workers, including the commit IDs recorded in their handoffs.

We closed the node-ID prompt-injection path at three boundaries: persistent identifiers follow a strict grammar, prompt fields are JSON-encoded, and live planning creates node IDs on the client instead of trusting model-supplied IDs. The final review also caught worker self-approval, handoff replacement during pending review, secret leakage into worker environments, and several judge-facing UI problems.

What we learned

Mocks were useful, but they were not enough. The most important failures appeared only in the actual browser, during a real Codex supervisor turn, or on the production host. We added real-browser, bridge, and live-worker checks as those layers became available, and kept the full trail, including false alarms, in PROGRESS.md.

What's next for MissionGraph

WebMCP does not currently provide subscriptions to page state. If a future browser API adds them, MissionGraph's ledger could supply the event stream.

Our next engineering step is leased server-side delivery for worker-control events. Beyond software development, the same task, handoff, and authorization model could also support research pipelines or operations runbooks.

Built With

Share this project:

Updates

Submission history