Inspiration
The agent can see a support queue. It still has to guess the clicks.
Understudy is a support triage console, not a storefront. You do the ticket once: assign someone, fill a canned reply. The page turns that demonstration into a WebMCP tool and registers it on the tab. After that the agent calls the tool.
A server sitting beside the browser cannot watch you operate this tab, cannot inherit the session already in the page, and cannot grow this tab's tool list mid-session. The teaching has to happen here.
Independent measurement, not ours. The WindTunnel benchmark (nekuda-ai/WindTunnel) measured GPT-5.6 on native WebMCP solving 48/49 tasks at a median of $0.002 and 2,596 tokens. The same model driving the same sites via DOM and vision cost $49.87 total, with 29.3s median agent time. We built this so the cheap path is the one that ships.
What it does
Forty tickets live in the page. There is no backend.
A person does the job once in the live UI. The page generalises that trace into a parameterised tool and registers it on the tab. The human teaches. The agent writes the name and description. The human approves. Then the agent runs the same procedure on a different ticket by calling the tool instead of guessing clicks.
The agent talks to three meta-tools: understudy_list_recordings, understudy_draft_tool, understudy_publish_tool. Drafts never register. Publish holds until the human hits Approve or Reject on the overlay. After 45 seconds the tool returns awaiting_approval and tells the agent to ask the user to return. poll: true is the same shape without waiting.
Entity ids default to parameters, never frozen constants. Otherwise the tool keeps hitting last week's ticket and looks fine while it does it.
When a taught tool runs, a bottom toast names the current command and that step takes is-highlighted. Mutated tickets get a badge on the queue row: performed by agent via tool [name]. After the first publish, a count button opens the library (l): author, invocation count, success rate, last failure. Disable frees a host slot. Revoke aborts the tool. Fail-and-repair re-teaches one broken step.
Sharing is a Tool Pack file. Ctrl+Shift+E downloads versioned JSON (understudy-pack.json). Drop a .json pack onto the page to import. Import checks that every command id exists in this catalogue, then publishes. The human dropped the file, so import does not reopen Approve. Two normal windows on the same origin share IndexedDB. A real share demo uses a private window or another profile.
If WebMCP is missing, the console still runs for a human. document.body.dataset.webmcp is degraded. Tools do not register.
This checkout is local-first. Persistence is IndexedDB on this origin. There is no live URL in this text. Source: https://github.com/DITlieD/understudy
The demo video uses in-page document.modelContext. It is not ChatGPT chrome. Same registerTool and execute path. Open the app in ChatGPT's in-app browser, or in Chrome with chrome://flags/#enable-webmcp-testing, if you want that host.
How we built it
Boot is small. src/main.ts opens IndexedDB and calls bootApp. Claims below are that path.
The host is a typed command bus. Every state change goes through a catalogue command with a stable id and a JSON Schema payload. Recording captures that sequence. Replay dispatches the same commands the UI dispatches. A taught tool may dispatch only command ids that appeared in its recording. That allowlist is enforced in the compiler execute path.
Live catalogue:
- Read:
filter_tickets,get_ticket,list_templates,list_assignees - Write:
select_ticket,set_ticket_priority,set_ticket_status,set_ticket_assignee,set_ticket_tags,apply_template
Registration is document.modelContext.registerTool only. We do not call unregisterTool, navigator.modelContext, provideContext, or clearContext. Those are gone from the shipped surface.
Each published tool gets its own AbortController. Publish registers with that signal. Revoke and Disable abort it. Re-teach and Enable register again. A separate controller holds the three meta-tools for the life of the page. Restore from IndexedDB re-registers stored procedures on load.
The registry listens for toolchange and re-reads getTools() so the inspector and the context-budget meter stay truthful.
Annotations are computed from the command catalogue, never hand-set. readOnlyHint is true only if no step mutates. untrustedContentHint is true if a step returns ticket bodies or other user-generated fields. The spec has no destructiveHint. Destructiveness lives in the description and in the approval overlay.
Execute honours the incoming AbortSignal between steps. Output is compact JSON, projected under 1.5K. Chrome budgets (name 30, parameter description 150, tool description 500) fail validation before publish.
Elicitation is userland. The spec has no primitive. understudy_publish_tool holds a promise on in-page Approve / Reject. Timeout is 45 seconds, then awaiting_approval.
Hard cap: 16 taught tools. The three understudy_* meta-tools do not count against that cap.
Stack: TypeScript, Vite, Vitest, IndexedDB. Feature check: /probe.html reports whether document.modelContext is present, then tries a no-op registerTool named understudy_probe.
Nearby work records a demonstration and emits something consumed elsewhere. Understudy records a demonstration to grow this tab's own live tool surface. Schrute emits MCP tools that run on an external server. Narada is a Chrome extension for a proprietary agent runtime. Cursor learn mode and Lotus MCP emit skill files for a coding-agent harness. Preview-before-commit on a fixed tool list is already covered (webMCP-Legit-exploration). We change the tool surface itself.
Challenges we ran into
We built on the WebMCP surface that shipped, not the one we wished for.
There is no elicitation primitive, so approval is the in-page overlay and the execute callback holds a promise (or the agent polls with poll: true). There is no unregisterTool, so revoke only works by aborting the AbortController passed at register time. There is no focusTab(). If the tab is backgrounded, the human may never see Approve. After 45 seconds the gate times out and tells the agent to ask the user to return.
Tool calls have no transient user activation. No file pickers, requestPermission, popups, clipboard writes, or PaymentRequest from inside a tool. Pack import is a human drop of a .json file. There is no progress reporting and no streaming. One compact result per call. Also absent, so we did not design as if they shipped: agent identity on execute (second argument is only { signal }), outputSchema, destructiveHint.
The naive build is a click recorder. Clicks are brittle, and a judgement that never became a catalogue command cannot be captured. The quiet failure is a frozen ticket id: the tool looks correct and keeps hitting one row. Entity ids default to parameters because of that.
Procedures are linear step sequences with parameters. No loops, no branching, no expressions. Conditionals are the agent's job: compose more than one taught tool. Tools are tab-scoped. Persistence is IndexedDB on this origin, plus Tool Pack files. No accounts, no multi-user sync.
Too many tools and the agent picks worse. The registry refuses a 17th taught tool. Disable frees that host slot until Enable, or until the next load re-registers every stored tool. Import publishes at once. It does not re-open Approve.
The console has no conversation thread, no SLA engine, and no related-item graph. That trim is deliberate. Every live triage command is sensitive: false. Compile / publish do not refuse unexplained selection; the generalizer records those errors and the panel can show them.
Honesty for judges who open the repo: this is a local console. The video cut is in-page modelContext, not ChatGPT's browser chrome. Do not treat a missing live host as a ChatGPT demo.
Accomplishments that we're proud of
The loop works on a real triage console. Assign someone, fill a canned reply, the page registers a WebMCP tool, the agent runs it on a different ticket. Same commands the UI already dispatches.
Revoke is real teardown. Each taught tool has its own AbortController, and aborting that signal is how the tool leaves the host. The approval overlay is real userland elicitation. Pack files move procedures without a server. The console still runs when WebMCP is absent.
We stayed honest about the demo: local IndexedDB, in-page document.modelContext, no live URL claimed.
What we learned
Chrome's guidance is that descriptions decide whether the agent picks a tool. So the agent writes the name and description, and the human still gates publish.
Missing primitives are design inputs. No elicitation means an overlay and a 45 second timeout. No unregisterTool means revoke is abort. No focusTab() means a backgrounded tab can miss Approve.
If a judgement never became a catalogue command, it cannot be captured. The silent failure is a frozen ticket id.
Sixteen taught tools is a cap, not a suggestion. Tool-surface bloat degrades selection.
Do not overclaim the host. Local-first means IndexedDB and pack files.
What's next for Understudy
Tool Packs are the local-first version of team sharing. They already move procedures as versioned JSON. Import checks that every command id exists in this catalogue, then publishes.
The production shape is a procedure registry: per-team access control, review before a procedure becomes callable, and usage telemetry to retire dead ones. The eight-day plan put that production ACL out of scope. This build has no accounts and no multi-user sync.
The later registry can sit on what already shipped. Procedures are portable, declarative documents. Their only dependency is the command catalogue they were recorded against.
Built With
- chrome
- css
- html
- indexeddb
- javascript
- jsdom
- json
- node.js
- typescript
- vite
- vitest
- webmcp
Log in or sign up for Devpost to join the conversation.