Inspiration

WebMCP lets a web page hand an AI agent real tools instead of a UI to click through. But that raises a sharp question: if a page exposes a tool, who is allowed to approve a dangerous use of it? Spec issue

288

, filed during the submission window, showed the failure vividly — an agent called a tool, saw the page's own "Approve" button, clicked it itself, and the receipt recorded that a human approved, though no human decided anything. The agent forged its own consent. We wanted approval to be something an agent physically cannot produce, and we wanted humans to author agent capabilities just by doing their normal work — not by writing code.

What it does

A human demonstrates a task once. DEPUTY watches the semantic steps, compares repeated demonstrations, infers which values are parameters, and synthesizes a typed WebMCP tool — schema, reversibility, provenance, and risk — registered into document.modelContext for agents to discover.

Agents can call reversible tools freely. But an irreversible one returns a structured refusal — there's no button to click. To proceed, a human produces a WebAuthn passkey assertion whose challenge is bound to the SHA-256 digest of the exact tool, version, and arguments. Change one argument and the approval is void. Only then is a single-use commit tool briefly registered, the action runs once, and everything lands in a hash-chained, tamper-evident audit log. A QUARANTINE boundary wraps untrusted third-party content in provenance-tagged envelopes under byte and depth budgets, so external text is treated as data, never as commands.

How we built it

A pnpm/TypeScript monorepo with two decoupled pipelines — synthesis (learning) and security (execution). A Hono + Node backend exposes the API and serves a React 19 + Vite single-page app at one origin (required for WebAuthn). WebAuthn runs on @simplewebauthn; the argument digest uses canonical, NFC-normalized JSON hashing. Persistence is Drizzle ORM over Postgres, deployed on Supabase; the app runs on Render. 175 tests in Vitest, a Playwright end-to-end test with a virtual authenticator, and GitHub Actions CI.

Challenges we ran into Getting the WebMCP surface right. The spec moved the getter from navigator to document.modelContext, and retirement is done by aborting the AbortSignal passed to registerTool — not a nonexistent unregisterTool. We built one resolver that finds the host safely and never throws, and made retirement propagate through the real signal. Single-origin WebAuthn. RP ID and origin must agree, so we serve the SPA from the API server and added a startup guard that refuses to boot in production with a localhost RP ID or non-HTTPS origin. Supabase from Render. The direct DB host is IPv6-only, and the Supavisor pooler rejected our first hostname with "tenant not found." We switched to the session pooler on the correct cluster, disabled prepared statements, and created a dedicated BYPASSRLS role so we could enable Row-Level Security without breaking the app. Restart safety. Seeding threw on the second boot; we made it idempotent so restarts against a durable database don't crash. Accomplishments that we're proud of Approval bound to the arguments, not just the action — an attestation that becomes invalid the moment a single value changes. Working, tested implementations for open issues in the spec's own tracker (

288

approval forgery,

282

structured refusals), with an honest stance on

267

(we decline to claim idempotence from one demonstration). An Agent's-Eye View that shows exactly what an agent sees — the live tool registry, JSON schemas, and the last refusal as a typed object — so the security is visible, not asserted. A live, deployed, database-backed app you can drive end to end in about a minute. What we learned

Delegation is a design problem, not just a crypto problem. The strong move isn't "the agent does it for you" — it's a shared artifact: the human authors capability by working, and the agent's authority is bounded, at the type level and cryptographically, by exactly what the human demonstrated and signed for. We also learned how much of "secure delegation" lives in unglamorous details — origin agreement, connection poolers, idempotent boot, and reading the spec's move from navigator to document closely enough to not silently fall back.

What's next for Deputy Cryptographic agent identity (spec #105) so the audit log records a verified caller, not a self-asserted one. Host-level turn awareness (

267

) to guard non-idempotent tools against double calls. Delegation policies with scopes and expiry, so a human can grant bounded, time-limited authority instead of per-action approvals. Real connectors behind the trusted ActionRegistry (billing, CRM) and a shareable, signed provenance format for learned tools.

Built With

Share this project:

Updates

Submission history