Inspiration

ONCE explores how one human-agent collaboration can become reusable procedural memory.

People repeatedly tell agents which inputs change, which rules remain fixed, which corrections are one-off, and when human judgment is mandatory. UI macros capture clicks and chat history captures prose, but neither creates a reliable shared procedure.

What it does

A human and an agent work together on the same live application state, with the human using the UI and the agent using WebMCP.

ONCE records their work as one readable semantic trace, then the human selects Teach this routine. ONCE deterministically separates:

  • changing inputs, such as budget and candidates;
  • durable policies, such as ordered criteria and approval requirements;
  • repeatable agent procedures;
  • outputs that must be generated again; and
  • example-only corrections that must not become general rules.

The routine then runs with new inputs. It restores the learned policy, starts with fresh evidence, scores, uncertainty, and recommendations, and pauses at the learned human approval boundary before completion.

The competition MVP uses deterministic Vendor Evaluation to demonstrate this interaction model. It does not claim arbitrary workflow learning or live vendor research. All four vendor dossiers contain fictional, first-party challenge fixtures.

How it works

Human UI actions and native WebMCP actions route through the same synchronous TypeScript semantic command bus. The bus validates each command, applies state changes, and records the actor, channel, phase, outcome, and a readable summary.

A deterministic, domain-specific compiler turns candidate names and budget into variables, preserves criteria and approval policy, compiles observed procedures, and excludes literal evidence, scores, recommendation text, and one-off corrections from durable rules. The replay engine instantiates that routine with new inputs and fresh outputs.

Only the human UI can Teach, start replay, approve, reject, or reset. Required criteria need a score of at least 3. An eligible recommendation before approval is rejected with APPROVAL_REQUIRED.

How WebMCP is used

ONCE uses native WebMCP directly through document.modelContext.registerTool(...) with no wrapper. Ten native WebMCP tools let an agent read and mutate the live application through named domain actions. Those handlers and the React human UI call the same semantic command bus, so agent and human work becomes one structured history.

WebMCP gives the agent semantic application actions rather than click or DOM recordings. That shared semantic substrate is what makes the collaboration structured enough for ONCE to compile into a constrained routine.

State, trace, the taught routine, and replay persist locally in the browser. ONCE has no login, database, backend service, embedded LLM, external research dependency, API keys, or paid service requirement.

Challenges

The hardest problem was deciding what ONCE should learn without overgeneralizing. Candidate names and budget are safe variables; criteria and approval are durable policy; evidence, scores, and recommendations are regenerated outputs; and a single human correction is recorded as Example only — not generalized.

A second challenge was enforcing a genuine approval boundary while keeping the flow deterministic and easy to inspect. Bulk agent tools also had to stay atomic while retaining granular semantic events.

Native WebMCP calls were verified against production in Codex's in-app browser. A separate ChatGPT Work invocation could not be independently verified because the client repeatedly blocked browser access with an admin-policy verification error.

Accomplishments

  • Completed the full Collaborate → Teach → Replay → Human approval → Complete loop.
  • Registered ten native WebMCP tools and routed them through the same command system as the human UI.
  • Verified native production calls, persistence, policy rejection codes, human approval and rejection, and two consecutive completion dry runs.
  • Passed 215 deterministic tests across seven files, plus lint, type checking, and the production build.
  • Shipped a publicly accessible HTTPS deployment with no login required and a public MIT-licensed repository with no required environment variables.

What we learned

Semantic domain actions are a stronger foundation for procedural memory than gesture recording. A credible learning system must show both what it retained and what it deliberately refused to generalize. Human approval is most meaningful when it is compiled as policy and enforced during replay, not added as a decorative confirmation at the end.

What's next

The next step is to explore other constrained domains where inputs, policies, generated outputs, and approval boundaries can be modeled explicitly. Future work could add human-reviewed routine editing and broader native-client verification while preserving the current claim boundary rather than pretending to infer arbitrary workflows.

Testing instructions

  1. Open https://once-webmcp.vercel.app/ in a WebMCP-capable browser and select Reset demo.
  2. Ask the agent: “In ONCE, evaluate Aegis Cloud and BeaconStack for a $24,000 annual budget. Use Security, Integration, and Cost as criteria in that priority order. Read the ONCE vendor dossiers, attach evidence, and score each candidate for every criterion using 1=does not meet, 2=materially below, 3=meets, 4=strong, 5=excellent. Flag any meaningful uncertainty and set an initial recommendation.”
  3. Make Security Required, keep priority 1, enable Require human approval before final recommendation, and correct BeaconStack's security evidence to emphasize that SAML SSO requires the Enterprise add-on.
  4. Select Teach this routine and review the variables, fixed policies, repeated procedures, generated outputs, approval gate, and example-only corrections.
  5. Open Replay with new inputs with Northwind AI, Orchid Systems, and an $18,000 budget.
  6. Ask the agent: “Run the active ONCE routine with the new inputs. Use the replay plan and ONCE vendor dossiers. Complete the required evidence and scores, then follow the routine through its approval boundary.”
  7. Review the paused evaluation and select Approve. The agent can then set the final recommendation and complete the replay. Reject ends the run without a recommendation.

The app should report Native WebMCP registered · 10 tools. Registration confirms browser capability; a successful tool call confirms agent invocation.

For Chrome testing, use Chrome 149+ with chrome://flags/#enable-webmcp-testing enabled and restart the browser before opening ONCE.

Built With

Share this project:

Updates

Submission history