The problem

A release manager wants an AI agent to help with one exact release. The failure mode is simple: the human approves it, then the evidence or the decision changes, but the agent still retains the consequential capability.

EvidenceBound makes the capability follow the current human decision. The agent starts with three safe WebMCP tools. It can inspect the release and request approval, but it cannot approve, revoke, correct, or restore its own authority. When the release manager approves the exact current release snapshot, a fourth tool — execute_authorized_release — appears. If the human corrects the security evidence or revokes approval, that tool disappears. Execution also re-checks current authority, so stale discovery state cannot succeed after the decision changes.

Judge moment: 3 tools → human approval → 4 tools → human correction/revocation → 3 tools.

The controlled demo records externalSideEffect=false; it does not claim to perform an actual external deployment or customer mutation.

What the judge will see

  1. The agent starts with three safe tools and inspects the current release.
  2. The agent requests approval. The request becomes pending, but the AI still cannot approve itself.
  3. The human clicks Approve this exact release.
  4. execute_authorized_release appears as the fourth WebMCP tool.
  5. The agent invokes the controlled action; EvidenceBound revalidates current authority before recording a receipt.
  6. The human corrects the security evidence or revokes approval.
  7. The release tool disappears. A stale retained descriptor is rejected by execution-time revalidation.

The public page separates Human and Agent actions and keeps the next step visible so this behavior can be understood without reading the architecture first.

Why WebMCP is load-bearing

This is not a generic MCP wrapper. EvidenceBound uses the real browser-native document.modelContext.registerTool(...) interface so current human authority changes which capability the agent can discover.

Three baseline tools are available for safe inspection, approval requests, and receipts. execute_authorized_release is registered only while the exact current release snapshot is AUTHORIZED. Correction, revocation, stale evidence, or a hard evidence/policy failure removes or blocks that capability.

That is different from exposing a consequential tool permanently and hiding a denial inside the handler. Here the human decision changes the model-visible browser capability surface itself. Execution-time revalidation is a second fail-closed boundary in case an agent retained stale discovery state.

Human-only grant, revoke, correct, and restore controls are deliberately not WebMCP tools. The AI cannot mint or repair its own authority.

Human + agent collaboration

The useful division of labor is explicit:

  • Agent: inspect the release, explain what is missing, request approval, and use the release capability while it is current.
  • Human: approve one exact release, correct evidence, revoke approval, or restore fresh evidence.
  • EvidenceBound: project the current human decision into the WebMCP tool surface and revalidate it at execution time.

The result is easy to see: the human changes their mind, so the agent's consequential tool disappears.

Who this is for

The primary user is release engineering / DevOps / platform operations. Their recurring job is letting agents help with consequential operational actions without letting yesterday's approval become today's authority.

A common alternative is to expose sensitive tools continuously and rely on prompts or a late authorization check during execution. EvidenceBound adds an authority-aware capability layer: the agent sees the consequential tool only while current evidence and explicit human approval justify it, and execution verifies again before success.

A reusable version of this pattern could support release and administrative workflows that need revocable, evidence-bound agent capability. We are not claiming current customers, revenue, certifications, or production external deployments.

Implementation

The WebMCP surface lives under site/webmcp/. A dedicated adapter feature-detects document.modelContext; when the real WebMCP API is unavailable, the app does not emulate or label a fallback as genuine WebMCP.

Internally, EvidenceBound distinguishes AUTHORIZED, HUMAN_REQUIRED, STALE, INVALIDATED, and BLOCKED. Tool registration follows that state. Execution revalidates authoritative state and includes race-condition hardening for authority correction during asynchronous verification and concurrent execution attempts.

Automated tests cover state transitions, stale and invalidated authority, execution-time revalidation, concurrent execution behavior, and WebMCP registration/removal contracts. GitHub Actions runs both the WebMCP-specific suite and the broader EvidenceBound Core CI. The public site is deployed on Vercel and the repository is Apache-2.0 licensed.

What was added during the challenge

EvidenceBound Core existed before the WebMCP Challenge. It already contained evidence, provenance, policy, stale-state, invalidation, receipt, persistence, and fail-closed primitives.

Challenge-period work added the WebMCP-specific authority projection and judge experience:

  • browser-native WebMCP tool registration;
  • dynamic 3 → 4 → 3 tool visibility driven by current human authority;
  • agent request vs. human grant/revoke/correct separation;
  • execution-time authority revalidation;
  • correction/revocation propagation into the visible capability surface;
  • adversarial stale, correction-race, and concurrent-execution tests;
  • a public WebMCP judge surface and production acceptance evidence.

The repository documents pre-existing work and challenge-period work separately so prior EvidenceBound primitives are not presented as new competition work.

Try it

Open https://evidencebound.org/webmcp/ in ChatGPT's in-app browser or Chrome with WebMCP enabled. The page contains the 60-second judge path and links to the public source and judge guide.

Built With

Share this project:

Updates

Submission history