Inspiration

AI evaluations can disagree even when testing the same models on the same benchmark. Small protocol differences like reasoning budget, parsing, retries, or tool access can change the conclusion.

We wanted to answer one question:

What is the smallest protocol change sufficient to reproduce the conflicting result?

What it does

Minimal Protocol Witness (MPW) compares conflicting evaluation setups through controlled counterfactual tests.

In our demo, Lab A concludes MODEL_A while Lab B concludes MODEL_B on the same 400-item synthetic benchmark.

MPW tests protocol changes, finds the minimum witness, verifies it across the full exposed protocol space, and produces a replay-verifiable certificate.

How we built it

MPW uses React, TypeScript, Vite, and WebMCP.

It exposes four semantic WebMCP tools:

  • read_dispute
  • run_counterfactual
  • inspect_evidence
  • verify_witness

The agent decides what to investigate, while a deterministic engine computes results, verifies minimality, and issues certificates.

Challenges we ran into

The hardest part was making the result rigorous instead of just visually convincing.

Protocol effects can be non-monotonic, so we could not assume that adding more changes preserves a result. We built exact search and verification instead.

We also had to keep agent reasoning separate from scientific verification.

Accomplishments that we're proud of

  • Exact minimum-witness verification
  • Full 16-subset verification in the canonical example
  • Shared state between humans and WebMCP agents
  • Experiment-ID-bound evidence inspection
  • Replay-verifiable certificates
  • Extensive automated and adversarial testing
  • Explicit limits on causal and universal claims

What we learned

WebMCP is most powerful when a website exposes meaningful domain operations, not just UI actions.

We also learned that agents and deterministic systems work best together: the agent explores, while the system verifies.

What's next for Minimal Protocol Witness

Next we want to test MPW on real model evaluations, larger protocol spaces, more external harnesses, and repeated stochastic runs.

Longer term, we want evaluation reports to become machine-inspectable and counterfactually testable by both humans and agents.

Built With

Share this project:

Updates

Submission history