VERDICT — AI Decision Engine for WebMCP
Inspiration
AI can give you an answer. But should it be allowed to make the decision?
As AI agents become capable of researching information, using tools, and taking actions, I became interested in a boundary that becomes increasingly important as agents become more capable:
Where does recommendation end, and authorization begin?
A traditional decision application can calculate a score and recommend an option. A chatbot can explain its reasoning. But neither necessarily provides a structured process for an agent to evaluate, challenge, negotiate, and then hand control back to a human before a consequential decision is committed.
I wanted to build that missing layer.
That idea became VERDICT.
"Let AI do the reasoning. Make AI defend the reasoning. Keep final authority with the human."
WebMCP made this particularly interesting because it allows websites to expose structured capabilities directly to AI agents instead of forcing agents to infer how to operate a user interface. The official challenge specifically asks builders to explore what becomes possible when humans and agents can interact and collaborate through WebMCP.
Rather than expose a generic "make a decision" action, VERDICT exposes the actual stages of its decision process as native WebMCP tools.
That led to the core protocol:
Evaluate → Explain → Challenge → Negotiate → Authorize → Commit
What it does
VERDICT is an AI Decision Engine built around native WebMCP.
An external AI agent can interact with the live application through six structured WebMCP tools:
research_options— retrieve available options and decision-relevant data.score_options— calculate weighted scores and rank the available options.challenge_top_pick— challenge the current winner against the runner-up, exposing score gaps, vulnerabilities, counterarguments, and tipping points.adjust_priority— change the importance of a decision criterion and recompute the result.request_commit— prepare the current recommendation for commitment while placing it behind a human-approval boundary.commit_decision— finalize the decision only after explicit human authorization has been granted.
The agent is therefore not given a single autonomous "make decision" button.
Instead, it works through a structured decision process.
A recommendation is not the same as a decision
In the recorded hardware-selection workflow, the agent initially recommended Laptop B with a score of 91.0.
It then challenged its own recommendation.
After the agent changed the decision priorities, Laptop B remained the winner at 89.1. However, the challenge engine surfaced an important counterargument:
Laptop A had a 5-point performance advantage over Laptop B.
The engine also calculated that if performance became sufficiently dominant — approximately 78% priority in that decision state — Laptop A would overtake B.
That is more informative than simply saying: "Laptop B wins."
VERDICT can instead expose the conditions behind the recommendation:
"Laptop B wins under the current priorities, but Laptop A has a significant performance advantage and would become the preferred option if performance became sufficiently important."
The objective is not to manufacture uncertainty. It is to make the recommendation's assumptions, alternatives, and sensitivity visible.
Human authorization is a separate step
After the agent requested commitment, VERDICT entered a pending human-approval state.
The agent then attempted to call commit_decision without approval.
The application rejected the operation with:
HUMAN_APPROVAL_REQUIRED
No committed decision was created.
Only after the human explicitly clicked APPROVE DECISION could the agent successfully commit the decision and produce a frozen decision record containing the final decision state.
This is the central rule of VERDICT:
The agent can recommend. The agent can challenge. The agent can negotiate. The agent cannot authorize itself.
The authorization boundary is enforced by VERDICT's application logic; WebMCP provides the structured interface through which the agent interacts with those application capabilities.
More than one decision domain
VERDICT currently demonstrates the same decision protocol across two scenarios:
Hardware Selection
"Which laptop should I choose for AI development and daily commuting?"
Housing Decision
"Which apartment should I lease for hybrid work and city commute?"
The laptop scenario is the primary recorded external-agent demonstration.
The housing scenario demonstrates that the underlying engine is not fundamentally a laptop selector. The reusable primitive is the decision protocol itself:
Evaluate → Explain → Challenge → Negotiate → Authorize → Commit
How we built it
Shared decision state
The most important architectural decision was creating a single canonical decision state.
Both the human interface and the WebMCP tools operate against the same application state.
That state contains:
- current scenario
- available options
- decision criteria
- priority weights
- scores
- rankings
- challenge results
- approval state
- commitment state
- final decision record
When an agent changes a priority, the resulting state is reflected in the application. When a human changes the decision state, subsequent agent operations reason over that updated state.
This avoids maintaining a separate agent-side representation of the decision.
Native WebMCP tool layer
VERDICT registers six native tools through the browser's model context API using the WebMCP registration pattern:
document.modelContext.registerTool(...)
Each tool has a defined name, description, input schema, and execution function.
The tools are intentionally separated around meaningful application capabilities rather than exposing one large autonomous operation.
The resulting agent workflow is:
Research → Score → Challenge → Adjust → Request Approval → Commit
This makes the interaction between the agent and application observable and gives each capability a clear responsibility.
Deterministic decision engine
The decision engine handles:
- weighted scoring
- option ranking
- dynamic weight normalization
- adversarial comparison
- counterargument generation
- sensitivity analysis
- tipping-point calculation
The scoring itself is deterministic for a given decision state.
The tipping point is calculated algebraically from the scoring model rather than relying on arbitrary trial-and-error values.
This means the system can explain not only which option currently wins, but also how the recommendation changes when priorities change.
Authorization boundary
The commitment workflow deliberately separates request_commit from commit_decision.
request_commit prepares the recommendation. It does not grant authorization.
Before commit_decision can create the final record, the application checks whether explicit human approval has been granted.
The important design principle is therefore enforced in application state, rather than being left to an instruction such as: "Please ask the human first."
That makes the authorization boundary part of the application's behavior.
External-agent verification
The deployed application was connected to an external Gemini agent.
The agent discovered the six WebMCP tools exposed by VERDICT and used them against the live application to execute the decision workflow:
Research → Score → Challenge → Adjust priorities → Re-score → Challenge again
→ Request approval → Attempt unauthorized commit → Human approval → Commit
The unauthorized commit attempt was rejected with HUMAN_APPROVAL_REQUIRED.
After explicit approval through the UI, the final commitment succeeded.
This is the key WebMCP demonstration: an external agent can discover and operate the application's structured capabilities while the human remains the authority over the final consequential state.
Challenges we ran into
Keeping the UI and agent synchronized — the hardest engineering problem was ensuring that human interactions and agent tool calls operated on the same canonical state. A separate state model for the agent would have made the demonstration easier to fake but weaker as a product. Instead, the UI and WebMCP layer were designed around the same decision store. That made state synchronization an architectural requirement rather than a visual effect.
Designing the right WebMCP tools — the simplest implementation would have been one large make_decision() tool. But that would hide the most interesting part of the agent interaction. Instead, I decomposed the process into six capabilities with clear boundaries. This makes the agent's interaction with the application understandable and gives the human a clear authorization boundary.
Preventing autonomous commitment — a capable agent should be able to perform substantial analytical work without automatically receiving permission to finalize the result. The solution was to separate preparation from commitment. Testing this boundary was particularly important because the expected failure is actually evidence that the boundary is working: HUMAN_APPROVAL_REQUIRED.
Making recommendations challengeable — a weighted scoring model can create a false sense of certainty. The challenge engine therefore asks: What is the strongest argument against the current recommendation? The tipping-point calculation shows when a change in priorities would reverse the recommendation. This turns the decision from a static score into something that can be inspected and questioned.
Accomplishments that we're proud of
1. Turning WebMCP into a complete decision workflow Six native tools form an end-to-end interaction model: Research → Score → Challenge → Negotiate → Authorize → Commit. This is the core of the project.
2. Demonstrating an external agent operating the live application An external Gemini agent discovered the native WebMCP tools exposed by the deployed VERDICT application and used them to execute the decision workflow against the actual live decision state.
3. Making the recommendation defend itself VERDICT does not stop when it finds a winner. The challenge engine explicitly looks for arguments that weaken the recommendation. In the recorded run, the challenge surfaced Laptop A's 5-point performance advantage and identified the approximate priority at which that advantage would change the outcome. Not just an answer — a defensible answer.
4. Enforcing the human authorization boundary
The external agent attempted to commit before receiving approval. VERDICT rejected the operation with HUMAN_APPROVAL_REQUIRED. After the human explicitly approved the decision, the commit succeeded and produced an immutable decision record.
5. Building beyond a single demo scenario The same underlying protocol works across the hardware-selection and housing-decision scenarios. The laptop is the demonstration domain. The decision engine is the product.
What we learned
The biggest lesson was that agent capability is only one part of agentic product design.
The more important question is:
What should the agent be allowed to do, in what order, and where should human authority begin?
Building VERDICT made me think about WebMCP as more than a way to expose functions to an agent. A good WebMCP interface should expose the application's meaningful capabilities and make the boundaries between those capabilities understandable.
I also learned that recommendations should expose their assumptions. A winner without context is often less useful than a winner accompanied by: the priorities that produced it, the score gap, the strongest counterargument, the criterion where the runner-up has an advantage, and the condition under which the recommendation changes.
Another important lesson: human-in-the-loop does not have to mean human-doing-everything. The agent can perform substantial work — research, analyze, compare, challenge, negotiate, prepare — while the human retains authority over the consequential final action.
That feels like a more useful model for capable agents than either extreme: manual everything or autonomous everything.
What's next for VERDICT
1. More decision domains — travel planning, vendor selection, purchasing, project prioritization, hiring workflows, resource allocation, and career decisions. The goal is to separate the reusable decision protocol from the underlying domain data.
2. Richer adversarial reasoning — multiple challenge perspectives including worst-case analysis, hidden-assumption detection, sensitivity analysis, alternative objectives, risk-focused challenge, and confidence analysis. The objective: make an AI recommendation harder to accept blindly.
3. Decision history and replay — a structured audit trail showing who changed what, how the recommendation changed, what the agent challenged, and what the human approved.
4. More granular authorization — permission levels such as Read → Analyze → Negotiate → Prepare → Request Approval → Commit, allowing applications to expose increasingly powerful agent capabilities while maintaining precise control over consequential actions.
5. Multi-agent decision negotiation — multiple specialized agents reasoning against the same canonical decision state while VERDICT coordinates the process and the human remains the final authority.
The Bigger Vision
As websites become increasingly accessible to AI agents, applications will need more than buttons and APIs.
They will need clear agent capabilities, understandable state transitions, meaningful boundaries, and explicit authorization models.
WebMCP creates an opportunity to make those capabilities directly accessible to agents.
VERDICT is my exploration of what that could look like for decision-making.
The goal is not to make AI less capable.
It is to make increasingly capable AI more inspectable, more challengeable, and more accountable.
Don't just ask AI what to choose. Make it defend the choice.
Built With
- ai
- gemini
- javascript
- openai
- react
- typescript
- webmcp
Log in or sign up for Devpost to join the conversation.