Inspiration
Most WebMCP demos we looked at treat "the agent" as strictly upside: it fills things in faster, it clicks buttons for you. Nobody was building the other half of the story, the part where an agent's authority has to have a hard edge somewhere. High-stakes paperwork is the clearest place to make that argument: a visa or residency application has dozens of fields an agent can genuinely help with (dates, addresses, transcribing from a passport) and a handful it categorically should not be able to touch, because they're attestations. Signing your name is not a transcription task.
So we built a fictional application (the invented "Skyward Residency Program," no real agency or form referenced anywhere) specifically to make that boundary concrete and testable, not just asserted in a README.
What it does
Consequence is a long, branching application where every field carries a permission class: agent_fillable, human_only, or evidence_required. An agent can fill addresses, employment history, and dates directly. If it calls answer_question on an attestation field, it doesn't fail quietly, it gets back a structured error naming the field and telling it to hand back to the human. For evidence_required fields like a bank-balance summary, the agent can only propose_answer, which lands as pending confirmation, never as answered, until the human accepts it.
Every field also carries provenance (who set it, when, from where), and the whole mutation history is hash-chained and replayable, so a final review screen can show exactly who filled what.
Conditional branches make the same point about tool existence. Answering that the applicant is self-employed reveals a Self-Employment Details section, and only then do fields like business_income_summary become part of what get_section and list_open_questions read from. The tool surface reconfigures itself based on answers already given.
How we built it
- React + TypeScript + Vite on the client, Cloudflare Workers + a Durable Object as the authoritative backend, same architecture as our other entry, Cadence.
- webmcp-kit for tool definitions, shared across both apps.
- ~63 fields across 9 sections (7 base, 2 conditional), each declaring a
PermissionClassinseed/schema.ts. - One reducer (
src/shared/reducer.ts) enforces the permission class on every mutation path, client UI, Durable Object, and every WebMCP tool call alike, so there's exactly one place the boundary is defined and no path that can skip it. check_consistencydoes real cross-field validation: an expired passport, overlapping employment dates, an unexplained income gap, checks that are genuinely tedious for a person and genuinely tractable for an agent, run against real answers insrc/shared/derive.ts.submit_applicationis confirmation-gated and independently refuses while anything required is still open, on top of the per-field refusals.
Challenges we ran into
- Making the refusal mean something. It would have been easy to add a client-side check that just hides the "sign" button for agents. That's not a real boundary, it's a UI suggestion. The actual fix was routing every write, human or agent, through the same reducer, so the refusal is enforced at the one place a mutation can happen, not duplicated (and driftable) across the UI and the tool layer.
- Seed data with nothing to catch. Early on, the seed application started blank, so
check_consistencyhad zero contradictions to ever find, meaning the tool looked broken on every fresh demo even though the logic was correct. Fixed by seeding two fields that genuinely overlap (an employment start date before the prior job's end date), so the tool has real work to do the moment a judge opens the app. - Shared-board demo entropy, the same problem we hit on Cadence: because a Durable Object's seed loads exactly once, repeated testing against one shared instance permanently consumed the demo state. Fixed with per-visitor isolated application instances plus a visible in-app reset button.
Accomplishments we're proud of
- A refusal that's actually structural, not cosmetic:
answer_questionagainst ahuman_onlyfield is rejected by the same reducer a human's own edit goes through, verifiable by readingsrc/shared/reducer.tsdirectly rather than taking our word for it. - A tool surface that changes shape based on the application's own content, not just page navigation, extending the same dynamic-registration idea from Cadence into conditional business logic.
- A tamper-evident provenance trail (hash-chained mutation history) that answers "who actually filled this field" with cryptographic backing, not just a database column that could silently be edited later.
What we learned
The interesting design constraint wasn't "what can the agent do," it was "what can the agent be structurally prevented from doing, in a way that can't be bypassed by a differently-worded tool call." A permission check living only inside one tool's execute function is trivially circumvented by any other code path that touches the same state. Putting the enforcement in the shared reducer instead of the tool layer was the difference between a real guarantee and a suggestion.
What's next
A real evidence-extraction pipeline behind read_evidence (currently a clearly-labeled stub), a dedicated multi-party reviewer invite flow (the data layer already supports scoped reviewer tool grants), and extending the consistency checks beyond the current date/income rules.
Built With
- cloudflare-durable-objects
- cloudflare-workers
- css3
- html5
- model-context-protocol
- node.js
- react
- typescript
- vite
- webmcp
- websockets
Log in or sign up for Devpost to join the conversation.