Inspiration I kept coming back to a specific failure. Ask an agent to help with a floor plan and it looks at a picture of one. It can tell you a desk "looks close" to a walkway. It cannot tell you the desk leaves 20 cm of centred clear width where 90 cm is needed, that the desk is desk-3 at (600, 130), that you locked desk-4 ten minutes ago, or that moving it costs two seats.
Accessibility planning is exactly the domain where that gap matters, because the answer is a measurement, not an impression. And it is also the domain where you least want an agent acting alone — a layout change is a physical, consequential decision that belongs to a person.
So ClearPath is built around one claim: the agent and the human should reason over the same structured model, and only the human should be able to commit.
What it does ClearPath models one classroom — North Hall, 9 × 6 m — in centimetres. Not a picture of a classroom: walls, an entrance, a route polyline, thirteen objects with footprints, rotations, locks and seat-capacity contributions, a 150 cm turning zone, and a door approach/recovery zone.
A deterministic geometry engine recomputes, on every change:
segment/rectangle intersection along the route centreline-to-obstacle distance and centred clear width destination-wall clearance with terminal-contact semantics turning-space overlap and entrance approach conflicts active capacity, route length, and changed-object count On the shipped fixture that yields a baseline of score 61, 20 cm centred clear width, 24 seats, 1,048 cm route, and one critical issue: route-corridor:desk-3.
A bounded beam search then looks for a way out. On this fixture the top result moves Desk 3 from (600, 130) to (530, 130) — 70 cm — reaching 90 cm centred clear width, clearing every issue, preserving all 24 seats, and scoring 98.
The agent gets twelve page-scoped tools:
get_plan_summary get_plan_geometry audit_access_routes focus_audit_issue set_planning_constraints generate_route_alternatives stage_route_proposal compare_plan_versions apply_staged_plan reject_staged_plan undo_plan_change get_audit_history apply_staged_plan is the important one, because of what it does not do. It never commits geometry. It records an approval request, opens the visible comparison, and returns approvalRequired: true, committed: false. Only a human clicking Approve and apply can change the committed plan — and the commit function re-checks the actor and re-validates the staged geometry at runtime, so the boundary is enforced in the domain layer rather than in the UI. Undo restores the exact stored prior plan rather than reconstructing it.
How I built it Geometry first, UI second. lib/planning-engine.ts owns the domain: types, geometry primitives, the audit, the scoring heuristic, and the search. It has no React in it. The SVG canvas renders straight from that model — the viewBox is plan.width × plan.height, the rectangles are the real object coordinates — so the picture and the tool output cannot drift apart, because they are the same numbers.
The score is explainable on purpose. It starts at 100 and subtracts 22 per critical conflict, 10 per review conflict, 0.25 per centimetre below the 90 cm threshold, 2 per changed object, and 4 per lost seat. I wanted a number a person could argue with, not a black box — and it is labelled a planning heuristic, never a compliance score.
Ranking is lexicographic, not weighted. Proposals sort by: route blockers → clear width met → clear-width deficit → other critical issues → review issues → capacity feasibility → capacity loss → number of changes → movement distance → route length. That ordering encodes a value judgement: capacity loss is ranked ahead of disruption. Lowering the minimum seat count to 22 makes a desk-removal option legal, but the 24-seat solution still ranks first. The agent is never nudged into quietly deleting a seat because you relaxed a constraint.
Search is bounded and deterministic. Beam width 72, max depth 4, 1,500-state evaluation cap, duplicate-state elimination, and every candidate validated for locks, bounds, overlap and capacity before it is scored. It runs in about 2 ms.
Verification is the deliverable. 76 Vitest unit and contract tests, 9 Chromium end-to-end workflows through a mock WebMCP host, and CI running typecheck, lint, tests, build and browser workflows on every push.
Challenges I ran into The destination was cheating. The route has to reach the presentation wall, so my first audit treated the wall as non-obstacle geometry — which quietly exempted a 500 × 55 cm physical object from every clearance check. The fix was terminal-contact semantics: the wall is audited on every non-terminal segment, and only the final 45 cm (one corridor half-width) of a verified perpendicular contact is exempt. A parallel, penetrating, or misplaced final segment gets no exemption at all. Getting that right was most of a day.
Two rounding paths, one measurement. Issues computed clear width as round(d × 2) while the metric tile computed round(d) × 2. For a clearance of 10.4 cm those give 21 and 20 — the issue card and the headline metric could disagree by a centimetre. In a tool whose entire pitch is "exact measurements," that is not a cosmetic bug. Now it rounds once.
My audit trail was lying. This was the one that stung. The undo_plan_change and reject_staged_plan tools called handlers that hardcoded actor: 'human'. So an agent-initiated undo was recorded in the history as a human action — in the project whose central claim is human-in-the-loop provenance. No test caught it, because the tests called the session functions directly and never went through the tool path. The actor is now threaded explicitly, and there are tests for both routes.
An accessibility tool with inaccessible text. I measured my own palette and eight muted greys failed WCAG AA at 9–12 px — the worst at 3.10:1 against 4.5:1 required. Fixing the colours was easy; making sure it stayed fixed was not. There is now a test that parses the JSX with the TypeScript compiler, resolves every text-[#…] literal against its nearest ancestor background, and fails below 4.5:1 — with icon-only exemptions listed explicitly at 3:1, and an assertion that every listed exemption is actually still in use, so stale ones cannot rot.
Cancelling destroyed the human's work. When I added per-execution AbortSignal support, aborting a search cleared the staged proposal and every alternative, and logged it as a failure. So an agent cancelling mid-search would silently throw away the comparison a person was reading. A cancellation is now a genuine no-op on state, with its own 'cancelled' result in the audit trail.
Cutting scope. ClearPath originally had a clinic and a café. They looked like three scenarios; they were one geometry engine wearing three skins. I deleted both. One honestly modelled classroom is stronger evidence than three demos that share coordinates.
What I learned Structured beats plausible. The moment tools returned real coordinates instead of prose, the agent's answers became checkable — and several were wrong in ways a screenshot would have hidden. Exactness is what makes an agent auditable.
"Human-in-the-loop" is a property you test, not a sentence you write. Mine was false in the audit log for days while the README described it correctly. If a safety boundary matters, there has to be a test that fails when it breaks.
Schemas are not validation. Every tool declares additionalProperties: false and bounded lengths, and every tool also re-validates inside execute, because the schema is a hint to the model and the execute boundary is the actual trust boundary.
Dynamic tool surfaces are the useful part of WebMCP. stage_route_proposal only exists after generation; apply_staged_plan and reject_staged_plan only exist while something is staged; undo_plan_change only exists once there is something to undo. The agent cannot attempt a nonsensical action because the action is not there — and proposal IDs are bound to a fingerprint of plan identity, geometry, locks, capacity and constraints, so a stale ID cannot stage or apply after anything changes underneath it.
What's next Rotated-polygon collision (rotation is modelled but collisions are axis-aligned), importing real plans instead of one fixture, jurisdiction-specific rule packs instead of transparent defaults, and multi-user persistence.
Honest limitations ClearPath is a planning aid, not accessibility certification, and no substitute for a qualified local review — thresholds vary by jurisdiction. It ships one deeply modelled scenario. The search is bounded and deterministic, not a general CAD optimiser. Collision checks are axis-aligned. State is local to the browser session. WebMCP itself is a draft API, and the automated end-to-end results come from a mock host — they are not evidence of native tool selection, and the evaluation document keeps those two things separate on purpose.
Built With
- accessibility
- ai-agents
- beam-search
- cloudflare-workers
- computational-geometry
- human-in-the-loop
- next.js
- playwright
- react
- svg
- tailwindcss
- typescript
- vite
- vitest
- wcag
- webmcp
- wrangler
Log in or sign up for Devpost to join the conversation.