The problem
Your accessibility CI is green. Your checkout is still broken.
Not because the rules are wrong — because of where they ran. Lighthouse and axe-in-CI load a cold, unauthenticated page and score the first paint. The defects that reach real users live somewhere else entirely: in step 3 of a logged-in flow, inside a modal that only exists after two clicks, in a computed colour that only resolves after the CSS cascade, in an accessible name that JavaScript removed during hydration.
You cannot audit that from outside the tab. There is no server that can see it, no scraper that can compute it, and no static analyser that runs the JavaScript which caused it.
DOM-Surgeon audits the page where the bug actually is — inside your running session, in the state you are currently looking at.
Why this is a strong fit for WebMCP
Everything being measured exists only in the live document: getComputedStyle after the cascade, the ARIA name after hydration, the real tab sequence, exceptions thrown in this session. WebMCP is what lets an agent execute inside that document instead of guessing at it from a screenshot or a DOM dump.
So the developer keeps their session — logged in, three clicks deep, mid-flow — and the agent inspects it in place. No fixture, no re-login, no reproduction script, no "works on my machine".
What people and agents can do together that was impossible before
You can prove a fix before you write it.
Every scanner tells you #pay measures 2.38:1 and stops. Ask DOM-Surgeon's agent to fix it and apply_fix walks the colour until it clears 4.5:1 in your live DOM, re-audits to confirm the finding cleared, and hands back the equivalent source change. You watch the defect go green in the browser, then take the diff to the codebase — instead of editing CSS, rebuilding, reloading, and re-checking on a hunch.
make_page_usable does it for every repairable finding in one pass. Verified against the demo: 6 findings → 0, then revert_fixes restores the original defective state exactly, so a before/after is one prompt apart. suggest_code_patch emits the source-level change list for the whole set.
That loop — human asks in plain language, agent measures runtime state the human cannot see, repairs it in the shared surface, and returns a committable diff — is the collaboration WebMCP makes possible. It is not a report. It is a fix you can watch working.
How it creates a better developer experience
Accessibility triage today means opening DevTools, hunting an element, reading a rule ID, and translating it into a source edit yourself. Here you ask:
- "Why can't a keyboard user finish this?" → the agent chains
get_focus_order,reproduce_interactionandget_console_log, and reports that a positivetabindexpulls the card field to tab position 1 and the submit handler throws. - "What's wrong with this modal?" → scoped audit of the state on screen, not the homepage.
- "Fix it and show me the patch." → live repair plus the diff.
Findings carry the measured ratio, the WCAG criterion, a unique CSS selector and the concrete fix — and highlight overlays bloom over the offending elements so you and the agent are pointing at the same button.
Closing the loop for coding agents
A terminal agent (Claude Code, Cursor, Codex) has a shell and a filesystem, not a document — so it can never call a WebMCP tool. That was the harder half of the problem, and it is why the audit engine is UI-free and also ships as a CLI:
npx dom-surgeon audit http://localhost:5173 --json --fail-on critical
Playwright loads your page, injects the engine, audits the real hydrated DOM, and exits non-zero. The agent writes UI, gets told which selectors fail and why, edits the source, and re-runs until it exits 0. This is the post-generation quality gate for generated UI — the thing that stops an agent shipping three unlabelled inputs and a 2.38:1 button. --baseline reports only new findings, so you can adopt it on a large existing app without a wall of pre-existing failures, and multi-route sweeps catch the "we only ever checked the homepage" gap. Same command works as a pre-commit hook or a CI step.
One engine, three surfaces: WebMCP tools for interactive triage in the browser, a CLI for the coding agent and CI, a bookmarklet for any page you can open.
How we implemented WebMCP
15 tools on document.modelContext.registerTool, each execute returning a plain string capped to the documented 1.5K output budget, with readOnlyHint on inspectors and untrustedContentHint on everything that reads live page text — a tool whose whole job is reading the DOM is a prompt-injection surface, and the spec's security guidance says to mark it as one.
Inspect: audit_page, get_a11y_tree, find_contrast_failures, find_missing_labels, get_focus_order, describe_element, get_console_log, reproduce_interaction, highlight_elements.
Repair: apply_fix, make_page_usable, revert_fixes, suggest_code_patch.
Adapt: describe_barriers, set_reading_preferences.
All native browser API underneath — getComputedStyle for post-cascade colour, ARIA reflection plus label-association walking for accessible names, real tabbability resolution (correctly excluding tabindex="-1", disabled and hidden inputs), the WCAG relative-luminance formula including the large-text 3:1 threshold, and a console.error/onerror patch installed at load. Zero backend, no runtime dependencies, one static bundle. Every capability is also reachable by button, so the app is useful with no agent attached.
Where this goes next
The same repair engine points at a second audience: the person a broken page is failing, who cannot fix the site and cannot wait for its next sprint. describe_barriers and set_reading_preferences exist for exactly that, and work today on any page the engine can be injected into. We are not overclaiming it — reaching that user at scale needs a browser extension so the engine loads without them installing a bookmarklet first, and that is the honest next milestone rather than something we are shipping here.
Honest limitations
Automated rules catch a real but partial slice of accessibility. Reading order, whether alt text is meaningful, and whether a flow makes sense still need a person — make_page_usable deliberately refuses to touch those and reports them for a human instead of silently papering over them.
We also pointed it at ourselves, and it found 11 defects in its own interface — including the primary "Make this page usable" button sitting at 3.77:1, which is an embarrassing thing for an accessibility tool to ship. All 11 are fixed; the self-audit now scores 0 while the demo's 6 seeded barriers remain intact, and npm run verify reproduces both numbers.
And a note on engineering method: verifying the repair cycle by running it rather than trusting a clean TypeScript build surfaced a real bug we would otherwise have shipped. cssPath was emitting #id without checking uniqueness, so a duplicated id — precisely the defect one of our own rules reports — produced a selector that resolved to the wrong element. Repairs were being applied to the wrong twin and revert could never find them again. Fixed, and the round-trip is now verified to restore identical state.
Built With
- aria
- css
- node.js
- playwright
- typescript
- vercel
- vite
- wcag
- webmcp