Inspiration

Coding interview transcripts takes weeks, and most of it is clerical work. The part that matters is small: the handful of quotes that could honestly belong to two different themes. Choosing between them isn't data entry. It's the argument the paper will make.

Every AI tool I tried resolved those quotes for me. Hand over a corpus, get back clean labels, and the disagreement that should have shaped the analysis is gone, filed under whichever theme won by a rounding error.

I wanted the opposite: the machine grinding through the volume, and the disagreements coming back to me instead of getting quietly filed.

What it does

Sensemaking Studio is a shared canvas. A researcher drags quotes between themes. An agent works on the same document at the same time: it reads the whole corpus, finds quotes between two themes, ranks them by how contested they are, and presents them to the researcher.

Then it stops. The agent has no way to settle a contested quote. This isn't a policy or a line in a prompt. The capability does not exist. It can argue for a reading. It cannot record one.

Why WebMCP fits

An agent that drives a web app by clicking has to guess what it is looking at, and everything it does arrives as an anonymous mutation. The record cannot tell a researcher's judgment from a stray click.

In qualitative research, that record is the product. An analysis is defensible because you can show who decided what and why. Lose the provenance, and you have lost the thing you were making.

WebMCP lets the page hand off the operations to an agent instead of the buttons. When the agent moves a quote, it invokes the same command that the researcher's mouse does, so both appear in one log with the author attached. That shared path is also why the guardrail holds. There's no backdoor into the analysis, so the agent's reach stops at the commands that exist.

What people and agents can do together now

Scoring every quote against two themes and ranking them by how close the contest is: mechanical, and infeasible by hand, somewhere past a few hundred quotes. Deciding any single one: a theoretical judgment that a model should not be making quietly.

Those two halves could not sit in the same room before. Batch tools returned a spreadsheet with the ambiguity already resolved and nobody's name on it. Click-driven automation could not read the analysis at all.

Now they interleave. The agent surfaces "I don't know whether ChatGPT is actually accurate", filed under policy uncertainty, and points out that the evidence is really about trust. I look. I agree or I don't. I move it. That decision immediately changes what the agent compares against on its next pass, so it gets better at finding my arguments because I keep making them.

The workspace also keeps score, running in the right direction. It tracks how often my decisions matched the agent's ranking, so I learn how much to trust it based on my own record rather than its confidence.

How I built it

Nuxt and Vue, entirely client-side, are deployed as static assets on Cloudflare Workers.

One store holds the analysis, and every change goes through it, whether a human or an agent asks. That single write path is what makes attribution reliable instead of aspirational.

Thirteen tools register with the browser's model context when a WebMCP host is present, and the page behaves normally when no host is present. Every tool is a research operation rather than a UI action: nothing in the set clicks a card or drags a rectangle. Read-only tools are marked as such and flagged as carrying participant-authored text, so a host knows not to treat an interview quote as an instruction.

The scoring is deliberately conservative. A quote has to be clearly relevant to both themes before closeness counts for anything, otherwise two weak scores look like a hard case when they are just noise. Theme definitions are built only from examples that a human confirmed, so the agent's proposals never quietly redefine the thing it is measuring against.

24 unit tests, plus a browser test that registers the full tool set and runs an agent through a complete session from a cold start. I run that against the deployed site too, so the live demo is verified the same way as the local build.

Challenges

Telling two kinds of hard cases apart took the most thought. A quote that genuinely sits between two themes needs a human. A quote that makes two separate points belongs to both, which is not ambiguity, just dual membership. Measure only how close the scores are, and the two look identical, but they need different handling and a different place in the queue.

What I learned

My first version had eight tools and a hole I could not see. One of them needed two theme IDs, and nothing in the set returned the list of themes. Reading the docs, it looks complete, because I already know my own theme names. An agent arriving cold was stuck at the first step. Whether an agent can orient itself from nothing turned out to be the only test that mattered.

The other lesson is about where a rule has to live. Writing "the agent should not decide" into a prompt is a request, and a model that wants to be helpful will eventually find a way around it. Not building the tool is a different kind of thing, and it's the version I'd rely on.

What's next

First the dull necessary things. Decisions don't survive a refresh, so the agreement score can't build up over the course of a study. Scoring could be sharper than word overlap. Several researchers should be able to work the same corpus. And it should export into NVivo and ATLAS.ti, where this work actually lives.

The bigger thing sits further up the pipeline. Interviews start as audio, and transcription and speaker diarization are both commodity at this point, so putting them on the front of this is mostly integration work. Then you hand it a folder of recordings and get back a live analysis with the contested passages already queued, rather than transcripts you still have to sit down and read.

That's the service I'd want to use, and I'd build it under the same restriction. Once most of a pipeline runs unattended, the record of who decided what is doing most of the work of making the output defensible at all.

Built With

Share this project:

Updates

Submission history