Inspiration

I've lost hackathons before because the winning team had more depth than me — more real testing, more honest evidence, fewer unproven claims. So this time I picked a problem where "it works" and "it's actually safe" are different questions, and built toward the second one.

Voice agents for clinic intake answer the same handful of questions all day: how long to fast, when to arrive, what's covered, what to bring. A standard RAG agent grounds its input — it retrieves the right document — but nothing stops the model from confidently repeating a stale or contradicted fact once it's generated. A patient who fasts for twelve hours instead of eight because an old prep sheet was still sitting in the index isn't a UX bug. It's a wasted appointment.

What it does

Hesitate checks every sentence the agent is about to speak against the clinic's actual current policy, before it reaches text-to-speech. A typed extractor pulls out check-worthy claims — fasting duration, arrival time, insurance coverage, cost, required documents. Moss retrieves the current authoritative record for that claim in sub-10ms. A resolver compares values directly, not by similarity score. If they match, the sentence is spoken as-is. If they contradict, the sentence is suppressed and replaced with a correction built only from the real policy record — never from the model's own wrong number. If nothing relevant is found, the agent says so instead of guessing.

How I built it

LiveKit handles the real-time voice connection. Deepgram transcribes speech. Groq's gpt-oss-20b generates the reply. Before that reply reaches ElevenLabs for speech, it passes through the verification gate — the actual novel part of this project. Moss holds the clinic's policy corpus and is queried per-sentence; this only works because Moss resolves in-process, fast enough to sit inside a live conversational turn without breaking the pipeline. A hosted vector database's ~400ms round trip would make this architecture impossible.

Challenges I ran into

The real one: retrieval similarity does not mean a claim is supported. A passage about fasting scores high whether it says 8 hours or 12 — ranking by similarity would approve the hallucination, because the retrieved chunk is topically perfect and factually wrong. The gate has to compare actual values, not scores.

I also found real bugs while testing live, and I'm listing them here instead of hiding them: a markdown-formatted LLM reply (**12 hours**) broke every extraction regex because none of them treated * as a word boundary — a claim just silently vanished instead of being caught. And GuardedTTS.stream() — my own quota guard for ElevenLabs — was unimplemented for the actual streaming code path AgentSession uses in production, so a live call would connect and simply produce no audio, no error, nothing. Both were caught by actually running the full pipeline against real accounts, not by unit tests against fixtures.

What I learned

That "it passed my tests" and "it works on a real account, live, end to end" are different claims, and the gap between them is where the actual bugs live. Every bug that mattered in this project was found by running the real thing against real services, not by reasoning about the code.

What's next

Broader claim-family coverage (right now it's fasting, arrival, coverage, documents, cost, appointment windows — not every kind of claim a model might assert), and testing the gate against a real clinic coordinator rather than my own judgment of what a coordinator needs.

Built With

Share this project:

Updates

Submission history