Dental Triage Voice Agent

Built at Healthcare Hack NYC + Twilio Searchlight.

Inspiration

Dental offices lose patients to voicemail. A patient with a swollen jaw or a knocked-out tooth calls after hours, gets a machine, and either goes to the ER unnecessarily or — worse — waits until morning with a real emergency. At the same time, front-desk staff spend most of a call on the same handful of things: confirming insurance, finding a slot, and figuring out how urgent the problem actually is. We wanted a voice agent that could handle that triage-and-book flow end-to-end in a single phone call, with hardcoded safety guardrails so a hallucinating LLM could never be the thing standing between a real emergency and 911.

What it does

A caller dials a Twilio number and talks to a conversational agent (ElevenLabs STT/TTS + Claude as the turn-by-turn decision-maker) that sorts the call into one of five buckets: routine, urgent non-emergency, true dental emergency, non-dental red flag, or an insurance/cost question. If the caller's number matches a patient record, the agent personalizes the greeting with their last visit and insurance plan. For routine and urgent calls, it checks live appointment slots and insurance coverage and books directly into the mock EHR (Supabase). If the caller describes a non-dental red flag — chest pain, trouble breathing, stroke signs — a deterministic regex/keyword detector (not the LLM) interrupts the conversation and routes them to 911/ER, no booking attempt. If it's a true dental emergency (avulsed tooth, uncontrolled bleeding, jaw trauma, swelling with fever), the agent books an emergency slot and fires a live transfer or urgent SMS before the call ends — "someone will call you back" is never an acceptable outcome. Every call is logged with transcript, classification, and tool calls, viewable afterward, and callers can also open a web link mid-call to type or upload a photo of an insurance card if it's easier than saying it out loud.

How we built it

Four of us split the build into parallel branches after one person locked down the Twilio webhook contract solo: telephony (Twilio number, RequestValidator signature validation), data/knowledge (Supabase schema for patients/slots/insurance, structured logger, post-call SMS), triage conversation (the Claude prompt, personalization, barge-in), and safety escalation (the regex red-flag detector, built and isolation-tested before it was ever wired into the live call flow). We ported patterns — not code — from an existing scheduler reference repo: the agent-turn webhook shape, the live-tool-call pattern, and the logger/error-recovery structure. Everything runs behind an ngrok tunnel, backed by a mock Postgres EHR seeded with 5 patients, 10 slots, and one insurance plan — no real patient data, no real clinical claims, by design.

Challenges we ran into

The biggest surprise was that ElevenLabs' native Twilio integration takes over the number directly (voice_url pointed at ElevenLabs' own endpoint), which meant our own /voice webhook — including the <Connect><Stream> handoff we'd built — never actually saw live call traffic; a real call died instantly with Twilio error 31921 because our stream was speaking the wrong websocket protocol. We had to accept that ElevenLabs owns the voice leg and re-scope our webhook to the tool-call surface it does hit. The red-flag detector also had two nasty near-misses that only showed up under real phrasing: it flagged "swollen but no fever" as an emergency because it matched "fever" without checking for the preceding negation, and it missed real stroke-sign phrasing ("face drooping," "speech is slurred") because the regex only matched one rigid word order. Both got caught and fixed before the demo, which is exactly why we isolation-tested the detector before wiring it into the call flow instead of trusting it end-to-end from the start.

Accomplishments that we're proud of

Keeping the safety escalation completely deterministic — no path where the LLM can talk its way out of flagging a real emergency — while still shipping a full conversational booking flow in a 3-4 hour window. We're also proud of catching both red-flag detector bugs (negation handling, word-order/inflection) through isolation testing before they could have silently misfired on stage, and of the fact that every tool call — booking, insurance check, escalation — fires before the call ever ends, so no caller is left with just a promise of a callback.

What we learned

Reference integrations (like ElevenLabs' native Twilio import) can silently reroute traffic around the exact webhook you built and tested — verifying signature validation against a simulated request isn't the same as verifying it against a real call, and we didn't catch that gap until a live test call failed. We also learned that safety-critical regex logic needs adversarial phrasing tests (negations, reordered clauses, alternate inflections), not just the happy-path phrases you'd write from a script, before it's trustworthy enough to gate a real emergency response.

What's next for Dental Triage Voice Agent

Resolve the Invariant 8 open question properly — either formally accept that ElevenLabs' side owns Twilio signature validation for the live voice path, or find a supported way to route real calls through our own validated webhook. Beyond that: expand past the single-concurrent-caller MVP scale (swap the in-memory web-input session store for Supabase/Redis), add OCR for uploaded insurance card images instead of filename-only capture, and broaden the red-flag test suite with more adversarial real-world phrasing before trusting it with a wider patient population.

Built With

Share this project:

Updates

Submission history