About CareLine
Inspiration
Hospitals are noisy — not just with machines, but with phones. Every floor nurse knows the rhythm: a patient's daughter calls at 9 AM, her brother calls at 9:15, the neighbor calls at 9:30. Each time, the same question — "How are they doing?" — and each time, the same dilemma: HIPAA says I can talk to you, but who are you? Are you on the list? And if you are, how much am I allowed to say?
We built CareLine because families deserve better than a busy signal, and nurses deserve to spend their time on care, not call-answering. There had to be a way to let a family member call any time — middle of the night, during shift change, during a code — and get a real, grounded, permission-aware answer. Not a recording. Not "please hold." A conversation.
The dream: a voice agent that knows exactly who you are, exactly what you're allowed to know, and speaks about your loved one with the warmth and accuracy of a good nurse — without ever guessing, embellishing, or leaking a single detail it shouldn't.
What We Learned
Building a voice agent for healthcare is a different muscle than building a chatbot. A lot of lessons came the hard way:
Latency is everything. On a phone call, a 5-second wait feels like an eternity. We learned to measure every millisecond — the STT round-trip, the LLM inference time, the TTS synthesis — and optimize the weakest link first. Switching from OpenRouter to Groq cut our LLM inference from ~4 seconds to under 1 second. Going from 3 retries with exponential backoff to a single fast fallback shaved off another potential 5-10 seconds on failure.
Prompt engineering for voice is different. Text-to-speech doesn't handle markdown, bullet points, or abbreviations. A response that reads fine on screen ("Patient is S/P ORIF L hip, VS WNL") sounds terrifying on a phone call. We rewrote every prompt to output short, warm, 2-3 sentence responses in plain English — no jargon, no symbols, no room numbers to the wrong caller.
Guardrails are hard to test in the moment. It's one thing to write a permission system on paper; it's another to watch someone rephrase "How much pain is she in?" six different ways until the model slips. We built a CLI test harness (
test_brain_cli.py) that let us hammer the guardrails from every angle before we ever answered a real call — and we caught multiple failures there that would have been expensive to debug over Twilio.State management on a phone call is fragile. A caller might hang up mid-sentence, stay silent for 20 seconds, or sneeze at the wrong moment and trigger Twilio's speech detection. Our
CALL_STATEdictionary perCallSidbecame a mini state machine that had to handle silence retries, session timeouts, and multi-turn coherence gracefully.
How We Built It
The system is a layered architecture with four main components:
1. The Brain (brain.py + data.py) — An in-memory patient database with encounter histories, medication lists, and treatment plans. Each patient record is annotated with auto-generated care summaries via a dedicated LLM call. The HospitalBrain class is the single source of truth for patient lookup, contact verification, and permission gating.
2. The Voice Server (main.py) — A FastAPI server with four endpoints: /voice handles inbound calls, /voice-retry re-prompts after silence, /handle-speech processes transcribed speech and returns TwiML responses, and /refresh-summary lets doctors regenerate care summaries on demand. The per-call state machine maintains conversation history using a dictionary keyed by Twilio's CallSid.
3. The LLM Layer (llm_client.py) — An async HTTP client that supports three backends (Ollama, OpenRouter, Groq) with automatic fallback chaining. Each backend gets up to 3 retries with exponential backoff before failing over to the next. The system prompt template enforces four non-negotiable rules: only use authorized facts, never give medical advice, decline out-of-scope questions, and keep responses short and warm.
4. The Dashboard (streamlit_app.py) — A real-time clinician UI that displays all patients, their current status, and auto-generated care summaries. Nurses can update patient status with a single click, which triggers immediate care summary regeneration via the LLM — and the next call automatically reflects the update.
5. The Telephony Layer — Twilio Voice handles all telephony: inbound PSTN calls, speech-to-text via <Gather input="speech">, and text-to-speech via <Say voice="Polly.Joanna">. The architecture is a simple request-response cycle: caller speaks → Twilio transcribes → POST to FastAPI → LLM generates reply → TwiML instructs Twilio to speak the response and listen for the next utterance.
Challenges We Faced
Permission gating without leaking. The hardest problem wasn't building the permission system — it was ensuring the LLM respected it. A status_only caller should get nothing but "Your father is stable." But an LLM that knows the room number, diagnosis, and medication list is under enormous pressure to answer follow-ups naturally. We solved this by feeding the model only the information it's allowed to share — the context for status_only callers literally excludes room number, diagnosis, and medications — rather than trusting the model to filter itself.
Latency in the voice loop. Our first end-to-end test measured 8-10 seconds per turn — completely unusable. The optimization journey was:
- Replacing Ollama (local) with OpenRouter (cloud) wasn't the magic bullet we expected — routing overhead added latency
- Moving to Groq (specialized inference hardware) dropped LLM latency from ~3-5s to ~500ms
- Removing exponential backoff on retries prevented silent 10-second delays on transient failures
- Reusing HTTP client connections eliminated TLS handshake overhead on every call
Realistic demo data. Healthcare demos live or die on their data. A fake patient that sounds fake breaks the spell. We wrote five patients with real-feeling admission notes, post-op checks, care plans, and medication regimens. Elizabeth Chen (74, hip fracture) and David Ruiz (8, asthma flare) felt like real people — and that made the demo land emotionally.
Twilio trial account gotchas. There's a ritual of setup — verifying phone numbers in the Twilio console (trial accounts can only call approved numbers), buying a number with Voice capability, pointing ngrok at the right port — that has to go perfectly or the demo just... doesn't work. Half the setup time went into this plumbing, and we built a detailed runbook (see RUN.md) to avoid fumbling under judge pressure.
TLDR : AI-powered voice assistant that answers family phone calls to hospitals, giving permission-aware patient updates without adding to nurses' workload. Instead of making nurses repeat the same status updates all day, callers can have a natural conversation with an AI that first verifies who they are and only shares information they are authorized to receive. The system combines a FastAPI voice server, Twilio telephony, an LLM with strict guardrails, and a clinician dashboard. To prevent privacy leaks, the AI is only given the information each caller is allowed to know rather than relying on the model to censor itself. We also optimized the voice pipeline for real-time conversations, reducing response latency from roughly 8–10 seconds to about 2.5–4.5 seconds by switching inference providers, eliminating unnecessary retries, and reusing network connections. The result is a fast, natural phone experience that protects patient privacy, keeps families informed, and lets healthcare staff spend more time caring for patients instead of answering repetitive phone calls.
Log in or sign up for Devpost to join the conversation.