Inspiration
Doctors spend 6 hours in the electronic health record for every 8 hours of patient time, and nearly half of that work happens after hours (PubMed).
AI scribes were supposed to fix this, but they leave three problems unsolved:
- The note still has to be proofread. 94.7% of AI notes are free of significant errors, but the rest can cause serious harm, so doctors read every line (PubMed). In the first randomized trial of AI scribes, the best tool saved just 41 seconds per note (NEJM AI).
- Accurate notes still miss things. Across seven commercial scribes, 83.8% of errors were omissions, and none produced an error-free summary (PMC). You can't transcribe a question nobody asked.
- Everything after the note is manual. Between 6.8% and 62% of outpatient lab results go without documented follow-up (J Gen Intern Med), and patients forget 40–80% of what they're told almost immediately (Kessels, 2003).
Carryover isn't another scribe. It sits on top of the scribe and checks its work.
What it does
- Click-to-source: every note sentence links to the exact transcript lines and audio moment it came from.
- Confidence scored in code, not by the AI: sentences lose confidence for weak citations, failed audits, mismatched numbers, or unverified medication details.
- Omission checks: before sign-off, the visit is checked against a visit-type checklist and last visit's unresolved items. Missed items can be completed, dismissed with a reason, or carried forward. Required items block sign-off unless overridden with a logged reason.
- Live copilot: a rare, quiet "Consider asking..." card during the visit, at most three per visit.
- Follow-through on sign-off: office tasks, carry-forward items, a plain-language patient summary in English or Spanish with a readability score, and a PDF report.
Spanish matters: adverse events among patients with limited English are more likely to cause physical harm (49% vs. 30%) (PubMed). All outputs come only from the signed note, so errors the doctor fixed can't leak through.
How we built it
- Transcription: ElevenLabs Scribe v2 Medical transcribes and separates speakers into numbered, citable transcript lines.
- Drafting and auditing: Gemini drafts a SOAP note in which every sentence must cite its source lines. A separate Gemini call audits the draft, so no model grades its own work.
- Deterministic scoring: each problem's score is
$$ \text{Score} = 0.5 \times \min(s_i) + 0.5 \times \operatorname{mean}(s_i) $$
One weak sentence lowers the score without sinking an otherwise solid section.
- Live copilot: ElevenLabs Scribe v2 Realtime over a raw WebSocket. Gemini checks roughly every 20 seconds, with a 45-second cooldown, a limit of 3 suggestions per visit, and a 0.8 confidence floor.
- Storage: Postgres + Drizzle on Tiger Cloud, with vitals and audit events as TimescaleDB hypertables. PGlite handles local development.
- Memory: a pre-visit brief from Backboard, falling back to Gemini, then to a basic chart summary.
- Frontend: Next.js 16 and React 19.
Challenges we ran into
- Real conversations don't pause cleanly. Realtime transcription waits for 1.5 seconds of silence, which interrupting humans rarely give it. We commit segments manually instead.
- Teaching the copilot to stay quiet. Noticing gaps was easy; tuning it down to zero or one useful interruption per visit was the hard part.
- Spoken numbers. "One thirty-eight over eighty-eight" has to match
138/88, so we wrote our own parser. - Free-tier limits across three AI services. Fallbacks and fixture data keep the demo running if any one of them fails.
Accomplishments we're proud of
- Click any sentence, hear the moment that supports it. No guessing where a statement came from.
- We catch what's missing, not just what's wrong. Omissions are the most common scribe error, and Carryover catches a missed allergy check before the prescription goes out.
- Failure is visible and recoverable. Rate limits and outages fall back to something useful instead of breaking silently.
What we learned
- Confidence means nothing unless you can explain it. That's why ours lives in code: same inputs, same reasoning.
- AI notes aren't the problem. One study even found they outperformed handwritten notes (Research Square). The problem is that doctors can't tell which parts to trust.
- Restraint matters as much as intelligence. Getting AI to notice things is easy; getting it to know when not to speak is hard.
What's next for Carryover
- Clinical validation: test the checklists with real clinicians, measuring review time, sentences flagged, gaps caught, and carry-forward items closed.
- Compliance: a BAA with ElevenLabs (Scribe v2 Medical already supports Zero Retention Mode) before any real patient data.
- More visit types beyond our current three.
AI scribes are already in roughly 30% of physician practices, and adoption is outpacing validation (npj Digital Medicine).
A medical AI system shouldn't just produce a note. It should make sure the note is supported, complete, and actually followed through.
Built With
- backboard
- bcrypt
- drizzle-orm
- elevenlabs
- gemini-api
- google-gemini
- jwt
- next.js
- node.js
- nodemailer
- pglite
- postgresql
- react
- react-pdf
- recharts
- resend
- tailwindcss
- tiger-data
- timescaledb
- turbopack
- typescript
- vitest
- websocket
- zod
Log in or sign up for Devpost to join the conversation.