Inspiration

In most Indian outpatient clinics, a doctor sees a patient, talks for two minutes, and scribbles a prescription in handwriting only the pharmacist next door can decode. Voice-to-text tools exist, but transcription alone doesn't solve the real danger: two drug names that sound almost identical can point to completely different organ systems.

The seed for MedScribe was a real look-alike/sound-alike (LASA) pair: Amlodipine (a blood-pressure medication) and Amiloride (a diuretic for fluid retention). Said quickly, in a Hinglish accent, over a phone mic, these two blur together, and a pipeline that just writes down what it hears will hand a patient the wrong drug with full confidence. We wanted to build the layer that catches that moment: not smarter transcription, but a system that treats every extracted medicine as unverified until it's checked against a real formulary and flagged if it's ambiguous or dangerous.

What it does

MedScribe listens to a live doctor–patient consultation and turns it into a safe, structured prescription draft in real time.

  • Extracts medicines, dosages, frequency, and instructions from the transcript as the consult happens.
  • Resolves each spoken drug name against a seeded formulary (220 brands, 120 salts, 20 conditions) using deterministic phonetic and fuzzy matching, never the LLM's own judgment.
  • Flags anything ambiguous, low-confidence, or clinically contradictory (like a LASA collision) directly on the draft, blocking approval until a doctor resolves it.
  • Locks any field the doctor manually edits so later re-extraction passes can't quietly overwrite a human decision.
  • Delivers the doctor-approved prescription and dosage reminders straight to the patient over WhatsApp, no app to install.

Nothing reaches a patient without a doctor's explicit sign-off, and nothing about identity or safety is ever decided by the language model.

How we built it

The pipeline is a strict separation of concerns:

  1. Extraction: an LLM (Mantle/Bedrock's openai.gpt-oss-120b, chosen after testing five models for zero rule violations on Hinglish transcripts) turns raw transcript text into structured medicine mentions, returning spoken_name completely verbatim: no normalization, no spelling correction, since correcting it early would destroy the exact signal the matcher needs downstream.
  2. Resolution: resolver.py scores each spoken name against the formulary using phonetic bucketing plus fuzzy matching, and only assigns identity when the top score clears a minimum bar and beats the runner-up by a minimum margin:

$$ \text{score}(m_1) \ge \text{MIN_SCORE} \quad \text{and} \quad \text{score}(m_1) - \text{score}(m_2) \ge \text{MIN_MARGIN} $$

If either fails, the item comes back as RESOLVE, not CONFIRM: visible and blocking, never guessed.

  1. Validation & LASA detection: a rules-based validator checks the resolved item's condition mapping against the patient's chart, and a post-validation pass flags look-alike/sound-alike collisions as CONTRADICTS.
  2. The merge layer (pipeline.py): a pure function, process(transcript_text, prior_state) -> dict, that unions new extraction results with the doctor's existing draft, respects locked fields, tombstones deleted items, and fails safe to the prior state if the LLM call errors. Persistence lives entirely outside it, with the caller owning storage.
  3. The doctor dashboard: every RESOLVE and CONTRADICTS flag sits directly in the draft, blocking the approve button until a human clears it.
  4. Delivery: once approved, the confirmed prescription and reminders go out via the Meta WhatsApp Cloud API.

Challenges we ran into

  • Threshold tuning under real clinical stakes. Every number we chose (MIN_SCORE, MIN_MARGIN, AUTO_SCORE) trades off over-flagging (doctor fatigue, ignored warnings) against under-flagging (a wrong drug slipping through as confirmed). At MIN_SCORE = 62, paracetamol confidently resolved to Stamlo 5mg, a blood-pressure drug, as an auto-confirmed match. We raised it to 72 and pinned the value with a contract test so no future change silently drifts it back toward the dangerous side.
  • Getting the LASA flag to actually fire. Our seed data initially made this impossible: a dead LASA_PAIRS list, a missing condition mapping for Amiloride, and Amlodipine landing in a different phonetic bucket meant the two drugs never even got compared. Fixing the demo's centerpiece meant adding a condition, remapping a salt, and rebuilding the seed data.
  • Infrastructure churn mid-build. A Bedrock account gate blocked every model in every region, and Twilio's new WhatsApp trial flow allowed only canned templates. Both forced pivots, to a Mantle-based OpenAI-compatible endpoint, and to the Meta Cloud API directly, without slipping the timeline.
  • Keeping the model boxed in. Getting an LLM to reliably never touch brand identity, never correct a spoken name, and never silently drop an item, even when that would make its own output look cleaner, took real prompt discipline, and ultimately a code-level guarantee (extractor.normalise() strips any id the model tries to emit) rather than trusting instructions alone.
  • State merging across a live consultation. A transcript is re-extracted every few seconds as the doctor keeps talking. Naively overwriting the draft each time would erase edits mid-consult. Every medicine carries a stable med_key (its resolved brand ID, or the normalized spoken name before resolution) so edits reattach correctly as the list grows.

Accomplishments that we're proud of

  • A working end-to-end demo of the exact failure case that inspired the project: Amlodipine and Amiloride, spoken in Hinglish, correctly caught as a contradiction instead of silently auto-confirmed.
  • A hard architectural guarantee, not just a prompt instruction, that the LLM can never assign clinical identity. Only deterministic code touches brand_id.
  • 156 offline tests, including a contract test that pins the resolver's exact threshold behavior so a future "helpful" tweak can't quietly reopen the paracetamol/Stamlo bug.
  • A real WhatsApp integration verified end to end (inbound linking, outbound delivery, read receipts) on Meta's Cloud API after Twilio's flow turned out to be a dead end.

What we learned

  • LLMs are excellent at extraction, terrible at authority. The moment we let the model "helpfully" normalize a drug name or resolve it to a brand, it made confident, silent, wrong guesses. The fix was architectural: the LLM never resolves identity, period.
  • Thresholds are a clinical decision, not a tuning knob. One number was the difference between "the system said it was fine" and "the system asked a human."
  • Silence is the real failure mode. A flagged, blocking item on screen is safer than a missing one; a doctor notices something wrong on screen, but never notices something that was never shown.
  • A safety feature is only as good as the data underneath it. The LASA check was correct code sitting on broken seed data; the bug wasn't in the logic, it was in what the logic had to work with.

What's next for MedScribe

  • Move the WhatsApp integration off its 24-hour temporary access token onto a permanent system-user token, and get the medicine_reminder template out of PENDING approval.
  • Build the Step Functions pipeline to run extraction continuously during a live consult instead of on manual triggers.
  • Add the reminders flow (scheduled WhatsApp nudges tied to dosage frequency) and the patient-facing dashboard with prescription history and chat.
  • Expand the formulary beyond the 220-brand seed set and start closing the verified_by gap so Tier 0 data actually earns its "verified" label.
  • Revisit condition matching, currently naive substring matching, which over-matches short strings like "gas" or "BP": fine for a demo, not for production.

Built With

  • amazon-cognito
  • amazon-web-services
  • api-gateway
  • aws-lambda
  • aws-sam
  • aws-step-functions
  • bedrock
  • dynamodb
  • gemini
  • moto
  • openai-gpt-oss-120b
  • python
  • sarvam-ai
  • whatsapp-cloud-api
Share this project:

Updates

Submission history