Inspiration AI scribes that listen to a consultation and write the clinical note are spreading fast. Their most dangerous mistakes are the quiet ones: the sentence reads fluently but is wrong about who, when or whether.

The son says "My father has asthma", and the note says "Patient has a history of asthma." The patient says "Coughing for a week… no, actually ten days", and the note keeps "one week." The doctor says "If the fever continues tomorrow, come back for an X-ray", and the note says "X-ray ordered." Vietnamese makes the first mistake more likely. Speakers often drop the subject ("Sốt hai hôm rồi", "fever for two days"), and kinship words double as pronouns (con, cháu, em can mean "I" or "the child"). Standard metrics do not see this: when we swapped the patient and the relative in 400 notes, ROUGE-1 stayed at 1.0000.

What it does MediTrace Sentinel turns a Vietnamese consultation, typed or recorded, into a draft clinical record that a doctor reviews line by line.

Every line links back to the exact turn of the conversation it came from, with a verbatim quote. Each piece of information is placed in one of three spots: the draft body; a "Needs confirmation" section, with the reason it was moved there; rejected, kept in the trace but out of the draft. It separates who said it from who it is about. It marks old values as superseded when a speaker corrects themselves, and keeps plans and hypotheticals apart from facts. It suggests up to 3 clarifying questions for the doctor. Audio in: local speech recognition (PhoWhisper). Optional speaker separation (pyannote); the doctor then assigns a role to each speaker. It does not diagnose or prescribe, and the draft is not an official medical record.

How we built it conversation → 1. extract statements (LLM) → 2. link people → 3. update state → 4. check evidence → 5. route → draft Extraction is the only step that uses a model. We use Qwen3-4B fine-tuned with QLoRA, running locally, with JSON output constrained by lm-format-enforcer. Each statement records content, subject, speaker, certainty, negation, fact/plan/hypothetical, time, medication fields, evidence turn numbers and a quote. Linking, state updates, evidence checks and routing are readable rules. The evidence checks are six tests against the cited turn: the quote exists; the content overlaps; person words match the subject; certainty was not raised; the time appears in the turn; the dose appears in the turn. Data. We wrote a rule-based generator for 5,000 synthetic Vietnamese consultations from 100 case templates, split by template. On top of those we built 300 paired challenge dialogues: subject swap, self-correction, speech-recognition noise and regional dialect. No real patient data was used. Scorer. It counts errors per statement by type: wrong person, wrong certainty, stale value, unsupported, medication, time. It also reports F₁ and slot error rate, with confidence intervals from a bootstrap over case templates. Web app. React + Express front end with per-line keep / edit / drop, invite-code accounts and SQLite storage. An optional external-LLM switch exists for side tasks only; it is off by default and never used for the draft. Hardware. Training ran on Kaggle T4 GPUs and one RTX 3060. More than 1,200 automated tests. Challenges we ran into Our own measurements fooled us several times. One results table was scored against the answer key of a different data generation. Two challenge sets were run under different settings than the others. The answer key for a dialect phrase ("nóng hầm hập") said "fever, certain" although our own rules say "ask again". Each time we added an automatic check so the mistake cannot repeat silently. Timestamps. The speech model gives only word start times, so silence between turns was invisible and whole consultations collapsed into one line. We now cut lines from voice-activity spans; on synthetic audio with 0.4–1.0 s pauses, mixed-speaker lines fell from 100 % to 15 %. Honest comparison. A fine-tuned model that writes the note directly scored higher on our attribution metric than our pipeline (0.939 vs 0.748 on the subject-swap set). Part of the gap comes from that model copying the generator's reference notes word for word in 30–63 % of cases, and part is real. We report both. Accomplishments that we're proud of Measured results on the challenge sets:

Measure Before After F₁, subject-swap set (plain extraction → with checks) 66.1 % 70.8 % Slot error rate, same set 45.4 % 36.9 % Superseded values left in the draft, correction set 11.48 % 0.80 % Evidence checks blocked 0 of 30,136 correct statements.

A scorer that sees the error ROUGE misses. On the patient-relative swap, ROUGE-1 did not change at all, while our attribution score dropped by 0.1381.

Everything runs locally on a single consumer GPU, so consultation audio never has to leave the clinic.

What we learned Talking to clinicians changed our assumptions. A head of clinical nutrition told us a record takes her 15–30 minutes. The slow part is finding the core problem and counselling the patient, not typing. She also said she did not need automatic flags for missing or contradictory information. A medical student told us records follow the reason for the visit and lab tables matter a lot. At teaching hospitals, students write the record and supervisors re-check it with the patient. So "saves doctors time" is a claim we will not make until we measure it. A strong commercial model assigned subjects more accurately than our 4B model. Our case is traceability and local deployment, not raw accuracy. Report negative results next to positive ones. Otherwise the numbers mislead you first. What's next Evaluate on dialogues written by real people and on real clinic audio, with ethics approval. Two independent human raters, to check our automatic scorer. Structure the draft by the reason for the visit, instead of fixed sections. Read lab result tables and link abnormal values to the conversation. A time-and-error study with clinicians before claiming any workflow benefit.

Built With

Share this project:

Updates

Submission history