How we built it
The phone is the only input device, and it works offline. The chat is a WhatsApp-style web app the midwife installs on her home screen.
A service worker keeps the app on the phone, so it opens with no network. Every photo is first encrypted on the phone (AES-GCM, with a key generated on the phone that can't be exported) and kept in local storage. When the network returns, the phone uploads the photos oldest first. It deletes each one only after the server confirms it has stored it. Closing the app, losing the network or the server going down loses nothing.
The processing server is a FastAPI backend with an encrypted record store and a strict state machine: captured → pending AI → processed → to review → validated → linked → saved → synced. Failure states always have a way out, an illegal jump is rejected, and every transition is logged.
Reading the page. The system first recognises the page type, then extracts every field with a value, a status and a confidence. Statuses are explicit (KNOWN, UNKNOWN, NOT_PROVIDED, ILLEGIBLE, NOT_APPLICABLE, NEEDS_REVIEW), so a blank, a dash and an unreadable word are never confused.
The organisers' specimen layout: templates built from the specimen PDF, page alignment, ink detection and checkbox scoring. The real pink booklet: a template-free reader. It finds printed labels and table lines in tilted, curved phone photos and builds each cell. It then removes the printed lines and print by ink colour, and joins writing that runs diagonally across rows. Field-aware reading: instead of taking the OCR's first guess, we score every plausible value for a field (a blood pressure in cmHg, a real date, "16SA+3j", a clinical word) with the recogniser's own character probabilities, plus weak priors from the synthetic data. It's the same model asked a better question. On the real photos this raised correct fields from 74 to 85 out of 127. A fine-tuned vision model: Qwen3-VL-8B, trained with QLoRA on a free Colab T4 to read one field at a time. We trained it first on specimen field crops, then on 6,000 synthetic crops we generated in the real midwife's formats ("11/7" pressures, "NF", "Reçu", the French "1"). It plugs into the same pipeline as a GPU reader on our own machine, not a cloud service. Privacy by design: names, ID numbers, phone numbers and addresses are never read or stored. They're painted out before any model sees the page, and patients are linked only by the code the midwife writes on the booklet.
Evaluation:
ground truth for 5,640 specimen fields, extracted from the PDF itself; held-out patients; simulated phone photos (blur, shadow, perspective, JPEG); 127 fields of the real photos transcribed by hand; a crop editor we built to inspect and correct how each word is cut out. Challenges we ran into Real handwriting is not the specimen. The specimen pages use handwriting fonts. The real booklet has fast French cursive, a "1" the OCR reads as "Λ", blood pressures in cmHg, words spilling over thin rows, and "RAS" written diagonally across whole sections. A method at 98.9% on clean specimen pages scored 67% on real photos. A bigger model wasn't automatically better. We tried a larger handwriting OCR (TrOCR), and it did worse on our crops than the small model given field context. Hardware. Our development laptop had 6 GB of RAM and no GPU. We hit memory crashes in the OCR runtime and models that didn't fit. Colab's GPU quota ran out before we could benchmark the second version of our vision model on real photos. Privacy versus training. The brief forbids sending real patient data to a third party, so we couldn't train on the real handwriting in the cloud. Instead we built a generator that reproduces the midwife's formats with synthetic values. When we tested on real crops, identifier areas were masked first. "Offline" had to mean the phone. Our first version simulated offline mode on the server. We rebuilt it so the photos live, encrypted, on the phone itself. That ran into a browser rule: offline apps and encryption only work on secure (HTTPS) pages. Keeping the numbers honest. The vision model scores 94–95% on synthetic and specimen crops but about 50% on real cursive (CPU OCR: 46%). We report both, and we don't claim the synthetic number for the real booklet. What we learned Start from the schema, not the OCR. Knowing what each field can contain was worth more than a bigger model. Asking beats guessing. The metric that matters is the silent error: a value accepted that is wrong. On real photos about half the fields go to the midwife for a quick check, and only 1 in 127 was accepted wrongly. A wrong blood pressure entered silently is worse than a question. Missing is information. Treating "unknown", "not provided", "illegible" and "not applicable" as first-class states makes the uncertainty visible instead of hiding it. Measure on real data early. Synthetic benchmarks flattered every method we tried; 127 hand-transcribed real fields told us the truth. Offline-first is about where the data lives, not a toggle in the interface. Privacy rules shape the whole machine-learning pipeline: what you can train on, where it runs, and what you're allowed to measure. What's next Benchmark the fine-tuned model on the real booklets on a GPU laptop, and train it on real handwriting inside the health system, where that data may stay. Arabic handwriting, plus the real booklet's delivery and postpartum pages. A real WhatsApp Business connection, and a native app with the key in the phone's secure keystore.
Log in or sign up for Devpost to join the conversation.