Inspiration
In many low-resource clinics, midwives record antenatal care, delivery and postpartum follow-up in a multi-page paper registry, filled in by hand in French. Once it is written down, the information stays on paper: it does not follow the woman from one visit to the next, and it never reaches the health system.
The usual answer is "give them an app" or "retype everything". Both change the midwife's workflow, and both fail when the network drops. We wanted the opposite: keep the paper registry exactly as it is, and add the digital layer on top. The midwife already has a phone and WhatsApp. If she can take a photo, she can feed a structured, longitudinal record, even when she is offline.
What it does
The midwife photographs a registry page and sends it on WhatsApp. The agent then:
- Reads the page into a structured record. Every field has a value, an explicit status (KNOWN, UNKNOWN, NOT_PROVIDED, ILLEGIBLE, NOT_APPLICABLE, NEEDS_REVIEW) and a confidence score. There is never a vague "N/A".
- Never hides its doubts. Each uncertain field becomes a question, one at a time, with a thumbnail of the exact cell. She confirms, picks between two readings, types the correct value, or marks the cell as empty. A summary follows: Confirm / Edit / Retake photo.
- Works offline-first. With no network, WhatsApp keeps the photos on the phone. When the connection returns, they all arrive at once, out of order and sometimes duplicated. The bot stores each one encrypted before acknowledging it, de-duplicates by message id, sorts by capture time, sends a single acknowledgement, groups the pages by registry and processes them in the background. If the AI or the database goes down, nothing is lost: pages wait, are retried automatically, and the conversation goes on.
- Links visits to the right woman with a random linking code (e.g. K7P-4M2) that the midwife writes on the registry. The code has a check character that catches typos. The agent suggests possible matches (Patient 1 · Type the code · None, create · I don't know) and never creates a patient automatically when a plausible match exists.
- Protects privacy by design. Names, spouse, national ID, phone and address are masked on the image before any AI sees it, and are never stored. Health data and original photos live in a local SQLite database encrypted with AES-256-GCM, on the clinic's computer, with a tamper-evident audit log that contains no health values. Nothing goes to the cloud.
How we built it
- Schema first, not generic OCR. One JSON template per page type: 8 for the synthetic registry and 5 for the official Ministry booklet. Each field has a type (checkbox, date, blood pressure, closed list such as Neg/Pos…), allowed ranges and rules (whether RAS / IDEM are allowed).
- Pre-processing with OpenCV, no AI. Quality checks, perspective correction, dewarping of curved booklet pages using the table grid, page-type recognition from the printed layout, and slicing into checkboxes, lines and table cells.
- Masking first. Identifier zones are blacked out on the working copy before anything else. Tests check that no identifier pixel ever reaches the model.
- A local vision-language model. Qwen2.5-VL runs through Ollama on the clinic's GPU: the 7B model when it fits, with an automatic switch to the 3B one on a 6 GB card. Each handwritten cell is read twice, a guided read with schema-constrained JSON output and a free transcription, and their agreement feeds the confidence. Checkboxes are read from ink density, without AI. Multi-cell annotations (a large "RAS" across several rows, "IDEM" ditto marks, strike-throughs) are detected after reading, only in cells that are still empty, then propagated by explicit rules and confirmed in a single grouped question.
- Confidence is the product of several signals: the probabilities of the tokens of the value read, agreement between the two reads, validation rules, local image sharpness, and how much the dictionary had to correct. Below 0.8, the agent asks.
- An explicit state machine. Each step returns an event, and an orchestrator applies a transition table. The challenge's lifecycle states (captured, waiting for AI, processed, needs review, validated, patient linked, saved, synced, plus the failure states) map onto ours. Every in-progress record is saved encrypted after each step, so after a hard crash the bot resumes at the exact question it was asking.
- A WhatsApp-faithful simulator. It plays the phone, the Graph API and signed webhooks, in the exact WhatsApp Cloud API format, and it enforces WhatsApp's real limits: 3 buttons, 10 list options, text lengths, the 24-hour window, image compression. It also simulates network loss, shuffled delivery, duplicate webhooks and a clock. Switching to real WhatsApp means changing a URL and some secrets, not the code.
- Stack: Python, FastAPI, OpenCV, Ollama, SQLite and the cryptography library, plus a vanilla-JS web chat. There are 176 backend tests and 12 simulator tests.
Challenges we ran into
- Synthetic ≠ field. The synthetic pages are clean and flat. Real booklet photos are tilted, curved, on dark backgrounds, with thin pink rows and fast handwriting that overflows into the neighbouring rows. We built templates for the real Ministry booklet from one photo per page, and dewarped the pages using the printed grid.
- Reading real handwriting with a small local model. The 7B model didn't fit on a 6 GB GPU. Tiny cell crops gave the model only a handful of image tokens, so we upscale them. The model sometimes copied the field label from the prompt instead of reading the cell. It "took refuge" in RAS or IDEM on hard cells, so we now require the free read to confirm those markers. Token probabilities were dominated by the JSON syntax, so only the tokens of the value itself now count.
- False IDEM / RAS. Overflowing writing looked like multi-cell annotations. We moved annotation detection after reading, and restricted it to cells that are still empty.
- Offline is messy. WhatsApp delivers queued messages all at once, out of order, sometimes twice. This forced persist-before-acknowledge, idempotency everywhere, batching by capture time, grouping by patient, and handling the 24-hour messaging window with pre-approved templates.
- Many "improvements" didn't improve anything. We tried four OCR engines (DeepSeek-OCR, Tesseract, PaddleOCR, TrOCR, which transcribed "RAS" as the English word "has"), plus few-shot examples, multi-crop voting, ink enhancement, row-level reading and per-field vocabularies. Measured on hand-annotated real cells, none beat the baseline, and we removed the OCR engines.
Accomplishments that we're proud of
- A complete end-to-end prototype, from a WhatsApp photo to a verified, linked, encrypted record. It runs entirely on a clinic computer, with no cloud.
- On the 80 synthetic pages (ground truth from the reference PDF; the handwriting reading itself is left out of this measurement): 80/80 page types recognised, 1,970/1,970 checkboxes correct, every handwritten field detected in the right place, and 0 identifier leaks across 62 masked texts.
- An agent that says when it doubts: explicit statuses, and a value read only once is never marked KNOWN.
- Offline robustness: no record lost when the network drops, on duplicate or shuffled webhooks, on AI or database outages, or on a hard kill in the middle of a batch. The bot picks up at the same question.
- Patient linking that respects the midwife's judgement: random codes with typo detection, match suggestions, and no silent patient creation.
- Honest measurement on real photos. We built an annotation tool and measured raw handwriting accuracy on real booklet photos (28% on 180 cells with the 3B model), instead of relying only on synthetic results.
What we learned
- Start from the schema. Deciding what each field is, and which statuses it can take, mattered more than any reading engine.
- Missing information is data. Empty, illegible and not applicable mean different things for the health record.
- Measure on real data before believing an improvement. Synthetic results can be perfect while real handwriting stays hard. A small annotated set saved us from shipping "improvements" that weren't.
- When the model is weak, the UX of uncertainty matters most. One clear question with the cell's thumbnail beats a silent wrong value.
- WhatsApp's constraints shape the design: 3 buttons, 10 list items, no pre-send camera access, the 24-hour window. Building the simulator to refuse what WhatsApp refuses kept the demo honest.
- Offline-first is a state machine problem: persist first, make every step idempotent, never trust arrival order.
What's next for Day1 - The Offline Midwife
- Manual entry when the AI is unavailable for a long time: wire the existing WhatsApp Flow, or ask question by question.
- Fine-tune the vision model on annotated real cells. The QLoRA pipeline is ready; it needs 1,000–2,000 cells from several booklets and handwriting styles.
- More of the Ministry booklet: templates for the remaining pages, merging the pregnancy table that is split across two photos, and field-by-field merging when a page is re-photographed.
- Role-based access to records and original images, and stronger key protection (OS key store or TPM).
- A real WhatsApp Business sandbox, Arabic handwriting, a bilingual French/English interface, and an anonymised aggregate dashboard for epidemiology.
- A pilot with midwives to validate field conventions (what a diagonal strike or a multi-cell "RAS" really means) and to tune when the agent should ask.
Built With
- fastapi
- javascript
- numpy
- ollama
- opencv
- pymupdf
- python
- qwen2.5-vl
- whatsapp-cloud-api
Log in or sign up for Devpost to join the conversation.