Inspiration
Midwives in Morocco record pregnancy follow-up in a paper register. DayOne asked for a way to get that data out of the paper, at low cost, without sending patient data to a cloud AI.
What it does
The midwife sends a photo of a register page to a messaging bot. The bot identifies the type of page, then extracts each field of that type of page. Values read with high confidence are marked as sure; every checkbox and every uncertain reading is marked "to confirm"; an unreadable value is reported as unreadable instead of being guessed. The paper register stays the reference.
How we built it
- A Telegram bot written with the Python standard library only.
- Two passes: classify the page type, then extract under a checklist for that page type, with a forced JSON schema and a confidence per field.
- No third-party cloud AI. By default an open model (Qwen3.8-27B) runs on a private GPU server; in our tests this was a university GPU server reached through a tunnel, standing in for a clinic or district server. Without any server, a smaller model (qwen2.5vl 7B) runs fully on the local machine through Ollama.
- A calibrated confidence built from the token probabilities of the model, which is more reliable than the confidence the model reports itself.
Challenges we ran into
- WhatsApp cannot be connected to a local model, so the demo uses Telegram as transport. In production we would use a self-hosted web app on the clinic network.
- Resolution matters a lot. On 79 verifiable cells from 20 pages, exact reading goes from 66 % at 896 px to 92 % at 1280 px and 100 % at native resolution. The practical advice is to send the page as an uncompressed file.
- Large visit grids can still shift a value by one column; this is why the midwife reviews the result.
- A general-purpose model we tried (Gemma) read only 29 % of the cells and was dropped.
Accomplishments that we're proud of
- Results that are measured, not claimed: an independent audit of 20 pages found about 96 % of the extracted fields correct for the fully local model, with six named failure modes.
- The bot says "unreadable" instead of inventing a value (tested on a masked name).
- Honest limits: the two accuracy figures above do not measure the same thing, and throughput on consumer hardware was not measured.
What we learned
- A document-specialist model can beat a larger general model on handwritten digits.
- Confidence from token probabilities separates right from wrong readings better than self-reported confidence.
What's next for The Offline Midwife
- Fix the column shift in visit grids and the pages for newborns.
- Encrypted local storage and correction buttons in the chat.
- Linking the visits of the same patient by file number.
- A self-hosted, WhatsApp-like web app for the clinic network.
Log in or sign up for Devpost to join the conversation.