Inspiration

Midwives in Morocco record pregnancy follow-up in a paper register. DayOne asked for a way to get that data out of the paper, at low cost, without sending patient data to a cloud AI.

What it does

The midwife sends a photo of a register page to a messaging bot. The bot identifies the type of page, then extracts each field of that type of page. Values read with high confidence are marked as sure; every checkbox and every uncertain reading is marked "to confirm"; an unreadable value is reported as unreadable instead of being guessed. The paper register stays the reference.

How we built it

  • A Telegram bot written with the Python standard library only.
  • Two passes: classify the page type, then extract under a checklist for that page type, with a forced JSON schema and a confidence per field.
  • No third-party cloud AI. By default an open model (Qwen3.8-27B) runs on a private GPU server; in our tests this was a university GPU server reached through a tunnel, standing in for a clinic or district server. Without any server, a smaller model (qwen2.5vl 7B) runs fully on the local machine through Ollama.
  • A calibrated confidence built from the token probabilities of the model, which is more reliable than the confidence the model reports itself.

Challenges we ran into

  • WhatsApp cannot be connected to a local model, so the demo uses Telegram as transport. In production we would use a self-hosted web app on the clinic network.
  • Resolution matters a lot. On 79 verifiable cells from 20 pages, exact reading goes from 66 % at 896 px to 92 % at 1280 px and 100 % at native resolution. The practical advice is to send the page as an uncompressed file.
  • Large visit grids can still shift a value by one column; this is why the midwife reviews the result.
  • A general-purpose model we tried (Gemma) read only 29 % of the cells and was dropped.

Accomplishments that we're proud of

  • Results that are measured, not claimed: an independent audit of 20 pages found about 96 % of the extracted fields correct for the fully local model, with six named failure modes.
  • The bot says "unreadable" instead of inventing a value (tested on a masked name).
  • Honest limits: the two accuracy figures above do not measure the same thing, and throughput on consumer hardware was not measured.

What we learned

  • A document-specialist model can beat a larger general model on handwritten digits.
  • Confidence from token probabilities separates right from wrong readings better than self-reported confidence.

What's next for The Offline Midwife

  • Fix the column shift in visit grids and the pages for newborns.
  • Encrypted local storage and correction buttons in the chat.
  • Linking the visits of the same patient by file number.
  • A self-hosted, WhatsApp-like web app for the clinic network.

Built With

Share this project:

Updates

Submission history