Inspiration
Every pregnancy is a story worth preserving. Yet for many women around the world, good prenatal care is still a luxury. In 2023, about 260,000 women died during or after pregnancy or childbirth, and nearly 92% of those deaths occurred in low and lower-middle income countries. Behind these numbers are thousands of women whose care still depends on paper registers. In many health centers, the midwife writes by hand, page after page, the details of a pregnancy that can last many months. The WHO recommends at least eight prenatal checkups during pregnancy, but globally, only 70% of women get even half that amount (four visits). In West and Central Africa, that number drops to just 57%. Once written down, the information stays on the page. It doesn't follow the woman from one visit to the next, and it never reaches the health system. So we set out to solve a very concrete problem: how can we make paper-based information follow the patient, without changing the way the midwife works? That constraint shaped every decision we made: paper remains the reference, and the digital layer builds on it rather than replacing it.
What it does
FirstBreath is a conversational agent, built as a WhatsApp-style chat, that turns a photo of a maternal register into a structured digital record. The midwife keeps filling in her register as she always has, then photographs each page, even with no network. The AI reads the page against a 245-field schema and gives every field a value, a status and a confidence level. It only asks about what it is unsure of, one question at a time, showing the cropped photo so the midwife can see exactly what is being asked. She answers with Confirm, Correct, Blank on paper, or Retake photo. To link a woman's visits together, the app generates a random code at her first visit. The midwife copies it onto the register, and that code finds the patient again at later visits, without ever storing her name. When several records might match, the app suggests candidates but never creates a record on its own: the final decision is always human. A pregnancy thus becomes one continuous record instead of a series of isolated pages. FirstBreath covers all 8 page types of the register: identification, medical history, current pregnancy (up to 9 visits), delivery, early and late postpartum, mother and newborn.
How we built it
We started from the schema, not from generic OCR. We defined 245 fields from the register's 8 sections, with a validator shared across the whole team. For each page type, a prompt describes its layout. The model transcribes, and the code normalizes: dates, choices, and a blood pressure written as "11/7" converted to 110/70. The system is offline-first. The photo is stored locally in encrypted storage (IndexedDB, Web Crypto AES-GCM), placed in an "Awaiting AI processing" queue, and processed automatically once the network returns. Each record moves through a 12-state lifecycle, from CAPTURED to SYNCED, including 4 failure states, so that nothing is ever lost, even if the connection drops mid-processing. Every field receives one of six statuses: KNOWN, UNKNOWN, NOT_PROVIDED, ILLEGIBLE, NOT_APPLICABLE or NEEDS_REVIEW. An "N/A" never hides its reason. Confidence combines the model's own confidence, the plausibility of the value, and consistency across fields (gestational age, dates), with a 0.92 threshold drawn from calibration on difficult data. The system must never invent a medical fact: when in doubt, it asks the midwife to confirm. On the technical side, the backend runs on Python, FastAPI and SQLite, with duplicate detection through SHA-256 fingerprints. The interface is built with React and Vite, designed mobile-first. We use Claude Sonnet 5 to read the pages and Claude Haiku 4.5 to recognize the page type. Development was done with Claude Code, Git and GitHub.
Challenges we ran into
The main challenge wasn't recognizing text. Registers get photographed in poor light, at an angle, and the handwriting can be hard to read. The pregnancy page alone has 160 fields, including a table of 9 visits that we had to make the model read column by column. On degraded photos, the model sometimes kept reasoning until it hit the length limit and returned a truncated answer. An automatic retry with reduced reasoning effort recovered every page. Another difficulty: on clean pages, the model rarely doubts itself. We needed deliberately degraded photos to check whether its confidence really predicted its mistakes. We also had to balance caution against workload: a conservative threshold prevents silent errors but multiplies the questions on difficult pages. We then had to align our conventions (blank, illegible or not applicable; checkboxes; "None" versus "NAD") in a shared file, so that extraction and evaluation made the same choices. All of this in 24 hours, as a team of four, using a shared data contract and separate folders for each person. Privacy was another major concern. No direct identifier is stored: name, spouse's name, national ID, phone number and address are excluded from the schema and filtered out, including in page headers. Internal identifiers are random, and only synthetic data was ever sent to the API.
Accomplishments that we're proud of
We turned an entirely paper-based process into a complete digital chain: photograph, extract, verify, link and keep. On patients never seen during development (patients 7 to 10, measured once, model frozen): 98.92% of the values read are correct (547 out of 553); only 0.65% of fields are wrong without the agent flagging them (4 out of 619); across 128 degraded photos (blur, shadow, tilt, low light, compression), silent errors stay at 0.7%, and at 0% on tilted photos; the agent asks a median of 3.5 questions per page. We also made a discovery: how long the model reasons predicts its errors. The third of pages where it reasons the most contains five times more silent errors. Above all, we didn't try to replace the midwife with AI. We designed it as a tool that works alongside her. Paper stays the source of truth, and FirstBreath finally makes that information usable from one visit to the next. For instance, on page 9 the agent reads "A" instead of the handwritten "4" in the file number, but flags it itself. In the interface, Confirm and Correct carry equal visual weight, and we avoided green, yellow and red codes so as never to suggest clinical triage. The app gives no medical advice: no diagnosis, no risk prediction, no recommendation.
What we learned
In healthcare, the best technology isn't necessarily the one that automates the most, but the one that truly fits into professionals' daily work. Starting from the schema rather than from OCR makes every field measurable and every doubt explainable. It's best to let code do what code does best: the model transcribes, Python normalizes without errors. A model's self-reported confidence isn't enough, it has to be calibrated on difficult data. And in healthcare, the interface is part of safety: it should make verification easy without encouraging blind trust. Finally, a medical system must recognize its own limits. An uncertain value should never become a false one just because an AI felt it had to give an answer. In our system, "I don't know" is a valid answer. That is why we stand by a cautious threshold: one extra question is better than one silent error.
What's next for FirstBreath
Our next steps: evaluate the system on multilingual pages, since Arabic handwriting has never been measured for lack of available pages; lighten verification on very dense pages, section by section, with the cropped photo and values side by side; automatically straighten tilted photos before reading; replace the prototype with a real WhatsApp Business Platform integration; add authentication for midwives and supervisors; adapt the schema to the register actually used in the field. In the longer term, our AI client is interchangeable: an open-source model hosted by the health system itself could replace the API, so that data never leaves the health network. Anonymized data (any cell under 5 records is hidden) could also help health facilities spot trends in pregnancy follow-up and gaps in data collection.
Research References
WHO (2025). Maternal mortality [fact sheet]. World Health Organization. (About 260,000 women died during or after pregnancy and childbirth in 2023, and roughly 92% of those deaths happened in low and lower-middle income countries)
WHO, UNICEF, UNFPA, World Bank Group & UN DESA (2025). Trends in maternal mortality 2000 to 2023. (Interagency estimates behind the WHO fact sheet)
WHO (2016). WHO recommendations on antenatal care for a positive pregnancy experience. (Recommends a minimum of eight contacts with health providers during pregnancy, up from the earlier four-visit model)
UNICEF (2024). Antenatal care. UNICEF Data, drawing on The State of the World's Children 2024. (Worldwide, only about 70% of women get at least four antenatal visits, and in Western and Central Africa the figure was 57% in 2023)
FirstBreath team (2026). Evaluation results on held-out patients 7–10 [project repository]. GitHub, firstbreath-public. (98.92% of values correct, 0.65% silent errors, measured once on a frozen model)
Built With
- anthropic-api
- claude
- claude-code
- claude-haiku
- claude-sonnet
- dexie
- fastapi
- git
- indexeddb
- javascript
- matplotlib
- pandas
- pillow
- pymupdf
- python
- rapidfuzz
- react
- scikit-learn
- sqlite
- vite
- web-crypto
Log in or sign up for Devpost to join the conversation.