Inspiration
Every visit adds to a mother’s care journey. But when that journey is recorded on handwritten pages, important information can be hard to carry forward—especially where internet access is unreliable. DayOne is our effort to help change that.
What it does
DayOne is a local-first maternity-record prototype, not just an OCR tool. It brings together page capture, source-linked transcription suggestions, human review, visit linking, conversational access, privacy controls, and an anonymized dashboard.
We tested several models on 16 manually selected handwritten crops from one booklet. GPT‑6 achieved 14/16 exact readings with optimized instructions, compared with 11/16 with the initial instructions. On the particularly difficult remaining images, our team estimated that about 12% were genuinely ambiguous. These crop results measure transcription, not full-page field extraction.
DayOne also includes a Telegram bot, a WhatsApp Business-compatible sandbox, encrypted local storage, a processing queue designed to resume after a restart, image-quality checks, and bilingual interface elements. The prototype uses fictional records.
How we built it
We built a local-first web application using Python, FastAPI, SQLite, JavaScript, HTML, and CSS. The application includes local OCR, source-linked suggestions, human review, patient and visit linking, and a persistent processing queue.
The Telegram bot supports the conversational demonstration. The repository also contains a WhatsApp Business-compatible text-message sandbox. After dependencies and model weights are installed, local processing can run without sending records to a cloud OCR service.
Handwriting model results
| Model tested | Result |
|---|---|
| GPT‑6 — optimized | 14/16 exact readings |
| GPT‑6 — initial | 11/16 |
| PaddleOCR‑VL 1.6 (fine-tuned on handwritten) | 7/16 |
| PaddleOCR‑VL 1.6 | 4/16 |
| EasyOCR, without handwriting-specific training | 3/16 |
| Qwen3‑VL‑4B | 3/16 |
| Qwen3‑VL‑4B with image context | 5/16 |
| TrOCR | 0/16 |
| Groq Qwen — separate 4-crop test | 1/4 |
On the particularly difficult remaining images, our team estimated that about 12% were genuinely ambiguous. These crop results measure transcription, while full-page field extraction is a separate challenge.
Challenges we ran into
Reading a selected handwriting crop is different from finding and connecting fields across a full page. GPT‑6 produced our strongest crop result, but the local application pipeline found zero structured fields in five phone photos from one booklet. That exposed the gap between model transcription and a complete page-to-record workflow.
We also learned that ambiguous handwriting should remain visible as ambiguous. A plausible guess is not a reliable record.
Accomplishments that we're proud of
We built a connected prototype covering the challenge’s main workflow areas: extraction suggestions, conversational interaction, uncertainty handling, offline resilience, privacy and record linking, and code documentation.
The product also includes an anonymized aggregate dashboard, image-quality checks, bilingual interface elements, a Telegram bot, and a WhatsApp Business-compatible sandbox. Aggregate outputs suppress small groups.
Arabic handwriting recognition
Prototype capabilities
| Area | What is implemented |
|---|---|
| Capture and review | Page capture, source-linked suggestions, human correction and confirmation |
| Continuity between visits | Patient and visit linking with human confirmation |
| Conversational workflow | Telegram bot; WhatsApp Business-compatible synthetic-message sandbox |
| Offline resilience | Local storage and a processing queue designed to resume after restart |
| Privacy and reporting | Encrypted local records and anonymized aggregate dashboard |
| Bonus-oriented features | Image-quality checks and bilingual interface elements |
What we learned
Comparing models on the same selected handwriting crops helped us identify GPT‑6 as the strongest tested option. It also showed that crop-reading results do not establish full-page extraction performance. A usable record workflow needs provenance, human correction, visit linking, and a clear way to handle uncertainty—not only a transcription model.
What's next for DayOne (The Optimizers)
We plan to improve full-page field detection and visit attribution, then evaluate them on independently annotated pages. We also want to test handwriting-specific model adaptation and measure Arabic handwriting separately before making claims about its recognition quality.
Log in or sign up for Devpost to join the conversation.