Install and run

AI that runs entirely on your phone, fully offline, using ~0.91 GB RAM. Install the APK and import the model pack once to get started.

  1. Download Sanad: APK + model pack about 3 GB.
  2. Extract this ZIP. It contains Sanad-v0.2.0.apk, Sanad-models-v1.zip, and README.txt.
  3. On a 64-bit Android phone (Android 8 or newer), open the APK and allow installation from your file manager. Install Sanad.
  4. Copy Sanad-models-v1.zip to the phone's Downloads folder. Open Sanad → Settings (top right) → Import model pack and select that ZIP.
  5. Keep Sanad open until it says Models ready.
  6. Choose a registry image file when prompted. Use Saved records to reopen saved results.

Results

Organizer benchmark: 80 supplied pages. These are organizer provided synthetic samples, separate from our generated training data.

Extraction metric Result
Form classification 97.5% (78/80 pages)
Field value exact match 73.5% (3,458/4,705)
Field value match ignoring case, accents, punctuation and spacing 80.1% (3,770/4,705)

Form accuracy.

Form Normalized field match
Pregnancy/postpartum monitoring 70.5%
Identification and history 84.5%
Current pregnancy 78.1%
Labor and delivery 87.4%
Early postpartum — mother 83.2%
Early postnatal — newborn 90.0%
Late postpartum — mother 66.9%
Late postnatal — newborn 87.3%

OCR: date reader 94.8% exact match / 1.03% character error rate on 600 samples; numeric reader 78.4% / 4.74% on 900 samples.

Android measurements: one complete organizer sample took 37.016 seconds, process memory reached ~0.91 GB, and model storage is 2.94 GB.

Inspiration

Sanad (roughly “healthcare support” in Arabic) started with a simple question: how can we digitize maternal and newborn health records without asking midwives to abandon the paper forms they already use?

We wanted to support healthcare workers in rural communities, where internet access can be unreliable and budgets may not stretch to powerful smartphones or computers.

Our goal is to reduce manual transcription, making digital recordkeeping more accessible, regardless of location or resources

What it does

Sanad turns photos of handwritten maternal and newborn health forms into structured records for human review. It locates fields, reads Latin and Arabic handwriting, extracts dates and measurements, and interprets checkboxes.

Reviewers can inspect the source image alongside extracted values, correct mistakes, and confirm the record. Uncertain readings remain flagged.

Photos and records are encrypted and stored locally while offline. Sanad is designed to automatically sync confirmed records to a central database when internet access returns, allowing healthcare workers to continue documenting care through connectivity gaps

How we built it

We generated 4,800 synthetic training pages across eight form types using the supplied templates, Latin and Arabic handwriting fonts, and camera-style image augmentations. This gave us labeled examples without using organizer handwriting to train the models.

Our pipeline combines:

  • MobileNetV3 to classify forms and handwriting scripts. $$ \hat{f}=\arg\max_{k\in{1,\ldots,8}}p_{\theta}(k\mid I) $$

$$ \mathcal{L}_{\mathrm{cls}}=-\log p_{\theta}(f^{\star}\mid I) $$

  • YOLO detectors to locate printed labels and handwritten lines, then associate handwriting with nearby fields. $$ M^{\star}=\arg\min_{M}\sum_{i,j}M_{ij}C(a_{i},h_{j}) $$

$$ x_{i}=\mathrm{Crop}(I,\mathrm{Match}(a_{i},H)) $$

$$ A=D_{\mathrm{labels}}^{(\hat{f})}(I),\qquad H=D_{\mathrm{handwriting}}(I) $$

$$ \mathcal{L}_{\mathrm{det}}=\lambda_{\mathrm{loc}}\mathcal{L}_{\mathrm{loc}}+\lambda_{\mathrm{cls}}\mathcal{L}_{\mathrm{cls}} $$

  • TrOCR and Arabic Nougat to recognize text, with separately fine-tuned TrOCR models for dates and numbers.

$$ r_{i}=\mathrm{Route}(\mathrm{field}_{i},\mathrm{script}(x_{i})) $$

$$ \mathcal{L}_{\mathrm{OCR}}=-\sum_{t}\log p_{r_{i}}(y^{\star}_{i,t}\mid y^{\star}_{i,1:t-1},x_{i}) $$

  • Field-specific validation to check formats, units, and ambiguity while preserving the original reading.

$$ (v_{i},s_{i})=G_{i}(\hat{y}_{i}) $$

$$ c_{i}=\min(c_{\mathrm{form}},c_{\mathrm{label}},c_{\mathrm{hand}},c_{\mathrm{OCR}}) $$

$$ \mathrm{Output}_{i}=(\mathrm{value},\mathrm{status},\mathrm{confidence}) $$

Challenges we ran into

Accurate recognition starts with accurate cropping. A clipped word assigned to the wrong field can undermine the entire result. We introduced printed-label anchors and relative positioning to improve field association.

Dates and measurements also proved difficult for general handwriting models. We trained specialist readers and added field-specific validation.

Accomplishments that we're proud of

We’re especially proud of improving how Sanad locates and crops handwriting. We trained two independent YOLO models: one detects printed labels, and the other detects handwritten lines. By matching handwriting to nearby labels, we could associate text with the correct fields and crop it more accurately, even when its position varied.

We also brought a demanding AI pipeline onto Android. Through selective quantization and loading one OCR reader at a time, we reduced the model package to approximately 2.86 GB and limited memory pressure. This technique made it possible for Sanad to run on affordable Android devices with only 4 GB of RAM.

What we learned

Document extraction is a chain of decisions. Finding the right handwriting, assigning it to the right field, choosing a reader, and interpreting the result all matter.

We learned that specialist models can substantially improve narrow tasks. We also learned that compression needs output testing: a smaller model is only useful if it preserves the information people need.

Most importantly, human review belongs at the center of the product. Making uncertainty visible is part of making the system useful.

What's next for Sanad

Over time, the structured peripartum data collected through Sanad could support research into predictive models. For example, by identifying patterns associated with maternal or newborn complications and helping healthcare teams prioritize follow-up.

Built With

Share this project:

Updates

Submission history