-
-
Olu Igbo transcribing "ọ dị mma" in 1.4 seconds. It's fully offline, on a Redmi Note 10. Airplane mode confirms no network connection.
-
3-stage on-device pipeline: mel → encoder (93 MB) → cross-attn init (54 MB, once) → KV-cache decoder (172 MB). 319 MB total. No network.
-
Length-aware encoding cuts encoder latency from ~3,900 ms to ~405 ms (~9×). Total end-to-end drops from 6–8 seconds to ~1.2 seconds.
Inspiration
According to UNESCO, the Igbo language is an endangered one. There are at least 40 million speakers of the language — I'm one of them — and there's no real on-device speech-to-text for the language. Not on phones, not anywhere that doesn't require sending your voice to a cloud server first. Olu Igbo is my way of contributing to the sustenance of my language. That gap felt worth closing, and not with a cloud API wrapper, but with a model that actually runs on the kind of phone most Igbo speakers carry: mid-range Android, no NPU, no GPU, just a CPU.
What it does
Olu Igbo transcribes spoken Igbo completely on-device. You hold a button, speak Igbo, release, and the phone transcribes it locally in about 1.2 seconds. You don't need any internet connection, and your voice never leaves the device. This sample runs on a Redmi Note 10 (Snapdragon 678, 2021 mid-range) with no cloud dependency of any kind.
How we built it
First we fine-tuned a Whisper Small model for Igbo using LoRA, then a warm-start full fine-tune on FLEURS Igbo + Nigerian Common Voice + ~25k NaijaVoices utterances. Since Igbo isn't in Whisper's native 99 languages, the Yoruba token (<|yo|>) is used as a proxy at inference time.
Exported to a custom 3-stage INT8 ONNX pipeline:
- Encoder (93 MB, QUInt8) — converts mel spectrogram to hidden states
- Cross-attention initializer (54 MB, FP32) — pre-computes cross-attention KV cache from encoder output, runs once per utterance
- KV-cache decoder (172 MB, INT8) — greedy-decodes one token at a time, reusing the pre-computed cross-attention cache
Two optimizations drove most of the latency improvement. Length-aware encoding cuts encoder time from ~3.9 seconds to ~400ms by only processing the actual utterance length instead of a fixed 30-second window. KV-cache reuse cut decoder latency ~10× by building cross-attention tensors once per utterance instead of rebuilding ~110 MB of tensors on every single decode step. This also eliminated the OOM crashes.
On-device mel spectrogram computation required fixing three simultaneous bugs (power vs magnitude spectrum, STFT centering, Slaney mel normalization) and verifying against a numpy parity harness before trusting it — the Kotlin implementation now matches Whisper's feature extractor to ~1e-5 precision.
Challenges we ran into
ONNX export correctness — The decoder silently mislabeled two output tensors due to a K/V ordering mismatch between the declared output_names and what the model actually returned. The result was numerically plausible but completely wrong output. This made it far harder to catch than an obvious crash. Layer-by-layer binary search through all 12 decoder layers eventually isolated it.
Domain mismatch — Adding IgboSynCorp (a 40-hour annotated Igbo oral narrative corpus from Harvard Dataverse) consistently regressed FLEURS test WER across multiple controlled experiments, despite the extraction pipeline being correct. Oral narrative style is acoustically different enough from FLEURS' read-speech style that training on it pulled the model away from the test distribution. Documented honestly in the README.
Never trust val_loss alone — An early training run set forced_decoder_ids during training rather than only at inference, causing a 115% WER regression. Every checkpoint in this project was verified against the full 969-sample FLEURS test set before being published.
Accomplishments that we're proud of
- 41.95% WER on the FLEURS Igbo test set (969 samples), down from 68.95% zero-shot — a 27-point improvement on whisper-small
- ~10× decoder speedup from KV-cache reuse, eliminating OOM crashes in the process
- Length-aware encoding cutting encoder latency from 3.9s to ~400ms
- Fully self-contained on-device pipeline — mel spectrogram, encoder, cross-attention, decoder — all on the phone, verified to match the Python reference to ~1e-5
- 319 MB total model size, down from ~970 MB unquantized Whisper Small
- Honest, reproducible methodology: every result verified on the full test set, negative results (IgboSynCorp) documented rather than hidden
What we learned
Real on-device ML is mostly about correctness, not cleverness. The wins that actually mattered, like the mel parity fix, the KV-cache reuse, the export tensor ordering bug, all came from careful measurement and honest debugging. For a low-resource language like Igbo, the honest number (41.95% WER) is more valuable to the community than an inflated claim: it tells future builders exactly where the real difficulty lies. And val_loss is not WER. Ensure you verify on the actual test set every time.
What's next for Olu Igbo
NNAPI and GPU-DSP offload for the encoder (currently the latency bottleneck). More labeled Igbo speech data — NaijaVoices was the biggest single WER improvement, and more of the same would compound it. Eventually, a text input path for keyboard suggestions and translation, not just transcription.

Log in or sign up for Devpost to join the conversation.