-
-
Language without a connection. Real-time speech translation running entirely on-device with privacy at its core.
-
The internet is gone. The conversation isn't. Translate privately with 0 bytes of network traffic.
-
Speech → Translation → Speech. A fully on-device pipeline optimized for Arm CPUs.
-
Measured on real Arm64 hardware: latency, throughput, memory, and model performance across inference paths.
-
Product Home Page
-
Privacy Declaration Page
-
Real Test Metrics
-
Engine Used for Testing
Inspiration
What happens when the internet disappears, but communication can't?
That question became the starting point for OFFTONGUE.
Imagine a traveler stranded in an unfamiliar country, a journalist working during a blackout, a humanitarian worker in a remote region, or simply two strangers trying to understand each other when there is no signal.
Modern translation is incredibly powerful — but much of that power still assumes that the cloud is available.
We wanted to build something different.
We asked ourselves:
What if the interpreter could live entirely inside your phone?
That led us to OFFTONGUE — real-time speech translation designed to work without the internet, without a server, and without sending your voice away from your device.
Our vision is simple:
When the signal disappears, the conversation shouldn't.
What it does
OFFTONGUE turns a smartphone into a private, on-device interpreter.
A person speaks naturally, and OFFTONGUE performs the complete translation pipeline locally:
🎙 Speech → 🧠 Understand → 🌍 Translate → 🔊 Speak
There is no cloud API sitting between the user and the translation.
The system combines:
- Whisper for on-device speech recognition
- NLLB-200 for multilingual translation
- On-device TTS for translated speech
- ONNX Runtime for efficient inference
- Arm KleidiAI for optimized CPU kernels
- SME2 / I8MM / NEON acceleration where supported
- MediaPipe / XNNPACK for efficient audio processing
OFFTONGUE is designed around the principle that privacy and connectivity should not be prerequisites for communication.
The translation backbone provides up to 200-language text translation coverage, while speech input/output is exposed according to the languages actually supported and validated by the corresponding speech models.
Most importantly, OFFTONGUE can continue working in airplane mode.
0 bytes uploaded. 0 bytes downloaded. The conversation stays on the phone.
How we built it
We treated OFFTONGUE not simply as a translation application, but as an on-device AI optimization problem.
The core pipeline looks like:
🎙 SPEECH
│
▼
┌─────────────────┐
│ Audio / VAD │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Whisper │
│ ASR │
└────────┬────────┘
│
▼
┌─────────────────┐
│ NLLB-200 │
│ Translation │
└────────┬────────┘
│
▼
┌─────────────────┐
│ On-device │
│ TTS │
└────────┬────────┘
│
▼
🔊 SPEECH
The interesting part happens underneath this pipeline.
We optimized the computationally expensive stages for Arm CPUs, using quantization and Arm-optimized execution paths through ONNX Runtime and KleidiAI, with runtime dispatch to capabilities such as SME2, I8MM and NEON depending on the device.
Rather than relying on a single benchmark, we measure the complete experience:
- Time-to-first-audio
- ASR latency
- Translation latency
- End-to-end latency
- Tokens/second
- Peak memory
- CPU utilization
- Energy consumption
- Network traffic
- Translation quality
This lets us demonstrate not just that OFFTONGUE works, but why it works efficiently on Arm.
Challenges we ran into
The hardest part wasn't making one model run.
It was making three different AI workloads behave like one real-time conversation.
Speech recognition, translation and speech synthesis each have different computational characteristics. A system can have a fast translation model and still feel painfully slow if audio processing or speech synthesis introduces a bottleneck.
We therefore had to think about the entire pipeline rather than optimizing isolated components.
Language coverage was another challenge.
We discovered an important distinction:
200 translation languages ≠ 200 spoken languages.
NLLB provides broad multilingual text translation, while speech recognition and TTS have different language coverage.
Instead of hiding that limitation, we built around it and designed OFFTONGUE with explicit capability levels for:
- Speech input
- Text translation
- Speech output
This made the system more technically honest and much more robust.
Then came the mobile constraints.
Desktop AI gives you the luxury of large GPUs, abundant RAM and continuous power.
A phone gives you:
limited memory + limited battery + thermal constraints + CPU-only workloads + impatient users.
That changed how we approached optimization.
We couldn't simply ask:
"Can this model run?"
We had to ask:
"Can this model run fast enough, efficiently enough, and long enough to actually be useful?"
Accomplishments that we're proud of
The thing we're most proud of is that OFFTONGUE isn't dependent on the cloud to demonstrate its intelligence.
We can put the phone into airplane mode and still have a conversation.
That moment — seeing a phone with no connectivity still listening, translating and speaking another language — is what made the entire project feel real.
We're also proud of turning normally invisible systems-level optimization into something measurable and understandable.
Instead of simply saying:
"Arm makes AI faster."
we can show the actual execution path:
ONNX Runtime → KleidiAI → Arm CPU → SME2/I8MM/NEON
and measure the impact on:
latency, memory, energy and responsiveness.
We are especially proud that the project brings together multilingual AI, speech technology, mobile systems engineering and Arm CPU optimization into one end-to-end experience.
And perhaps most importantly:
We built an interpreter that doesn't panic when the network disappears.
What we learned
Our biggest lesson was that on-device AI is a systems problem.
It isn't enough to have a powerful model.
You need to make the entire system work within the physical constraints of the device.
We learned that:
- Model size matters.
- Memory bandwidth matters.
- Kernel efficiency matters.
- Quantization matters.
- CPU architecture matters.
- Streaming architecture matters.
- Battery matters.
- And ultimately, milliseconds matter to humans.
We also learned that optimization shouldn't be measured by a single impressive number.
A model being 3× faster means very little if it consumes twice the memory or produces significantly worse translations.
The real goal is:
More useful AI per watt, per byte, and per millisecond.
That became the engineering philosophy behind OFFTONGUE.
What's next for OFFTONGUE
This hackathon is only the beginning.
We want OFFTONGUE to evolve from a demonstration into a true offline communication platform.
🌍 More languages
Expand validated speech and TTS coverage, especially for low-resource languages that are often overlooked by mainstream translation products.
⚡ Better on-device performance
Continue optimizing the pipeline for newer Arm architectures and explore more aggressive quantization, kernel optimization and memory-management strategies.
🗣️ More natural conversations
Move from turn-by-turn translation toward continuous, low-latency conversations where OFFTONGUE feels less like a translator and more like an invisible interpreter sitting between two people.
📱 Smaller models
Reduce the storage and memory footprint so the complete system can run comfortably on a wider range of phones.
🌐 More real-world scenarios
We envision OFFTONGUE being useful for:
travelers · humanitarian teams · journalists · emergency communication · remote communities · education · field workers · privacy-sensitive conversations
And ultimately, we want to push the idea further:
AI shouldn't become less capable when the internet disappears.
With OFFTONGUE, we're trying to make the phone itself the interpreter.
No cloud. No signal. No barrier.
Log in or sign up for Devpost to join the conversation.