Inspiration

Every single day, millions of people fall victim to SMS phishing (smishing), financial scams, and text-based harassment. While solutions like Truecaller or Google Messages exist, they all share a fundamental flaw: they rely on crowdsourcing. This means they read your personal, intimate text messages and send them to the cloud for analysis. We believe that privacy is a fundamental human right. Our inspiration for Reveal was simple: to build a powerful guardian for your inbox that is smart enough to catch modern threats, but designed from the ground up to guarantee absolute privacy by never letting your data leave your device.

What it does

Reveal is a privacy-first, on-device SMS threat detection application for Android.

It acts as a silent guardian:

  1. Real-Time Interception: It instantly catches incoming text messages.
  2. On-Device Edge AI: It analyzes the text locally and classifies it as SAFE, SCAM, or HARASSMENT.
  3. Context-Aware Alerts: If a threat is detected, it floats a colored alert chip directly over your screen—but only when you are inside your messaging app, avoiding annoying pop-ups during games or videos.
  4. 100% Offline: Reveal does not even ask for Android's INTERNET permission. It is cryptographically and physically impossible for your messages to leave your phone.

How we built it

We built Reveal using a modern, performant tech stack focused on Edge AI.

  • The Model: We started with a DistilBERT architecture and fine-tuned it on an amalgamation of datasets including the UCI Spam Collection, Jigsaw Toxic Comments, SMS Phishing Dataset, and HateXplain.
  • Optimization: To run a transformer model on a phone, we used Hugging Face Optimum to perform int8-quantization, compressing the model down so it runs blazing fast on ARM64 mobile CPUs.
  • Inference Engine: We deployed the compressed model (model_quantized.onnx) into the Android app using ONNX Runtime.
  • Android App: The app itself is written entirely in Kotlin using Jetpack Compose (Material 3) for a beautiful, declarative UI.

For the actual text processing, since Python's tokenizers aren't available on Android, we wrote a custom WordPieceTokenizer entirely in Kotlin!

// Example of how we trigger classification locally
val classification = onnxMessageClassifier.classify(incomingMessageText)
if (classification.label == MlLabel.SCAM) {
    overlayService.showWarningChip()
}

Challenges we ran into

  1. Running Transformers on Mobile: NLP models are massive. Getting DistilBERT to run on a phone without draining the battery or taking 10 seconds to process a message was a huge hurdle. We had to dive deep into quantization to shrink the model weights from 32-bit floats to 8-bit integers without losing accuracy.
  2. Tokenization in Kotlin: Transformer models require text to be tokenized exactly the way they were trained. We couldn't just use a Python library, so we had to reverse-engineer and write our own WordPieceTokenizer natively in Kotlin.
  3. Android's Background Restrictions: Modern Android versions strictly limit background services. We had to carefully architect our SmsReceiver and DetectionOverlayService (using PACKAGE_USAGE_STATS and SYSTEM_ALERT_WINDOW) to trigger instantly and reliably when an SMS arrives, without violating OS battery-saving policies.

Accomplishments that we're proud of

  • Zero Cloud Dependency: We successfully built a highly accurate AI app that does not have an INTERNET permission in its AndroidManifest.
  • Lightning Fast Inference: By leveraging ONNX and int8 quantization, message classification happens in milliseconds.
  • Math on the Edge: We implemented the probability calculations natively. For instance, converting model logits to probabilities using the Softmax function locally on the phone's CPU: $$ \text{Softmax}(x_i) = \frac{e^{x_i}}{\sum_{j} e^{x_j}} $$
  • Modern UI/UX: We built a fluid, animated interface in Jetpack Compose that feels premium and native.

What we learned

  • Edge ML Deployment: We learned how to bridge the gap between Python-based model training (PyTorch/Hugging Face) and mobile deployment via ONNX.
  • Advanced Android APIs: We mastered dealing with complex Android permissions like usage stats, screen overlays, and broadcast receivers.
  • Quantization: We learned the mathematical trade-offs between model size, inference speed, and classification accuracy.

What's next for Reveal

  • Multi-lingual Support: We plan to train our model on a multilingual DistilBERT base to protect users against regional language scams.
  • On-Device Link Analysis: We want to add a compressed, local bloom filter of known malicious domains to flag dangerous URLs before the user clicks them, all while keeping the app 100% offline.
  • Accessibility Features: Integrating voice alerts for visually impaired users when a highly confident scam message is received.

Built With

Share this project:

Updates

Submission history