Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for SOUNDMARK EDGE

Inspiration

Most sound-alert apps ship a fixed vocabulary, but the sound that matters may be one old washer, a studio count-in, a lab timer, or a door mechanism unique to one place. Uploading microphone data to train a cloud model adds privacy, connectivity, latency, and service-cost risks.

What it does

SOUNDMARK EDGE is a Mobile AI proof of concept that lets an Arm64 Android phone learn a custom sound from 3–5 local one-second examples. The phone creates a tiny cosine prototype and class threshold, then recognizes the sound in a foreground stream. Top-two margin, UNKNOWN rejection, 2-of-3 temporal agreement, and cooldown suppress uncertain or duplicate events. Accepted events produce a visual alert, optional vibration, and deletable local metadata.

A login-free Guided Demo runs real bundled INT8 inference for three known synthetic sounds plus an UNKNOWN fixture. An Optimization Lab executes FP32 and INT8 on the same input and exports a JSON report containing device, runtime, model, APK hash, and an explicit physical-device verification flag.

How we built it

The reference path deterministically generates Apache-2.0 synthetic audio, computes a 16 kHz 64×96 log-mel tensor, pools it to 16×16, fits a compact 32-D projection, exports FP32 TFLite, and performs full-integer INT8 post-training quantization with 64 training-only calibration fixtures.

Kotlin reproduces the frontend, memory-maps the SHA-256-verified model, runs LiteRT 2.1.4 with XNNPACK and two CPU threads, and drives Jetpack Compose, AudioRecord, prototype personalization, an adaptive energy gate, temporal smoothing, app-private storage, and the benchmark runner.

Arm optimization and measured evidence

The Android package targets arm64-v8a only and includes LiteRT's Arm64 native runtime.

  • FP32 model: 34,384 bytes
  • Full INT8 model: 10,776 bytes
  • Serialized size reduction: 68.7%
  • 56-fixture synthetic 3-shot macro F1: 0.9353 for both
  • Known recall: 0.8929 for both
  • UNKNOWN rejection: 1.0000 for both
  • Minimum INT8/reference embedding cosine: 0.999554
  • Observed 3-shot metric delta: 0

These quality values come from deterministic synthetic fixtures and are not real-world accuracy claims.

Privacy and accessibility

The APK has no INTERNET permission or analytics SDK. Raw microphone windows stay in memory and are discarded. Only prototype vectors and up to 100 alert metadata records use app-private storage excluded from backup. Leaving the foreground stops capture. The high-contrast interface uses text plus color states, semantic headings, large controls, explicit errors, and a prominent non-emergency disclaimer.

Challenges and lessons

The hardest boundary was keeping Python and Android preprocessing identical. A byte-for-byte PCM fixture and float32 golden tensor gate the Kotlin FFT/mel frontend at 5e-4 maximum error. We also found and fixed an emulator identity bug that could have mislabeled a virtual Arm runtime as physical evidence. The final report now rejects emulator and non-Arm runs.

Verification boundary

No physical Arm Android benchmark was available for this P0 submission. Emulator screenshots and reports are integration smoke evidence only and carry armBenchmarkVerified=false. Apple M1 reference timings are also excluded from Android performance claims. Physical microphone, vibration, thermal, and latency verification remains documented follow-up work.

Try it

No login, API key, network, or external model download is required.

Screenshots

All screenshots below were captured on an API 35 Arm64 emulator and are integration evidence, not physical-device performance evidence.

Guided Demo with known and UNKNOWN fixture inference

Privacy and Trust screen

Benchmark JSON export explicitly labeled EMULATOR / UNVERIFIED

Built With

Share this project:

Updates

Submission history