Challenge Track: Physical AI

ReliefRelay is submitted to the Physical AI track.

It processes real-world audio sensor input on nearby Arm64 edge hardware and transforms it into human-reviewed incident intelligence that supports a physical emergency-response workflow.

Inspiration

Emergency response does not happen under perfect conditions. Field teams work through damaged infrastructure, unreliable connectivity, noisy radios, constrained computing resources, and rapidly changing information.

A report such as “three people trapped near Riverside Shelter requesting a rescue boat” contains everything responders need—but only if the location, severity, affected people, and requested resource can be recovered accurately.

Most speech systems stop at transcription and often depend on cloud connectivity. ReliefRelay began with a more operational question:

What if an Arm-powered device could turn degraded field audio into a reviewed response incident locally, even when the network is unavailable?

That became ReliefRelay: a local incident-intelligence layer designed around the operator who remains responsible for the final decision.

What it does

ReliefRelay converts radio-style emergency audio into structured, prioritized incidents on nearby Arm64 hardware.

An operator can:

  1. Record a report using the browser microphone.
  2. Upload an existing WAV recording.
  3. Test with included clear, radio, and severely degraded scenarios.
  4. Run local transcription through an optimized Whisper model and whisper.cpp.
  5. Review the transcript, confidence, warnings, location, incident type, severity, affected people, and requested resources.
  6. Correct or confirm the AI-generated draft.
  7. Acknowledge, assign, dispatch, resolve, or reject the incident.
  8. Inspect every original report, correction, assignment, and lifecycle event.

Every machine-created incident begins in needs_review. ReliefRelay never treats model output as authorization to dispatch. The AI structures the report; a human owns the operational decision.

Similar open reports can be consolidated without erasing their individual sources. Resolved and rejected incidents cannot absorb future reports, and later reports cannot silently lower the highest known severity.

Inference, extraction, review, and persistence operate locally without sending operational audio to an external AI API.

How we built it

ReliefRelay is a complete local response workflow rather than a standalone transcription demonstration.

The inference layer uses Whisper Tiny English through checksum-pinned whisper.cpp v1.9.2. Our optimization pipeline generates a Q5_1 model from a verified full-precision baseline and evaluates both variants under identical conditions.

A FastAPI service validates WAV files, controls the inference queue, applies concurrency and process timeouts, and passes transcripts into a deterministic emergency-report extractor.

The extractor handles:

  • known response locations and numeric addresses;
  • negated incident language;
  • compound number words;
  • severity indicators;
  • affected-person counts;
  • requested resources with explicit request intent; and
  • conservative incident deduplication.

The response dashboard uses dependency-light HTML, CSS, and JavaScript. It supports microphone recording, WAV uploads, editable review forms, incident prioritization, assignment, lifecycle controls, report history, and audit inspection.

SQLite in WAL mode provides durable local storage. Docker and Compose provide a non-root deployment path. Environment configuration controls authentication, storage, model paths, threads, inference concurrency, queue limits, and timeouts.

We also created a reproducible Arm optimization workflow:

  1. Download and SHA-256 verify the baseline model.
  2. Build or provision pinned whisper.cpp.
  3. Generate Q5_1 from the verified baseline.
  4. Benchmark both models using the same device, runtime, corpus, threads, warmups, and run counts.
  5. Record every model hash, transcript, timing, environment property, and task result.
  6. Reject comparisons that violate provenance, accuracy, footprint, or latency requirements.
  7. Repeat the process on native Arm64 GitHub Actions runners.

Meaningful challenge-period work

ReliefRelay was meaningfully rebuilt during the Arm Create challenge period.

We expanded an early transcription prototype into a complete incident-operations product with optimized local Arm inference, persistent storage, human review, incident assignment and lifecycle management, report history, audit events, runtime protection, Docker deployment, native Arm64 CI, and reproducible before-and-after benchmarks.

Challenges we ran into

The first challenge was that transcription quality and operational correctness are not the same thing.

A transcript can have a relatively low word error rate while still damaging the one field that matters most, such as a location or requested resource. We therefore measure structured-field accuracy alongside WER and latency.

Noisy speech created additional risks. Negated phrases could create false incidents, number words could be interpreted incorrectly, and resource terms could appear when no resource was requested. We added deterministic safeguards and made uncertainty visible to the operator.

The second challenge was safe incident consolidation. Naive deduplication can merge unrelated unknown-location reports or attach a new emergency to a resolved event. ReliefRelay limits merging to similar open incidents inside a configurable time window and preserves every source.

The third challenge was proving the Arm optimization fairly. Running on Arm was not sufficient. Both models needed identical runtime conditions, verifiable provenance, repeated measurements, and explicit quality thresholds.

The final challenge was protecting responsiveness during CPU-heavy local inference. We introduced Arm-aware thread tuning, bounded concurrency, queue limits, and transcription timeouts.

Accomplishments that we're proud of

We turned an early prototype into a working incident-operations product with persistent data, human review, assignment, response states, report history, and auditability.

Our native Apple M4 Arm64 comparison produced:

  • 58.6% smaller model: 74.10 MiB to 30.68 MiB
  • Unchanged median inference: 0.273 seconds
  • Improved p95 inference: 0.307 to 0.302 seconds
  • Observed WER improvement: 11.23% to 8.79%
  • Structured-field accuracy: 97.78% to 100% on the submitted corpus
  • 126 measured native Arm64 inferences

The repository includes raw reports, individual transcripts and timings, runtime metadata, hashes, quantization provenance, screenshots, and reproduction commands.

ReliefRelay also has 35 passing automated tests and native Arm64 CI that rebuilds the optimized model, executes tests, benchmarks both variants, and enforces the quality guard.

Most importantly, ReliefRelay does not hide uncertainty or automate authority. Every AI-created incident remains reviewable, original reports remain preserved, and operators control every response decision.

What we learned

Meaningful edge-AI optimization is a system problem, not only a model problem.

A smaller model reduces storage and deployment cost. Stable tail latency matters because operators experience slow requests, not just the median. Task accuracy matters because responders need actionable fields rather than a transcript score. Queue controls matter because edge compute is finite.

We also learned that deterministic logic and machine learning complement one another. Whisper handles speech recognition, while explicit rules protect negation, resource intent, lifecycle transitions, and human authority.

Finally, transparent limitations make a product more credible. Our nine-fixture synthetic corpus demonstrates reproducibility, not universal speech accuracy. Real deployment requires consented field evaluation, multilingual testing, organizational security review, and authorized response-system integration.

What's next for ReliefRelay

Our next steps include:

  • validating performance and energy usage on Raspberry Pi 5 and Arm Neoverse;
  • building a consented multilingual and multi-accent radio corpus;
  • adding coordinate-aware geospatial resolution;
  • integrating authorized radio gateways and dispatch notifications;
  • replacing the shared token with organizational identity and role controls;
  • supporting PostgreSQL for multi-node deployments;
  • adding offline replication between disconnected response posts; and
  • piloting the workflow with emergency-management specialists.

ReliefRelay demonstrates a practical Physical AI loop:

real-world audio → optimized local Arm inference → structured intelligence → human review → physical response

When connectivity disappears, the response workflow should not disappear with it.

Built With

Share this project:

Updates

Submission history