Inspiration
Our inspiration came from a close friend who lives with sleep apnea and ultimately needed surgery as part of her treatment. Hearing how the condition disrupted her sleep, energy, and everyday life made us realize how serious—and easily overlooked—it can be. Because its clearest warning signs happen while someone is asleep, people may not recognize the problem until it has already affected their health and quality of life.
Her experience led us to examine sleep apnea as a local issue in Japan. Japanese clinical guidelines estimate that approximately 22 million people in Japan may have obstructive sleep apnea, including around 9.4 million with moderate-to-severe cases. Yet diagnosis normally relies on specialized sleep testing, which can make early recognition difficult.
Inspired by the theme “Local Problems, Global Solutions,” we wanted to address this major health challenge in Japan through technology that could also be accessible elsewhere. This idea became NemuruAI, a mobile-app concept powered by an experimental AI backend that analyzes overnight phone audio and highlights possible breathing irregularities that may be worth discussing with a medical professional. By using a device many people already own, we hope NemuruAI can offer a more accessible first step toward earlier awareness—both in Japan and around the world.
What it does
NemuruAI is designed around a simple overnight workflow. A user places their phone beside their bed, starts a recording before sleeping, and stops it in the morning. The app would then send the recording to NemuruAI’s AI-powered backend, which:
- Converts the recording into a consistent audio format.
- Divides the recording into overlapping time windows.
- Checks whether each window contains audio suitable for breathing analysis.
- Excludes sections dominated by speech, television, movement, clipping, or environmental noise.
- Uses a trained machine-learning model to assign a probability of a possible obstructive or mixed apnea-like sound pattern.
- Smooths predictions across neighboring windows.
- Combines overlapping flagged windows into timestamped possible events.
- Returns a structured screening report to the app.
The app concept would present the results through a clear timeline, showing when possible irregular breathing patterns occurred, which sections were excluded, and how much reliable audio was available. The backend also produces a JSON report that can be used to display the results within the app.
NemuruAI does not diagnose sleep apnea, calculate clinical Apnea-Hypopnea Index, or replace polysomnography or evaluation by a qualified medical professional. Nemuru AI primarily serves as an earlier screening suggestion.
How we built it
We first designed NemuruAI as a mobile-app concept with a straightforward user experience: record overnight sleep audio and receive an understandable summary in the morning. We then focused on building the AI model and backend needed to power that experience.
We developed the audio-analysis backend in Python using TensorFlow, Keras, TensorFlow Hub, YAMNet, Librosa, SoundFile, NumPy, SciPy, scikit-learn, Pandas, FFmpeg, and FastAPI. FastAPI allows the future mobile app to upload an audio recording to the backend and receive the model’s results as structured JSON.
The pipeline accepts common phone-audio formats, including WAV, MP3, M4A, AAC, FLAC, and OGG. Every recording is converted to mono audio, resampled to 16 kHz, and divided into overlapping windows. This allows the model to examine changes in breathing over time instead of treating an entire night as one audio sample.
Before classification, the backend runs an audio-reliability and contamination check. It measures characteristics such as volume, clipping, near-silence, signal activity, and environmental noise. YAMNet’s general-audio outputs also help identify speech-like, television-like, movement-related, or other unsuitable sounds.
YAMNet then converts each usable audio section into numerical representations called embeddings. Our custom classifier uses those embeddings to distinguish between:
- normal, event-free breathing audio
- possible obstructive or mixed apnea-like audio patterns
We trained and evaluated the model using ambient bedside-microphone recordings paired with expert-scored sleep-study annotations. Entire subjects were separated between training, validation, and testing so the model could not simply memorize one person’s breathing style, microphone, or recording environment.
We also used Claude and other AI-assisted development tools to support coding, debugging, testing, and integration of the machine-learning pipeline into the app’s backend. These tools helped us understand errors in our code, trace problems through the audio-processing pipeline, interpret model metrics, and identify why certain versions produced excessive false positives. They also supported tasks such as organizing the backend, connecting the trained model to FastAPI, checking JSON outputs, and improving compatibility with Google Colab.
The resulting backend can load a trained model, process uploaded audio, generate real probabilities, merge possible events into a timeline, and return a structured JSON report. We also created a Google Colab workflow so the inference system can be tested without installing the full development environment.
Challenges we faced
Our biggest challenge was making the model perform consistently across different people and recording environments.
One early model achieved high sensitivity, meaning it detected many labeled events, but it also incorrectly classified most normal windows as positive. For one difficult subject, nearly every negative window was flagged. After investigating, we found that the issue was not caused by a wrong microphone channel, broken annotation file, clipping, or incorrect resampling. Instead, the model struggled with subject-specific conditions such as quiet indoor noise, speech or television-like contamination, and differences between recording environments.
This taught us that thousands of audio windows do not necessarily represent thousands of independent examples. Many windows may come from the same person, room, microphone, and night, allowing a model to learn recording-specific patterns rather than patterns that generalize across people.
To address this, we narrowed the model’s target to obstructive and mixed apnea-like sounds, added contamination filtering, balanced subjects more carefully, and evaluated recall and specificity together. We also kept difficult subjects completely outside the training process so they could act as genuine tests of generalization.
Another major challenge was that clinical respiratory-event labels are created using multiple physiological sensors, while NemuruAI analyzes only audio. A real respiratory event may be visible through airflow, oxygen saturation, or breathing-effort sensors without producing a clear bedside sound. This limits what an audio-only system can reliably detect.
What we learned
We learned that creating responsible health-related AI requires much more than simply training a neural network.
Accuracy alone can be misleading. A model can achieve high recall by flagging almost everything, while an impressive overall score can hide extremely poor performance for individual subjects. Confusion matrices, specificity, per-subject results, and subject-separated testing gave us a much more honest understanding of the model.
We also learned the importance of:
- preventing data leakage
- testing models on completely unseen people
- treating contaminated audio separately
- preserving quiet breathing instead of automatically deleting silence
- balancing subjects rather than only balancing individual audio windows
- selecting thresholds using validation data rather than test data
- reporting limitations rather than presenting experimental outputs as medical conclusions
Most importantly, we learned that a model that technically runs is not automatically a model that should be trusted. Responsible AI development means being willing to reject a model when its predictions are not reliable enough.
What’s next
Our next steps are to train with more independent subjects, improve the balance of normal and event windows, test longer overnight phone recordings, and study how different phone microphones and bedroom environments affect performance.
We also hope to develop medically labeled phone-recorded datasets, improve the system’s ability to distinguish breathing from speech and environmental noise, and evaluate full-night recording results rather than relying only on short audio windows.
Built With
- ai
- audio-analysis
- claude
- data
- early-screening
- google-notebook
- health-technology
- json
- machine-learning
- numpy
- pandas
- python
- scikit-learn
- scipy
- sleep
Log in or sign up for Devpost to join the conversation.