AmbiSense
Passive Environmental Awareness of your workspace Your phone listens to the room, not to you.
Inspiration
We started with a conversation about focus apps and why none of them work well enough. The tools on your phone measure your screen, which apps were open, but lacks to factor your external environment’s noise.
So we went the other way: measure the environment instead.
We spent the first hour not designing, but working out what Apple's on-device AI
can genuinely do well. The Foundation Models framework gives you a small language
model with a limited context window and almost no external world knowledge. That constraint shaped the whole product:
we would do the sensing with SoundAnalysis, keep the data structured and small,
and use the language model for the one thing it does better than any code we
could write.
Then Harinie reframed it, and the project became something bigger than a focus tracker. If we are already sensing the external world, why not also capture the internal one? Then we had two variables the objective room and subjective mood
That is where it transformed to a mixture of a productivity and wellness tool.
What it does
The phone sits on your desk and listens. Once per second, Apple's SoundAnalysis
classifier labels the room. From that, a dwell-based state machine works out
whether you are focused, interrupted, or idle — and crucially, how long it takes
you to recover after each interruption. Recovery time became our headline
metric: it is the thing everyone feels and nobody tracks.
Alongside it, a floating emotion bar lets you log how you feel with one tap, building an emotional arc across the session.
At the end, Apple's on-device Foundation Model reads the whole session — focus spans, interruption causes, recovery times, the sound profile of the room, and the mood arc — and writes a plain-language debrief. The most interesting thing it finds is usually the divergence: the stretch where the room was quiet and the audio said you were focused, but your mood was sinking anyway. No sensor can see that. Only the combination can.
Nothing is recorded. Audio buffers are analysed one second at a time and discarded. No transcription, no camera, no keystroke tracking, and nothing ever leaves the phone. The entire app works in airplane mode — which we demo by putting it in airplane mode.
How we built it
AVAudioEngine tap → SNAudioStreamAnalyzer (.version1) + RMS volume
→ dwell state machine → FocusState (focused / interrupted / idle)
→ session log (transitions only) + emotion taps → JSON on device
→ Foundation Models → streamed debrief + chart annotations
We started with a throwaway diagnostic app — just the classifier and a live confidence readout — because we needed to know which real-world sounds the built-in model detects reliably before designing anything around them. Typing came back at 0.95. That single result told us the product was possible.
From there, four of us worked in parallel against fixed interfaces defined in the first thirty minutes: the audio pipeline and state machine, the session log and persistence, the visualiser and UI, and the prompt design. Nobody touched anyone else's file.
Everything is on-device: SoundAnalysis for perception, a hand-tuned dwell state
machine for interpretation, and Foundation Models with guided generation
(@Generable) for the debrief and the chart annotations.
Challenges we ran into
The first version thought you were never focused. Real sessions logged 2% focus while the screen clearly showed typing detected at 0.95 confidence. The dwell rules were asymmetric the wrong way — three seconds of speech to break focus, but eight seconds of silence to regain it. In a room with any background conversation, you leave focus instantly and never get back. Inverting it, and adding a rule that typing overrides ambient speech, took a session from 2% to 92%.
Two objects fighting over one microphone. Our debug screen and our focus engine were both installing a tap on the audio bus. Only one can. Opening the debug screen silently killed the state machine, which is why our counters sat at zero for hours before we found it. Collapsing everything into a single audio owner fixed three separate bugs at once.
Making the numbers honest. At one point the summary showed 9% focused and zero interruptions — mathematically implying 91% of the session vanished. It was idle time, tracked internally but never surfaced. A dashboard whose numbers do not add up destroys trust in everything else on the screen, so we made focused, interrupted and idle always account for the full session.
Getting the model to say something worth reading. Our first debriefs just paraphrased the statistics back at you: "Few interruptions and quick recovery times." Useless. We rewrote the prompt to demand that the observation state something the numbers imply but do not say — a cause, a comparison, a turning point — and banned generic advice. We tested it against three structurally different sessions and required that all three read distinctly. That acceptance test is what made the debrief feel like insight rather than a template.
What we learned
Constraints are a design tool. Knowing exactly what the on-device model is bad at — long context, world knowledge, open-ended chat — pushed us toward the one thing it is genuinely excellent at, and the product is better for it.
We also learned that on-device is not just a privacy feature here, it is the enabling condition. A tool that recorded audio in an office would never get past a works council. One that records nothing can be deployed anywhere.
And a judge told us something useful mid-build: they were more interested in the raw perception readout — what the app is hearing right now — than in the polished output. We restructured the whole main screen around that: hearing above, concluding below.
What's next
Community. A free tier of 45-minute listening sessions each day, aimed at students, people with ADHD, and anyone working somewhere they cannot control. The people most affected by environmental distraction are usually the ones with the least power to change their environment.
Enterprise. Aggregate, anonymous wellbeing and focus data for workplaces — which floors are actually usable, which meeting patterns wreck the afternoon. Desk sensors and badge data tell you whether a space is occupied. Nothing tells you whether it is workable. Because we record nothing, this is one of the only approaches that could ever be deployed at that scale.
Validation. The honest gap: we infer focus from the environment and have not yet proven the correlation. The next step is a study against self-reported focus, so recovery time becomes a validated measure rather than a plausible one.
Built With
- apple-intelligence
- audio-classification
- avaudioengine
- avfoundation
- claude-code
- codable
- combine
- core-ml
- foundation-models
- generable
- guided-generation
- ios
- ios26
- on-device-ai
- soundanalysis
- swift
- swift-charts
- swiftui
- xcode
Log in or sign up for Devpost to join the conversation.