Inspiration

A lot of meaning in conversation isn't in the words. It's in a pause before "sure," a flat "great," or a smile that doesn't match the sentence. Many neurodivergent people find these cues hard to read in real time, and the result is second-guessing, awkward follow-ups, or a misunderstanding neither person meant to cause. Either side can miss a cue, and neither person is the problem. We wanted a tool that treats these moments as a chance to connect rather than a social test someone failed, and that respects the privacy of everyone on the call.

What it does

CuedIn is a small floating overlay that runs next to a one-on-one video call on macOS. While you talk, it:

Transcribes both sides live, labeling your mic as "You" and the call window's audio as "Other person." Measures how things are said: pitch, pitch range, loudness, speaking rate, pauses, and response gaps, each compared to that speaker's own recent baseline. Reads facial expression over time. It tracks faces in the call window and shows a live valence meter (negative to positive), plus a soft glow around the speaker's video tile. Flags possible subtext. When the words, voice, and face don't line up, such as upbeat words delivered flat or a known sarcastic phrase with unusual emphasis, CuedIn shows a short insight card with the phrase, a tentative reading, and sometimes a gentle way to check in.

CuedIn presents every reading as tentative. It never claims to know what someone feels or intends, and it stays quiet when the evidence is mixed.

How we built it

macOS overlay (SwiftUI/AppKit): ScreenCaptureKit captures the selected call window's audio and video. AVAudioEngine captures the mic and resamples it to 16 kHz mono. Apple Vision detects face landmarks, head pose, and mouth and eye openness at 5 fps, and on-device text recognition reads participant names from their meeting tiles. Local Python service (FastAPI + WebSockets): it receives audio and small face crops over the loopback address only. Audio: energy-based voice detection feeds faster-whisper (base.en, int8 on CPU). Silero VAD, plus Whisper's no-speech, log-probability, and compression-ratio checks, filter out hallucinated text from silence. We extract prosody features in the same pass. Video: a ResNet50 encoder from EMO-AffectNet turns each face crop into a 512-value feature vector. An LSTM reads the latest 10 frames per face track and predicts 7 expression classes, which we map to a signed, smoothed valence score. Mismatch gate: a local heuristic flags possible sarcasm only when both the wording and the vocal delivery look unusual. Alignment pipeline: transcript, audio, and visual cues get session-relative timestamps and IDs, and are merged into 14-second windows with a one-second timeline index. Gemini: each window's transcript and numeric features go to Gemini with a strict JSON schema. Gemini must cite the evidence IDs it used, and we check its quote against the real transcript and the overlapping audio and visual cues before anything reaches the UI. Raw audio, video, and face crops never leave the machine, and all data lives in memory only.

Challenges we ran into

Whisper hallucinating on silence. Early builds turned background noise into full sentences. We fixed it with a two-stage filter: a permissive local voice gate, then Silero VAD plus per-segment confidence checks. Microphone capture. Some mic arrays cancel out when their channels are averaged, default-device switches left AVAudioEngine delivering silent buffers, and reinstalling an audio tap can crash the app. We ended up measuring each channel and routing the loudest one, and added a watchdog that rebuilds the audio engine when input stalls. Backpressure. Audio, video, and face crops share one socket. Under load, stale video frames delayed audio, so we built a send queue with priorities that keeps only the newest frame and never drops a mute command. Aligning timing across modalities. Whisper returns text seconds after the speech happened, so we had to align everything to capture timestamps rather than arrival time. Keeping Gemini grounded. The model sometimes quoted words nobody said or cut its JSON short once thinking tokens used up the output limit. Strict schemas, required evidence IDs, and our own check against the real transcript fixed most of it. Sequence models on live video. The LSTM needs 10 consecutive frames of the same face, so we had to keep face tracks stable across frames without identifying anyone.

Accomplishments that we're proud of

A full real-time pipeline running on a laptop: transcription for both speakers, prosody, facial valence, alignment, and LLM interpretation all happen within a few seconds. A privacy design we can stand behind: raw media stays on-device, only text and numbers go to the cloud, and nothing is written to disk. A multimodal gate that keeps the overlay mostly quiet. It speaks up when the words, voice, and face disagree, not every time someone pauses. Prompting and UI copy that stay tentative, avoid judgment, and don't assume neurotypical norms for eye contact, expression, or response speed.

What we learned

Deciding when to stay quiet matters more than detecting everything. A companion that interrupts too often gets turned off. Facial expression or tone alone is weak evidence. The useful signal is in mismatches between channels, and even that should be treated as a hint, not a conclusion. Real-time audio on macOS has many edge cases around devices, formats, and permissions. Diagnostics and heartbeat logging saved us hours. LLMs are much more reliable when you make them cite evidence and then check it yourself.

What's next for CuedIn

Test with neurodivergent users to learn which cues actually help and how they should be shown. A better mismatch detector: compare the emotion distributions from text, voice, and face directly instead of relying on hand-written heuristics. Speaker diarization so group calls work, plus detecting the meeting app's own mute state. A valence model trained for continuous valence rather than mapped from expression classes, with calibration and a CPU fallback. More languages and platforms, starting with Windows. Optional post-call reflection: a private summary of moments worth revisiting, deleted by default.

Built With

Share this project:

Updates

Submission history