Inspiration
In an online class, the student who is lost is the one who goes quiet. They don't unmute, they don't raise a hand, they don't type in chat. The signal that they need help is an absence of signals, which is exactly what a teacher managing thirty tiles cannot see. And the usual fix, calling on someone directly, is the thing that makes a struggling kid never speak up again.
What it does
Anchor is a macOS menu bar app that watches a live Zoom class and surfaces students who are falling behind, while the teacher can still do something about it.
Two signal sources feed one score:
- Live Zoom behavior: mute state, cumulative unmuted time, speaking duration, camera state, hand raises, chat message length, hesitation markers, and whether the student has asked a question.
- Google Classroom: missing assignments, grade average, grade trend, days since last submission, late submissions.
Those become a 16 feature vector scored on device by a Core ML classifier. The output maps to three levels: Engaged, Watch (0.40), and Needs attention (0.70), with a teacher adjustable sensitivity slider that shifts both cut-offs.
Then it says what to do about it. Anchor combines the score, the live lesson topic pulled from transcript capture, and that student's history on the topic to produce a specific question to ask a specific student. Apple's Foundation Models write the phrasing, on device. If the model is unavailable or times out, a deterministic fallback states the same facts in Anchor's own voice, so the feature degrades in phrasing quality and never in correctness.
Every score comes with a "Why this score" panel that attributes the probability to individual features by measuring each one's contribution against the loaded model, rather than asserting a hand written rule.
How we built it
- Swift and SwiftUI on macOS 26.5, built on NSStatusItem and NSPopover.
- Zoom Meeting SDK, bridged from its delegate based Objective-C API to Swift async/await. Mute state, camera state, raised hands and active speaker are all in-meeting client state, so the only way to observe them is to be a client in the meeting.
- OAuth for both Zoom and Google using authorization code with PKCE over a loopback redirect. A desktop app cannot keep a secret, since anything in the bundle is readable, so the proof of possession is a per request verifier that never leaves the process. Only refresh tokens are stored, in the macOS Keychain.
- Core ML gradient boosting classifier, selected over random forest by a 220 candidate, 5 fold cross validated search.
- Apple Foundation Models for topic extraction and recommendation phrasing, entirely on device. The input is a recording of children speaking in a classroom cross referenced with their grades, so there is no version of this feature that would be acceptable if the text left the Mac.
- NaturalLanguage for lemma overlap and word embeddings, matching a live lesson topic against assignment titles.
The model trains on 10,000 synthetic sessions from a latent variable generator. Each student has an unobserved engagement and academic state, the features are noisy observations of it, and the label is drawn from the state, so irreducible error arises the way it does in reality instead of being sprinkled on afterwards as label flipping.
Held out on 2,000 unseen rows: 81.10% accuracy against a 67.70% majority class baseline, ROC AUC 0.8676, recall on struggling students 63.62%.
Challenges we ran into
The one worth telling: we shipped a model, measured it, and found it returned 0.0% for an obviously struggling student. Every child read as fine, while the UI truthfully reported which model was doing the scoring.
It had been trained with missing_assignments constant at 0, which is zero variance and therefore unlearnable, and grade_trend on a raw -2 to 19 scale while production sends 70 to 130. Every live prediction was made on a value the model had never seen anything near. A fallback that fails silently in the exact direction this product exists to prevent is worse than no fallback, so it came out of the candidate list entirely.
Fixing it surfaced two more. The replacement data generator built every row as a noisy clone of one of 8 fixed archetypes, so the feature space was barely covered. And confidence_level was drawn from its own gaussian rather than computed with the formula production uses, so the column did not mean the same thing in training and at runtime.
Separately, scores had to be ramped. A class opens with everyone muted, cameras settling, nobody speaking yet. Feature by feature that is indistinguishable from a room full of disengaged students, so an unramped score lights the dashboard red in the first thirty seconds and cries wolf before the lesson has started.
Accomplishments that we're proud of
Catching the 0.0% bug at all. It took building an evaluation harness that checks the training data's feature ranges against the Swift that produces them, rather than trusting that a model which loads is a model that works.
Choosing the operating point on the right grounds. The default Watch threshold is deliberately the recall favouring one, because a missed struggling student is the failure this product exists to prevent, and a false alarm costs a teacher one question.
What we learned
That a model which loads, runs, and returns a confident number can be completely broken, and nothing in the app will tell you. The only thing that catches it is checking what the training data actually contained against what production actually sends.
That the honest comparison for a classifier on imbalanced data is the majority class baseline, not the raw accuracy. 81.1% sounds good until you notice that answering "coping" every single time scores 67.7%.
What's next for Anchor
Retraining on real labelled classroom sessions. Anchor already exports labelable rows with all 16 features; the struggle column has to be filled in by the teacher, because deriving it from Anchor's own prediction would just train the next model to agree with this one.
Publishing the Zoom Marketplace app, so teachers outside the developer's own Zoom account can authorize it and run a real pilot.
Built With
- core-ml
- google-classroom
- macos
- oauth2
- objective-c
- python
- scikit-learn
- swift
- swiftui
- zoom
Log in or sign up for Devpost to join the conversation.