Inspiration
Health data today is abundant, yet almost entirely disconnected. A single person generates heart-rate, heart-rate-variability, sleep, and activity data from a wearable; if mood and stress are tracked at all, they live in a separate check-in app; medical history and lab reports sit in PDFs or a clinic's own isolated system. None of these sources speak to one another, and the consequence is twofold. Individuals are shown numbers without meaning — a resting heart rate of 78 bpm, presented against a generic "normal" range of 60–100 bpm that says nothing about whether it is unusual for that specific person. And clinicians see only a snapshot, not a trajectory: a ten-to-fifteen-minute consultation captures a single point in time, while weeks of gradually accumulating deviation — the kind that often precedes a more acute problem — remain invisible unless the patient happens to mention it. This is not a marginal concern in the population we set out to serve. Exam-period stress prevalence rises to 43% among Indian medical students, accompanied by a near-doubling of salivary cortisol. 12.5% of surveyed Indian college students aged 18–26 show severe depression on standardized screening. We were struck by a simple observation underlying both findings: the gap is not a lack of data. It is the absence of a system that turns fragmented, individually meaningless numbers into a personalized, explainable signal — and does so honestly, without pretending to diagnose. That observation became NeuroVitals.
What it does
NeuroVitals is built on a single architectural decision that shapes everything else: instead of asking "is this number normal?" against a population threshold, it asks "is this unusual for you?" against your own historical baseline. The application ingests daily wearable vitals — resting heart rate, heart-rate variability, sleep, and activity — alongside a short daily mood and stress check-in, and treats these two categories of data as mechanistically linked rather than tracking them in isolation, which is how most existing tools handle them.
When a day's readings depart meaningfully from a user's own baseline, the system does not simply raise an alert. It generates an explanation in plain language, stating what changed, what data supports that reading, and why it is the most significant contributor to the flag — for instance, "Resting HR +9 bpm vs. your baseline, coinciding with reduced sleep. Most significant contributor: sleep consistency. This is not a diagnosis." No clinician involvement is required to use it, and, deliberately, the current version stops there: no doctor marketplace, no clinical decision support, and no emergency alerting layer. That boundary was a design decision made early and kept deliberately, not a limitation we ran out of time to address.
How we built it
The ingestion layer was designed around the schemas of Apple HealthKit and Google Health Connect, so that wearable and check-in data could be treated uniformly regardless of source device. On top of that data sits the baseline engine: a statistical model of each user's own rolling mean and variance, paired with Isolation Forest for anomaly detection and XGBoost for pattern modeling. Isolation Forest was chosen specifically because it tolerates the cold-start problem inherent to a personalized system — a new user has little historical data to learn from, and the model still needs to behave sensibly from day one. For each metric, we maintain a rolling average and a measure of typical day-to-day variation over the user's own trailing window — their personal baseline. Every new reading is compared against that baseline rather than a fixed clinical range, and the size of the departure is what determines whether a metric gets flagged as a deviation. When several metrics deviate on the same day, they are ranked by how far each one strayed from that person's own norm, which is what surfaces the "most significant contributor" shown in the insight card — a ranking the SHAP-based explainability layer formalizes further once the XGBoost pattern model is layered on top of the baseline.
Challenges we faced
The central technical difficulty was balancing sensitivity and specificity for a baseline that is, by design, personal rather than population-level — a cold-start problem that generic threshold-based models simply do not have to solve, since they never need to learn an individual at all. The harder challenge, though, was one of restraint rather than engineering: it was genuinely tempting to pitch the doctor marketplace and clinical decision support layer we had already designed, because a larger vision is more exciting to describe. Cramming that later-phase vision into a first-round submission would have weakened our feasibility standing rather than strengthened our ambition, so we chose to submit only what was actually built and demonstrable. We also spent real iteration time making "explainable AI" mean something concrete rather than serving as a buzzword — arriving at a fixed "what changed → evidence → why it matters" format took several passes before it felt honest rather than decorative.
Accomplishments that we're proud of
We are proud, above all, of having built a differentiator that is architectural rather than cosmetic. Comparing a user against their own history, rather than against everyone else's, is a genuine structural departure from how Apple Health, Whoop, and Oura approach the same data — not a different color scheme applied to the same underlying comparison. We are equally proud that explainability is load-bearing in the design rather than a summary bolted on afterward: every insight the system produces can point directly to the evidence behind it.
What we learned
Personalization in health technology turned out to be a modeling problem before it was ever a design problem — the dashboard was the easy part; the baseline was the real work. We also learned, somewhat unexpectedly, that scope discipline is itself a skill worth developing deliberately: deciding what not to present in a first submission mattered as much as deciding what to build.
What's next for NeuroVitals
The next phase introduces doctor-shareable health summaries on a FHIR-style interoperable schema, cross-device calibration to normalize signal quality across wearable brands, and — once sufficient longitudinal per-user data exists to support it — a transition from Isolation Forest toward LSTM or autoencoder-based anomaly detection for richer temporal pattern recognition.
Built With
- google-health-connect
- healthkit
- machine-learning
- python
- react
- scikit-learn
- shap
- supabase
- tailwindcss
- typescript
- xgboost
Log in or sign up for Devpost to join the conversation.