Inspiration
Healthcare is rapidly moving beyond hospital walls. Patients starting high-risk medication regimens or recovering at home often experience subtle, adverse physical deteriorations long before an acute emergency occurs. While continuous monitoring is ideal, clinicians cannot observe every patient 24/7 At the same time, centralizing live biometric telemetry onto external servers raises severe privacy concerns, latency risks, and complex regulatory hurdles.
We asked ourselves: What if an intelligent, clinically grounded triage system could live entirely on the patient's wrist and phone, keeping every vital strictly private while detecting deteriorations in real time? That vision inspired VitalSense.
What it does
VitalSense transforms raw wearable sensor data into an immediate, actionable early-warning assessment:
- Continuous & Seamless Sync: An Apple Watch captures physiological metrics (heart rate, blood oxygen, respiratory rate, wrist temperature) into Apple Health, which automatically syncs to the iPhone without requiring an active watch companion app to run.
- Zero-Cloud, Edge Inference: The trained machine learning model is a tiny 17 KB Core ML graph packaged directly inside the app bundle. It performs inferences in milliseconds without accounts, network calls, or cloud servers — patient vitals never leave their own devices.
- Transparent Provenance & Manual Inputs: Real wearables cannot measure everything (e.g., blood pressure, supplemental oxygen, consciousness level). VitalSense explicitly labels the source of every metric (
sensor,manual, orassumed) and lets users adjust unmeasured vitals on an intuitive "Other vitals" sheet, re-scoring in real time. - Explainable Clinical AI: Rather than returning an opaque numerical probability, the app pairs the neural network's band prediction (
Normal,Low,Medium,High) with an embedded Swift implementation of the National Early Warning Score 2 (NEWS2). It breaks down which specific vital contributed to risk and provides plain-English explanations.
How we built it
- Machine Learning Pipeline: We trained a fastai tabular learner on clinical risk observation data[cite: 5, 6]. Alongside raw vitals, the model is fed NEWS2 aggregate and severity cut-points to assist the architecture.
- Export & Core ML Baking: Normalization statistics (means and standard deviations) were baked directly into the exported Core ML
.mlmodelgraph. This prevents normalization drift across languages and allows Swift to feed raw numbers directly into inference. - iOS & watchOS Native Architecture: Built in Swift using SwiftUI and HealthKit. Both the iOS client and watchOS companion carry the Core ML model independently, allowing the watch to score vitals standalone via an
HKWorkoutSessionat 1 Hz and sync across WatchConnectivity. - Automated Verification: We developed Python project generation scripts (
generate_xcodeproj.py) and a verification runner (check_model.sh) that compiles the Swift and Core ML models on macOS and verifies them against 24 frozen PyTorch prediction tensors to ensure parity within $10^{-4}$.
Challenges we ran into
- Wearable Sensor Limitations:
During development, the physical watch we had on hand was an older model rather than an Apple Series 9 or newer. Because of this, we were unable to test and stream vital sensor data (such as wrist temperature and continuous background oxygen readings) directly from hardware sensors on the wrist in real time. To overcome this limitation and ensure the system could still be thoroughly evaluated, we engineered an intuitive manual fallback workflow directly in the iPhone app—allowing users to enter, adjust, or override any vitals, which immediately recalculates and refreshes the Core ML risk score.
Apple Watches do not measure systolic blood pressure, cannot reliably measure SpO₂ on all hardware configurations, and estimate wrist temperature as relative nighttime deviations rather than absolute core temperature. We engineered a custom baseline temperature estimator (
estimatedBodyTemperature()) that computes 30-day baseline deviations clamped to $\pm 3^\circ\text{C}$ to prevent false hypothermia/fever scores. - Model-to-App Feature Alignment: Tabular ML models silently return erroneous outputs if Swift feature arrays misalign with Python training order. We resolved this by building an automated test harness that cross-evaluates Core ML against PyTorch outputs before deployment.
- Accessible Clinical Design: Standard medical alert colors often fail accessibility guidelines We decoupled meaning from color alone: every status band is paired with a distinct SF Symbol, written label, and numerical severity score.
Accomplishments that we're proud of
- Packaging a fully functional, high-accuracy clinical risk triage engine into a 17 KB Core ML graph that runs instantaneously on-device.
- Designing a true privacy-first healthcare architecture where no Protected Health Information (PHI) is ever transmitted, removing cloud attack surfaces[cite: 5, 6].
- Achieving rigorous test parity between our PyTorch model and the Swift Core ML implementation down to a precision of $6 \times 10^{-7}$.
- Building a multi-platform app with transparent provenance, acknowledging the realities of wearable hardware without compromising user trust.
What we learned
- How to map raw sensor intervals from Apple HealthKit into clinical frameworks like NEWS2 without skewing baseline inputs.
- Why baking preprocessing and normalization transformations directly inside exported model graphs is critical for production ML safety.
- The immense privacy, latency, and reliability advantages of running healthcare models directly at the edge rather than depending on server-side APIs.
What's next for VitalSense
- Medication Regimen Logging: Correlating the timing of newly started prescriptions directly with shifts in risk band scores to detect adverse drug events faster.
- Clinician Portal Export: Enabling patients to export cryptographically signed, encrypted FHIR/HL7-compliant PDF reports to share with their doctors during follow-ups.
- Longitudinal Trend Analytics: Introducing time-series predictive modeling to detect gradual downward baseline drifts before acute score spikes occur.
Log in or sign up for Devpost to join the conversation.