Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for HeartTwin
What it does
HeartTwin is a 60-second daily check-in:
- 30-s selfie video: remote photoplethysmography (rPPG) extracts heart rate, HRV and breathing rate from tiny skin-colour changes.
- Voice task: a sustained "aaah" gives phonation time, harmonics-to-noise ratio (HNR) and jitter, which shift with pulmonary and laryngeal congestion.
- 3 symptom taps: breathlessness, swollen ankles, orthopnoea.
These signals feed a personal digital twin. When evidence of congestion persists, the care team gets a ranked dashboard entry with a one-line, non-diagnostic summary, e.g. "HR +12 bpm, phonation −4.8 s, weight only +1.6 kg – suggest call today & diuretic review."
How we built it
Digital twin. One latent congestion state \( c_t \) follows a random walk, and each standardised feature \( k \) is a noisy linear readout of it:
$$c_t = c_{t-1} + w_t, \qquad w_t \sim \mathcal{N}(0, q)$$
$$z_{k,t} = s_k \, \frac{y_{k,t} - \mu_k}{\sigma_k} = b_k \, c_t + v_{k,t}, \qquad v_{k,t} \sim \mathcal{N}(0, 1)$$
Here \( \mu_k, \sigma_k \) come from the patient's own first 7 days after discharge, \( s_k \) is the physiological direction of change (e.g. HR ↑, HNR ↓), and \( b_k \) weights each feature by its signal-to-noise ratio. A Kalman filter updates the posterior daily and handles missed check-ins naturally. An alert fires when \( \hat{c}_t / \sqrt{P_t} > \tau \) for 2 consecutive days.
Browser prototype (HTML/JS, no install). Camera rPPG with the POS algorithm (Wang et al., 2017), autocorrelation pitch and HNR (Boersma), the Kalman twin, and a care-team dashboard. All raw video and audio are processed on-device.
Verified signal processing. Unit tests on ground-truth synthetic signals: heart-rate error ≤ 0.5 bpm (62–115 bpm), breathing ≤ 0.1 /min, pitch ≤ 0.5 Hz.
In-silico study (Python). 1,000 virtual post-discharge patients over 90 days, with literature-based effect sizes, 85 % adherence, rPPG artefacts, confounders (colds, exertion) and 25 % rapid-onset decompensations. Thresholds were tuned on 500 patients and evaluated on the other 500 held-out patients, at a false-alarm budget matched to the weight rule.
| Detector | Events flagged early | Median lead time | False alarms / patient-year |
|---|---|---|---|
| Weight rule (+2 kg / 3 d) | 42 % | 2 d | 0.6 |
| Population thresholds | 66 % | 5 d | 1.0 |
| Video only | 54 % | 4 d | 0.9 |
| Voice only | 75 % | 5 d | 0.6 |
| HeartTwin fusion | 92 % (95 % CI 87–97) | 5 d | 0.8 |
Challenges we ran into
- rPPG is fragile under motion and poor light. We added a signal-quality (SNR) gate and robust clipping inside the twin.
- Octave errors in pitch tracking for higher voices, fixed with a "first autocorrelation peak within 85 % of the maximum" rule.
- No public voice dataset of HF patients exists, so we built a transparent, reproducible simulation with a conservative sensitivity analysis.
- Fairness: rPPG accuracy is known to drop on darker skin tones, so testing across Fitzpatrick I–VI is a mandatory next step.
What we learned
- Personalisation beats any single sensor. The personal baseline moved detection from 66 % to 92 % at a similar alarm burden.
- Fusion matters. Every single modality underperformed the combined twin.
- Honest limits (lighting, skin tone, adherence) have to be designed in from day one, not patched later.
What's next
- Validate rPPG on UBFC-rPPG / PURE and across skin tones.
- Native Flutter app with MediaPipe face mesh.
- Prospective pilot with a cardiology ward, targeting ≥ 20 % fewer 30-day readmissions.
- Extend the same engine to COPD, dialysis fluid status and post-surgical recovery.
Research prototype, not a medical device. Key references: Savarese 2023 (Cardiovasc Res), Khan 2021 (Circ Heart Fail), Chaudhry 2007 & Zile 2008 (Circulation), Murton 2017 (JASA), Amir 2022 (JACC HF), Wang 2017 (IEEE TBME), Koehler 2018 & Abraham 2011 (Lancet). Full list with DOIs in the attached PDF.
With every effect size halved, fusion still flags 68 % vs 9 %. These are simulation results, not clinical evidence. They are the hypothesis we want to test in a pilot.
Log in or sign up for Devpost to join the conversation.