Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for HeartTwin

What it does

HeartTwin is a 60-second daily check-in:

  1. 30-s selfie video: remote photoplethysmography (rPPG) extracts heart rate, HRV and breathing rate from tiny skin-colour changes.
  2. Voice task: a sustained "aaah" gives phonation time, harmonics-to-noise ratio (HNR) and jitter, which shift with pulmonary and laryngeal congestion.
  3. 3 symptom taps: breathlessness, swollen ankles, orthopnoea.

These signals feed a personal digital twin. When evidence of congestion persists, the care team gets a ranked dashboard entry with a one-line, non-diagnostic summary, e.g. "HR +12 bpm, phonation −4.8 s, weight only +1.6 kg – suggest call today & diuretic review."

How we built it

Digital twin. One latent congestion state \( c_t \) follows a random walk, and each standardised feature \( k \) is a noisy linear readout of it:

$$c_t = c_{t-1} + w_t, \qquad w_t \sim \mathcal{N}(0, q)$$

$$z_{k,t} = s_k \, \frac{y_{k,t} - \mu_k}{\sigma_k} = b_k \, c_t + v_{k,t}, \qquad v_{k,t} \sim \mathcal{N}(0, 1)$$

Here \( \mu_k, \sigma_k \) come from the patient's own first 7 days after discharge, \( s_k \) is the physiological direction of change (e.g. HR ↑, HNR ↓), and \( b_k \) weights each feature by its signal-to-noise ratio. A Kalman filter updates the posterior daily and handles missed check-ins naturally. An alert fires when \( \hat{c}_t / \sqrt{P_t} > \tau \) for 2 consecutive days.

Browser prototype (HTML/JS, no install). Camera rPPG with the POS algorithm (Wang et al., 2017), autocorrelation pitch and HNR (Boersma), the Kalman twin, and a care-team dashboard. All raw video and audio are processed on-device.

Verified signal processing. Unit tests on ground-truth synthetic signals: heart-rate error ≤ 0.5 bpm (62–115 bpm), breathing ≤ 0.1 /min, pitch ≤ 0.5 Hz.

In-silico study (Python). 1,000 virtual post-discharge patients over 90 days, with literature-based effect sizes, 85 % adherence, rPPG artefacts, confounders (colds, exertion) and 25 % rapid-onset decompensations. Thresholds were tuned on 500 patients and evaluated on the other 500 held-out patients, at a false-alarm budget matched to the weight rule.

Detector Events flagged early Median lead time False alarms / patient-year
Weight rule (+2 kg / 3 d) 42 % 2 d 0.6
Population thresholds 66 % 5 d 1.0
Video only 54 % 4 d 0.9
Voice only 75 % 5 d 0.6
HeartTwin fusion 92 % (95 % CI 87–97) 5 d 0.8

Challenges we ran into

  • rPPG is fragile under motion and poor light. We added a signal-quality (SNR) gate and robust clipping inside the twin.
  • Octave errors in pitch tracking for higher voices, fixed with a "first autocorrelation peak within 85 % of the maximum" rule.
  • No public voice dataset of HF patients exists, so we built a transparent, reproducible simulation with a conservative sensitivity analysis.
  • Fairness: rPPG accuracy is known to drop on darker skin tones, so testing across Fitzpatrick I–VI is a mandatory next step.

What we learned

  • Personalisation beats any single sensor. The personal baseline moved detection from 66 % to 92 % at a similar alarm burden.
  • Fusion matters. Every single modality underperformed the combined twin.
  • Honest limits (lighting, skin tone, adherence) have to be designed in from day one, not patched later.

What's next

  1. Validate rPPG on UBFC-rPPG / PURE and across skin tones.
  2. Native Flutter app with MediaPipe face mesh.
  3. Prospective pilot with a cardiology ward, targeting ≥ 20 % fewer 30-day readmissions.
  4. Extend the same engine to COPD, dialysis fluid status and post-surgical recovery.

Research prototype, not a medical device. Key references: Savarese 2023 (Cardiovasc Res), Khan 2021 (Circ Heart Fail), Chaudhry 2007 & Zile 2008 (Circulation), Murton 2017 (JASA), Amir 2022 (JACC HF), Wang 2017 (IEEE TBME), Koehler 2018 & Abraham 2011 (Lancet). Full list with DOIs in the attached PDF.

With every effect size halved, fusion still flags 68 % vs 9 %. These are simulation results, not clinical evidence. They are the hypothesis we want to test in a pilot.

Share this project:

Updates

Submission history