Inspiration
Nurses in British Columbia routinely work 12-hour shifts, often going five or six hours before their first real break. In a 2025 survey of 4,736 Canadian nurses, 67% said their workplace is regularly over capacity, and 65% said they fear repercussions for reporting safety concerns (CFNU 2025).
Everyone already knows nurses are overworked. What's missing is the record. Staffing changes when there's a paper trail: in Ontario, 109+ workload forms at one hospital led an independent committee to recommend 1:4 day and 1:5 night nurse-to-patient ratios (ONA). But after a 12-hour shift, the only evidence a nurse has is how tired they feel.
So we built WARD, the Workload And Rest Dashboard. It turns the watch a nurse already wears into a dated, unit-specific record of every shift, and it measures how many overloaded shifts never become a report.
What it does
- Nurses (Android app + website): once a day the phone uploads heart rate, steps and sleep from Health Connect. The nurse enters clock-in and clock-out, and WARD builds the shift: a physical-load chart, a stress indicator, breaks, watch-off gaps, sleep before the shift, and a plain-language card ("No break for 5 h 40 min"). A red shift drafts a workload report for the union's existing process. The nurse decides whether to send it.
- Managers: only anonymous, weekly unit data. That covers red-shift rates, a 12-week heatmap with unusual weeks flagged, side-by-side comparisons, an action log and a short AI-written summary.
- Joint union–hospital committee: the reporting gap (red shifts vs. reports actually sent) and an append-only log of who looked at what.
WARD documents workload. It is not a medical device and makes no diagnosis.
How we built it
Stack. A Kotlin + Jetpack Compose Android app reads Health Connect and uploads once a day with WorkManager. Two FastAPI services on Render (a nurse API and a manager API) sit over Tiger Cloud (TimescaleDB). The website is Next.js on Vercel, with Firebase email/password login. We have 45 backend tests passing.
Physical load. Step counts miss the hardest nursing work, like turning a patient or holding a limb. So we measure effort with percent heart rate reserve:
$$ \%\text{HRR} = \frac{HR - HR_{\text{rest}}}{HR_{\max} - HR_{\text{rest}}}, \qquad HR_{\max} = 208 - 0.7 \cdot \text{age} $$
Stress indicator. We estimate the heart rate that movement explains, and flag heart rate well above it while the nurse is stationary and below hard-effort levels:
$$ \widehat{HR}t = HR{\text{rest}} + \min!\big(70,\; 5 + 0.30 \cdot \overline{\text{steps}}_{3\,\text{min}}\big), \qquad r_t = HR_t - \widehat{HR}_t $$
A minute counts if $r_t$ is above our threshold, $\%\text{HRR} < 30\%$, and the nurse took fewer than 5 steps. This is currently a hand-set prior, not a trained model, and the app says so. The indicator never decides a shift's colour.
Banding. Each shift gets two bands anchored to published limits. Physical load is amber at $\geq 24.5\%$ HRR and red at $\geq 33\%$. Recovery is red at $\geq 300$ minutes without a confirmed break. The shift takes the worse of the two. Missing data is never green: a shift under 70% coverage turns grey.
Privacy by architecture.
- Raw heart rate is deleted when a shift is finalized, and unclaimed raw data after 7 days.
- The manager service holds a database role that can only read a published weekly schema.
- Cells with fewer than 5 nurses are suppressed, and so are cells whose nurse set differs from a neighbouring cell by fewer than 5 people. That blocks "compare Monday to Tuesday" differencing.
- Proportions are rounded to 10%, and counts get Laplace noise:
$$ \tilde{c} = c + \text{Lap}!\left(\tfrac{1}{\varepsilon}\right) $$
Anomaly detection. For each unit and shift type, we compare each day against a robust 28-day baseline:
$$ z = \frac{x - \text{median}{28}}{1.4826 \cdot \text{MAD}{28}} $$
We evaluated it blind on 50 randomly seeded synthetic cohorts, with the answer key hidden from the detector. It caught 54% of large, 36% of moderate and 20% of small planted anomalies, with 6.8 false alarms per 100 unit-days on cohorts with nothing planted. That false-alarm rate is the next thing we'd tune.
Gemini, with a leash. Gemini writes the weekly summary but never sees a number or a name. It gets placeholder keys and labels like "up / large", and our code fills in the values afterward. A validator rejects any draft with a digit, an unknown placeholder or cause-and-effect wording ("because", "due to") and falls back to a template. The manager page demos this rejection live.
Challenges we ran into
- "Standing still isn't resting." Our first design gated stress on step count, which would have labelled patient care as stress. Heart rate reserve fixed the physical side, and we accept that heart rate plus steps can't fully separate mental stress from stationary effort.
- Getting real watch data. Samsung Health passes data to Health Connect late, and slower still on our older test watch. We haven't yet processed a real-watch shift end to end, so the demo shift is a clearly labelled simulated day sent through the real ingest path.
- Live streaming wasn't feasible. We dropped it. Documentation doesn't need it, and uploading once a day means nothing streams during patient care.
- Anonymous isn't automatic. A manager who knows the roster can compare daily averages and isolate one nurse, so we released weekly data only, with suppression, rounding and noise. On small units, privacy noise is about the size of the signal.
- Missing data looks calm. Watches come off on the most chaotic shifts, so missing data shows as grey and as a red-rate range, never green.
- Gemini quota and data residency. We used up the free-tier quota during development, so summaries currently show our template fallback. The API also has no Canadian data-residency guarantee, which is what pushed us to number-free prompts.
- Watches aren't allowed everywhere. Fraser Health's hand-hygiene policy bans wrist watches for direct-care staff. The watch is our prototype; the intended product is an upper-arm band.
What we learned
- Overclaiming was our biggest bug. Almost every weakness came from claiming more than the data supported. Each fix narrowed a claim until the code could actually enforce it.
- Privacy is architecture, not a policy page. Who holds the database, what leaves it and when raw data is deleted matter more than any written promise.
- Documentation moves staffing, not dashboards. WARD's job is to feed processes nurses already have: the union's reporting process, joint safety committees and ratio oversight.
- Honest numbers beat impressive ones. We'd rather show a 6.8-per-100 false-alarm rate and a labelled untrained prior than a suspicious 0.95.
What's next
- Process a real-watch shift end to end, and train and test the expected-heart-rate model on open consumer wearable data.
- Move hosting to Canada and switch Gemini to a paid key.
- Run a single-unit pilot governed by BC's joint union–hospital ratio committee. Primary outcomes: the reporting gap, logged manager actions, and the share of shifts with no break.
Built With
- android
- claude
- claude-code
- fastapi
- firebase
- gemini
- google-ai-studios
- google-cloud
- health-connect
- jetpack-compose
- kotlin
- nextjs
- numpy
- pandas
- postgresql
- psycopg
- pytest
- python
- react
- render
- samsung-health
- tigerdb
- timescaledb
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.