GitHub: https://github.com/vshnu1/Relay
Link: https://relay-bayhacks.onrender.com/welcome/
LOGINS TO USE for testing : DOCTOR: [email protected] PASSWORD: DemoBuilder DEMO PATIENTS: [email protected] PASSWORD: DemoBuilder PATIENT 2: [email protected] PASSWORD: RelayPatient-2026
Elevator pitch (one line)
Relay watches the thirty days after a hospital discharge. It compares each patient's wearable readings with their own baseline, asks them what was happening when several move together, and hands the care team one reviewable, auditable evidence packet. A clinician decides.
Inspiration
The thirty days after a discharge are when a patient is most likely to be readmitted, and they are also when the line to the care team goes quiet. The patient does not know which change is worth a phone call. The clinician has forty patients and cannot watch any of them continuously. Meanwhile the watch on that patient's wrist has been recording the whole time, and nobody reads it, because a single reading out of context means nothing.
Medicare penalises six conditions on a thirty-day readmission window. Four of them — heart failure, pneumonia, COPD, and hip and knee replacement — are already recovery pathways in Relay. So is oncology: the post-chemotherapy pathway is a neutropenic-fever watch, and two patients in the demo roster were discharged from Moffitt's Malignant Hematology unit.
Who it is for. The buyer is the discharging hospital, which already carries the readmission penalty and already has a reimbursement route for this work: the CMS remote physiological monitoring codes (99453, 99454, 99457). The patient needs no new hardware, and the clinician gets two patients to look at instead of twenty-eight charts.
What it does
Turns a continuous stream into an evidence packet.
- Collects. Reads what the patient's existing wearable already records, plus an Apple Health export that is inflated and parsed entirely in the browser. No new hardware.
- Normalises. Cuts every signal into six-hour windows across a thirty-day monitoring window.
- Compares to the person, not the population. Each window is measured against that patient's own median and median absolute deviation, \(z = (x - \tilde{x}) / (1.4826 \cdot \mathrm{MAD})\). We measured why this matters: across the 66 subjects with usable baselines, resting heart rate spans 31 beats between people while one person varies about 2.4 day to day. A single population threshold cannot serve both.
- Requires coordination and persistence. For pneumonia, three of four counted signals must each stay past threshold for 24 hours with 24 hours of overlap. One odd reading does nothing.
- Gives the patient one next step. Not a number to worry about: one sentence written from their own readings, naming the devices it used, that leads to a check-in or back to their discharge letter. It never says what is wrong.
- Asks, rather than alarms. An ElevenLabs conversational agent runs a check-in whose questions are written for the discharge pathway: heart failure asks about swelling and breathing when lying flat, pulmonary embolism about one swollen leg, sleep apnoea about nights without the CPAP machine. Nothing reaches the care team until the patient reviews and approves it.
- Hands over one packet. The clinician gets the pattern, the patient's own words, a cohort comparison, a mock FHIR bundle, and the follow-up options.
Each recovery pathway watches different signals. Fifteen pathways. Sleep is counted for a COPD flare-up and watched for context in pneumonia. Weight counts for heart failure. The app says which, on screen, for every reading.
It says when it cannot see. Four states: review recommended, context needed, monitoring, and not enough data. Most tools cannot distinguish a quiet patient from one nobody can see.
How we built it
Stack. React and Vite on the front end, Express on the API, a Python analysis package, all deployed on Render.
Render Workflows. Render Workflow orchestrates the analysis: it runs the machine learning model and our evidence checks, then returns traceable results that help focus the patient's check-in and inform the clinician's review. The pipeline is six tasks — ingest and normalise, calculate baseline, detect coordinated deviation, request context, score with the Python model, compile the review item. It runs when a patient starts a check-in, when their answers are submitted, and when a clinician analyses or simulates a case. It does not run per wearable reading or per voice turn, so the live conversation stays fast. Every run returns an execution ID, and any number on the clinician's screen traces back to the run that produced it. If the Workflow is unavailable the app uses a clearly labelled fallback, and the interface never calls that a Workflow run.
ElevenLabs, in two places. A conversational agent runs the patient's check-in, with one client tool that drafts structured answers the patient then corrects and approves. And a spoken clinician briefing, text to speech, for a clinician between rooms.
The machine learning model. An Isolation Forest, 200 trees, unsupervised by design. No public wearable dataset follows people through a post-discharge decline, so there are no labels to train a classifier on, and we built a model that does not need them rather than inventing some. Per six-hour window it builds a feature row, eleven features for the pneumonia pathway — the robust deviation of each metric from that patient's own median, plus coverage, count of fresh signals, count of deviating core signals, and mean absolute deviation across core signals. It fits to that patient's own history with a chronological train/validation/test split, never a random one, so it cannot peek at the future. The threshold is the 95th percentile of that patient's own validation scores, and scores are mapped through a logistic centred on it, so 0.5 means exactly at the line. Where a patient has too little history it falls back to a per-pathway prior trained on 40 synthetic patients, labelled synthetic in the artefact's own metadata. Isolation Forest gives no feature importance, so explanations come from the robust deviations themselves rather than from the model.
The restraint ledger. The model's verdict passes four gates: signals moving together, persistence across windows, the anomaly score, and data quality. Each is displayed with what was required and what was observed. The clinician can see why Relay stayed quiet as easily as why it spoke.
The clinical boundary, enforced at runtime. Relay never diagnoses, predicts risk, grades severity, recommends treatment, or escalates by itself. Every generated sentence passes a language guard against five rule families before it leaves the server. A sentence that trips one is replaced with a description of the readings and the redaction is recorded and shown. The guard exists twice, in Python and in Node, and a test fails if the two disagree phrase for phrase, because a boundary enforced in one language and not the other is not enforced.
Security and HIPAA. There is a Security page in the product rendering all ten technical safeguards of the HIPAA Security Rule, 45 CFR 164.312, cite by cite. All ten are built, each with what it was before and what it still does not do. Two are read from the running process rather than asserted: whether the data is really encrypted, and whether the audit log verifies.
- Identity. Named accounts with scrypt-hashed passwords and revocable server-side sessions that end after fifteen minutes idle.
- Least privilege. A clinician account is for one care team and is refused every other team's records, on reads and on writes. Emergency access lifts that for one record, for fifteen minutes, with a written reason.
- Encryption by default. AES-256-GCM over state, the account store, the patient record and every audit line, held as encrypted JSON on a Render persistent disk. TLS and strict transport security in transit.
- An audit trail that is evidence, not a list. Every access names the person, not only the role, carried through the model's asynchronous work by an AsyncLocalStorage so one request cannot be attributed to another. Each entry carries the digest of the one before it, so an edited or removed line breaks the chain and the verifier names where.
- No shadow AI. The Apple Health import, 284 MB, is parsed entirely in the browser: nothing uploaded, no third-party model sees it. The content security policy names no third-party script origin at all, and typefaces are self-hosted, so not even a font request tells anyone who opened a record.
- Vendors held to the same bar. The voice check-in is the only path that leaves our origin, and the server refuses it for any record outside the synthetic roster, because neither vendor's current plan includes a business associate agreement. We treated AI procurement as security procurement.
- The patient's side of it. A patient can save a complete copy of their own record (45 CFR 164.524), and taking one is itself recorded.
Challenges
Calibrating without labelled outcomes. Nobody publishes wearable data from people declining after a discharge, so we measured what can be measured and kept the numbers separate. We injected a coordinated change of known size into 41 real subjects' own recorded data: at two personal standard deviations Relay sees it the next day, against a median of five days for that same person's ordinary variation. In a thirty-day readmission window that head start is the point. The false-alarm rate at the shipped 1.75 sd is 3.35% of subject-days against a whole-series baseline and 6.04% judged the way a deployment works, each day only against the days before it. We quote the stricter one. Sensitivity against real post-discharge decline needs a prospective study, and that is first on the roadmap.
Restraint is harder than detection. Alarm fatigue is the failure mode of monitoring, so the engineering went into coordination and persistence gates and into showing why Relay stayed quiet. In testing, the model's list of ordinary night-to-night variation could outrank what the patient's own readings pointed to when choosing check-in questions, so we fixed the questions per recovery pathway and let the model inform the priority, not the content.
Security that means something. Emergency access originally lifted no restriction, because every clinician could already reach every record. We marked it partial on our own page, then built care-team scoping underneath it and an integration suite that proves a record is refused, opened with a reason, and logged.
Regulatory status. On our own reading, Relay is probably device software. The Clinical Decision Support exemption in FD&C Act 520(o)(1)(E) requires all four criteria, and we do not meet the first: it excludes software that analyses a pattern from a signal acquisition system, which is exactly what reading a wearable stream and requiring 24 hours of persistence is. Knowing that on day one is what makes the path to clinical use a plan: the boundaries (never diagnoses, never scores risk, a clinician decides) and the roadmap are built around it.
Accomplishments
- A compliance page inside the product, not a markdown file: all ten technical safeguards built, two of them verified live from the running process, and what an adopting organisation adds listed under "Before clinical use".
- A model that shows why it stayed quiet, gate by gate.
- A measured, published difference between an in-sample and a deployed false alarm rate, and a measured detection lag.
- Two real applications on one record: a clinician workspace and a patient app built for readers in their fifties and up, with secure messaging between them that survives sign-out, restart and deploy.
- 103 JavaScript tests, 63 Python tests, and four HTTP integration suites, three of which run the server in the exact configuration the deployment runs.
What we learned
That the hardest part of a clinical product is not detection, it is restraint, and that a system which cannot say "I do not have enough data" will eventually tell somebody they are fine when nobody was looking.
What's next
A clinician to own the threshold table, which the code says in capitals is illustrative and not clinically validated, and then a prospective pilot with one discharge unit, reported against TRIPOD+AI and DECIDE-AI. Business associate agreements with every processor, a formal risk analysis, and a second authentication factor. FHIR conformance: LOINC codes, UCUM units, and a Device or Provenance resource saying these readings are wearable-derived. Billing-readiness counters for the CMS remote monitoring codes, which are derivable from the event log that already exists. A deletion endpoint for Washington's My Health My Data Act. And a live hardware feed, which is one line to change: the app reads from a source contract, and the simulated source is only today's implementation of it.
Where every number comes from
| Claim | Source |
|---|---|
| 31 beats between-person spread, 2.4 within-person | fixtures/calibration.json, 66 subjects with usable baselines |
| 3.35% and 6.04% false alarms | fixtures/calibration.json, 4,420 subject-days |
| Detection lag: next day at 2 personal sd, against five days for noise | fixtures/sensitivity.json |
| Reference cohort | LifeSnaps, public research data: 71 subjects, 4,453 subject-days |
| Engineering validation | One team member's own Apple Health export, 284 MB, de-identified under a secret not in the repository |
| Demo patients | 28 synthetic, 15 pathways, 13 hospital units. No patient data, ever |
| Model priors | Synthetic only, 40 synthetic patients per pathway, split 24/8/8, labelled in the artefact metadata |
| Tests | npm test (103), PYTHONPATH=ml python3 -m unittest discover -s ml/tests (63), and node tests/{api,accountsMode,messaging,careTeam}.integration.js |
| Ten safeguards built, care-team scoping, emergency access | The Security page in the product; docs/COMPLIANCE.md; tests/careTeam.integration.js |
| A patient's copy of their record | src/recovery/model/recordExport.js, tests/recordExport.test.js |
| Check-in questions per pathway | CHECKIN_QUESTIONS in src/recovery/model/profiles.js, tests/checkinPlan.test.js |
| Reimbursement codes | docs/COMPLIANCE.md, CMS remote physiological monitoring |
| Thresholds | src/recovery/model/profiles.js, which states in capitals that they are illustrative and not clinically validated |
Built With
- css
- elevenlabs
- express.js
- github
- html
- javascript
- node.js
- numpy
- python
- react
- render
- scikit-learn
- vite
Log in or sign up for Devpost to join the conversation.