Inspiration
In 2017, Zach Martin, a 16-year-old in Fort Myers, Florida, collapsed at an off-season football workout. His core temperature reached 107 °F. He died eleven days later. Florida's Zachary Martin Act (2020) came out of his death.
What makes it unbearable is how preventable it was. Heat stroke is close to 100% survivable when the athlete is cooled within 30 minutes, and it can be prevented before it starts. Yet about 9,000 high school athletes are treated for heat illness every year. Football's rate is ten times that of every other sport, and one in three of those cases happens with no athletic trainer present.
The way most programs manage this hasn't changed much: someone checks a weather app, then a coach decides how hard to push. A forecast is for the town, not the field. And it says nothing about a 125 kg lineman in full pads on day 2 of preseason versus a 95 kg linebacker on day 8, who sit in the same "Zone 2" but are in very different places on the way to the same line.
We asked: what if a coach could see each athlete's predicted core temperature through today's practice plan before practice starts, fix the plan, and then have the system keep watching and help if something still goes wrong?
What it does
HeatTwin is a per-athlete heat-strain simulation for high school sports. It works in three layers.
- Plan. The NWS forecast is turned into hourly on-field WBGT (Liljegren) and mapped to the FHSAA heat zone for each hour. A transient two-node thermoregulation model then predicts every athlete's core temperature, minute by minute, through the coach's practice plan. It takes body size, gear, drill intensity and acclimatization day into account. A constrained optimizer rewrites the plan (reorders drills, adds breaks, changes gear) so every athlete's predicted p95 stays under the planning line, while keeping as much training as possible and never breaking FHSAA rules.
- Watch. Live heart rate from straps athletes already own (Web Bluetooth, standard Heart Rate Service) recalibrates each athlete's model during practice and re-forecasts the rest of the session. When the re-forecast crosses the line, the coach gets a heads-up and a suggested change for that athlete only (at most two changes, in well under a second), with Apply and Undo.
- Respond. One-tap Collapse mode: a clock from the moment of collapse, a voice-guided cool-first protocol (Call 911, into the tub, stir the water, keep cooling, hand off to EMS), a live tub-water probe, and an EMS handoff timeline.
Plus a $30 sideline sensor node (ESP32/Arduino) that measures the field's real temperature, and the twin assimilates it to correct the forecast for the remaining hours.
Headline result (fixture forecast, 16 synthetic athletes, labelled as such in the app)
| Before | After Optimize | |
|---|---|---|
| Athletes over the 39.0 °C planning line (p95) | 16 of 16 | 0 of 16 |
| Hottest athlete's p95 | 41.65 °C | 38.98 °C |
| FHSAA rule issues | 2 | 0 |
| Training load kept | 72.9 % (12 changes) |
In our live heart-rate rehearsal, the app went from a heart-rate rise to a suggested fix in about 60 seconds.
Safety boundary: what it will never do
Estimated core temperature is for planning and early warning only. HeatTwin never diagnoses, never tells anyone an athlete is "safe," "fine" or "OK," never decides when to stop cooling, and never recommends medication. Rectal temperature is the only basis for treatment decisions (KSI / MHSAA guidance), and Collapse mode says so on screen. A language guard checks every generated sentence, and a semantic classifier backs it up. On held-out flagged text the rules alone caught 56 %, the classifier alone 96 %, and the two together 100 %, at a 6 % false-block rate. Blocked replies are held back rather than shown. That test text is synthetic, and we say so.
How we built it
- Physiology engine (Python, FastAPI, numpy/numba). We wrote our own transient two-node heat-balance model, vectorized across the roster, and cross-checked it against the JOS-3 reference from
pythermalcomfort. Weather comes from NWS, converted to WBGT with our own Liljegren implementation. - Optimizer. A search over candidate plans, with the simulation running inside the search loop, under hard FHSAA constraints. It is reproducible: it ends on iteration caps, not a wall clock, so the same inputs give the same plan on any machine.
- Calibration. Live heart rate is read against the plan's current drill and nudges each athlete's metabolic scale once a minute, with rest windows excluded and persistence gates so the system doesn't cry wolf.
- Web app (React 19, TypeScript, Vite). Live roster, practice-plan editor, athlete twin with body figure and forecast chart, response view, Collapse mode and a voice dock.
- Voice. A free, typed decision layer: a fastembed (ONNX) embedding classifier chooses intent, athlete, drill and intensity with calibrated probabilities, and abstains ("Did you mean…?") when unsure. The words you hear are always written by the engine, never freely generated.
- Hardware. An ESP32/Arduino field node streams JSON to the engine. It hot-plugs, and a reading more than 5 °C off the forecast is refused and labelled.
- Sourcing discipline. Every physiological, regulatory and physical constant lives in one file with a citation and a VERIFIED or TODO status, and a check fails the build if code uses a constant that isn't there. Ask "where does 39.0 come from?" and there's a one-click answer.
- Built with Claude Code, with sub-agents for physiology review, source checking and end-to-end demo QA.
Challenges we ran into
- Making the physics trustworthy. We reproduced Armstrong et al. 2010 (football-uniform heat study) in our model and in JOS-3. Our conservative planning mode matches control clothing (0.035 vs 0.037 ± 0.015 °C/min) and over-predicts the full-uniform rise (2.80 vs 2.37 ± 0.45 °C), which is the safe direction for a planning tool. JOS-3 and the ISO dynamic model under-predict both conditions by about 1 SD. Against field studies, our medians run 0.6–0.8 °C above the reported group-mean peaks, and we traced that mostly to how drill intensity is assigned.
- Forecast vs. reality. Our WBGT runs 2.0–4.4 °F above NWS's own WBGT layer over the demo window. We broke the gap down input by input and attribute most of the residual, by elimination, to NWS's unpublished clear-sky curve. We report it rather than tune it away.
- Not crying wolf. Standing around reads as "rest," not as the conditioning drill, so those windows are ignored. A reading that is implausible against the forecast is refused.
- A patent. The heart-rate → core-temperature Kalman method (Buller et al.) appears to be patented. Our core predictions come from the public two-node physics model. The HR filter exists only as an optional module, off by default and labelled "research mode — method appears patented."
- Honest labelling. Fixture forecasts, the synthetic roster, replayed heart-rate data and synthetic voice-test text are labelled as such in the app and in our results file. Nothing in that file is a number that code didn't compute from real, cited inputs.
Accomplishments we're proud of
- An optimizer that takes a 16-for-16 over-the-line plan to zero while keeping about 73 % of the training load.
- The whole loop on one screen: plan → optimize → watch → respond.
- A safety layer that is enforced in code and tested, not just stated.
- 430 engine tests and 65 web tests passing, plus an end-to-end walk of the demo script.
What we learned
A model that is right matters less than a model that is honest about how right it is. Putting the validation gaps, the labels and the sources into the product made it more convincing, not less. We also learned how much of "AI safety" in a health setting is language and boundaries, not just accuracy.
Limits, plainly
- This is a planning tool, not a medical device, and it has not been validated on adolescent athletes in the field. Our validation is against published adult studies.
- The roster, forecast and heart-rate replay in the demo are synthetic or fixture data, and the app labels them.
- Voice has not been tested with real speech. Its accuracy numbers are on synthetic text, and intensity classification is the weak spot.
- The live strap and Arduino flows depend on hardware we could not test in every real-world configuration.
What's next
- Real field validation with athletic trainers, using ingestible-pill or rectal data from real practices.
- Fitting per-athlete parameters over a season, not just within a session.
- Licensing the heart-rate → core-temperature filter for a production version.
- Manufacturing the $30 sideline node and a tub probe as a kit for schools without an athletic trainer on site.
Built With
- arduino
- esp32
- fastapi
- fastembed
- numba
- numpy
- nws-api
- onnx
- pydantic
- pytest
- python
- react
- typescript
- vite
- vitest
- web-bluetooth
Log in or sign up for Devpost to join the conversation.