Inspiration
Athletes hide how tired or hurt they are, because they're afraid of losing playing time. Coaches and athletic trainers rarely have objective movement data, so the usual "I'm fine, just over trained today" often goes unchecked. We wanted a cheap way to compare what an athlete's body shows with what the athlete says, using only a phone and to communicate it between coaches and players.
What it does
BreakPoint takes a side-view video of bodyweight squats and analyzes the fatigue signal.
- A pose model track the athlete's hip, knee, and ankle in every frame.
- We split the video into reps and measure three things for each one: tempo, squat depth, and ascent speed.
- Each rep is compared to the athlete's own first three reps, to ensure that the comparison is consistent to their physical build. The drift becomes the Rep Fatigue Index (RFI), from 0 to 100, which is measured through analyzing rep frequency, depth, and rise angles.
- After the set, the athlete reports effort (1–10) and any pain. Fixed rules compare the report with the measured fatigue and flag possible under- or over-reporting.
- Gemini explains the result in plain language for the athlete, coach, and trainer. Coaches and trainers get a dashboard with trends, flags, and feedback.
BreakPoint is a screening aid. Hence, it never diagnoses or says that someone is "safe," and any reported pain is always routed to a human.
How we built it
The math. Everything is divided by the athlete's leg length \(L\), so results don't depend on camera distance or body size. Each rep is compared with a baseline, the median of the athlete's first three clean reps. We turn each kind of drift into a score between 0 and 1 using
$$ \mathrm{clip}(x) = \frac{\min(0.30, \max(0, x))}{0.30} $$
which gives three scores: \(T\) for tempo, \(D\) for depth, and \(S\) for ascent speed. Each one only counts drift in the tired direction (slower, shallower, or slower rise). For each rep, the raw fatigue score is
$$ R = 100 \left( 0.35\,T + 0.30\,D + 0.35\,S \right) $$
The Rep Fatigue Index (RFI) is the rolling median of \(R\) over 3 reps. The mismatch check compares the athlete's reported effort with the expected effort
$$ \widehat{\mathrm{RPE}} = \frac{\mathrm{RFI}}{10} $$
and flags a gap of 3 or more.
Choosing the pose model. We didn't pick one by reputation and instead hand-labeled joint positions and rep timings on our squat video and split it into alternating 10-second dev and test blocks. Tuning used dev only, and rankings used test only. We scored 13 models from the YOLO11, MediaPipe, RTMPose, and ViTPose families on keypoint accuracy, rep detection, stability, robustness (blur, low light, JPEG compression, tilt, half resolution), and speed. Every model also had to pass a deployment gate of at least 10 FPS on a plain CPU.
- RTMPose-x was the most accurate but unfortunately never passed the CPU gate, so it's our accuracy reference.
- MediaPipe scored well, however it crashed on our demo machine.
- YOLO11n-pose tied for the best rep detection, ran at about 37 FPS on CPU in our benchmark, is only 6.3 MB, and was the most robust of the models that passed.
Its joint positions are less precise, but we measure the hip's change over time against the athlete's own baseline and smooth the signal with a Savitzky-Golay filter, and so small constant errors cancel out due to the data consistency.
The system.
- A React Native / Expo app for athletes, coaches, and trainers.
- A FastAPI server on a laptop runs the pose model, the analysis, and video rendering. The phone uploads the video and polls a background job for real progress.
- OpenCV + FFmpeg draw an annotated H.264 video with the skeleton, live metrics, and form cues.
- A Cloudflare tunnel lets the phone reach the laptop on campus Wi-Fi that blocks device-to-device traffic.
- Gemini writes the explanations. SQLite stores summary numbers for the weekly dashboard.
Rules decide, AI explains. The safety flag comes from deterministic code, because an LLM can be inconsistent or talked out of a flag. Gemini only writes the words. We also gave it structured JSON output, a banned-wording filter, and protection against instructions hidden in an athlete's comment.
Challenges we ran into
- Expo Go wouldn't open until we used the right team account. My teammate created the Expo project under her team account, so Expo Go on my phone wouldn't launch the app until I signed in with that account. Our frontend and backend were built on separate accounts and machines, so this blocked testing on a real phone until we sorted out whose account owned the project. After that, we settled on a repeatable start-up order: API server first, then the tunnel, then Expo, then scan the QR code.
- The phone couldn't reach the laptop. The phone and laptop weren't on the same network, and network firewalls (campus and guest Wi-Fi often block devices from talking to each other) stopped them from connecting directly. We solved it with a Cloudflare tunnel, which gives the laptop's API a public HTTPS address the phone can reach from any network. We also used an Expo tunnel so the phone could load the app itself.
- A model that crashed. Our best deployable model, MediaPipe, crashed in its GPU helper on the demo Mac. We re-ranked on what actually runs and shipped YOLO11n-pose.
- Pauses are ambiguous. A pause can be planned or a sign of struggle. A standing rest once pushed a clip's RFI to 42. We now score a pause by where it happens and whether it grows compared with the first reps. The same clip scores 19.
- Knee angles lie from the front. On one front-view clip, nearly identical depths came with 147° and 33° angles. We only show and use angles that pass a side-view check and a geometry check.
- The LLM crossed a line. Gemini once wrote that it was "clearing" an athlete. We now filter every output for clearance and diagnosis language, and fall back to a safety-only report if Gemini is down.
Accomplishments we're proud of
- Developing a measured, reproducible computer vision model choice.
- A safety design where rules make the call, AI explains it, and humans decide.
- A full pipeline from phone video to coach dashboard, with a one-tap offline demo.
What we learned
- Just because a model is accurate doesn't automatically make it the right one. Instead, rep timing, not raw keypoint accuracy, is what our pipeline depends on.
- Judging each athlete against their own baseline is simpler and fairer than a universal standard.
- Keep decisions deterministic and use LLMs for communication only.
Limitations
This is an early prototype. RFI weights and traffic-light cutoffs (Healthy < 35, Caution 35–59, Fatigued ≥ 60) are provisional and not clinically validated. Squat variations share the same calibration. Real use with minors needs consent and a data-handling policy.
What's next
Calibrate on many athletes and squat variations, validate against trainer judgment, add authentication and a proper database, and run pose inference on the phone so video never leaves the device.
Built With
- claude
- cloudflare-tunnel
- expo.io
- fastapi
- faster-whisper
- ffmpeg
- google-gemini-api
- hipergator
- mediapipe
- numpy
- opencv
- python
- pytorch
- react-native
- rtmpose
- scipy
- sqlite
- typescript
- ultralytics-yolo11
- uvicorn
- vitpose
- zustand
Log in or sign up for Devpost to join the conversation.