BodyQuest — Physical AI Coach

Inspiration

The origin of BodyQuest was simple frustration: while most fitness apps are able to keep track of the number of repetitions you do, none of them can watch you and point out problems with your technique as a genuine coach would. Since incorrect squat form is one of the main reasons for injuries in the gym, the kind of feedback that helps to prevent them — for example, "your knees are collapsing in", "you're not going deep enough", or "slow down and control the movement" — is generally only given by a personal trainer who is standing beside you. That's why we wanted to bring this kind of real-time, corrective coaching into everyone's pocket, working entirely on the device itself and with no need for a connection to the cloud or for any camera footage to leave the phone.

What it does

BodyQuest uses the camera on a phone to provide live form coaching and operates via a closed feedback loop.

$$ \text{SEE} \rightarrow \text{UNDERSTAND} \rightarrow \text{ANALYZE} \rightarrow \text{CORRECT} \rightarrow \text{VERIFY} \rightarrow \text{LEARN} \rightarrow \text{ADAPT} $$

For each rep of a squat, the app:

  1. It is able to detect and track 33 body landmarks in each frame using pose estimation performed on the device.
  2. Understands by converting the raw landmark movements into a rep state machine (this involving a descent, then reaching the bottom, followed by an ascent and ending with a lockout).
  3. It analyzes by converting the joint geometry into definite per-metric scores regarding knee alignment, torso lean, squat depth, tempo, and left/right symmetry.
  4. Makes corrections — if any metric falls below the threshold it gives exactly one spoken correction (it never provides a list, so the athlete doesn't get overwhelmed), with the corrections prioritised according to injury risk.
  5. Verifies by checking the other party's snapshot for the same metric to ensure that the correction had indeed been helpful.
  6. It learns/adapts by keeping a record of corrections and preparing a summary (including the number of repetitions, the trend in technique, and the quality over time) intended to influence subsequent sessions.

One of the fundamental principles we adhered to throughout was that form scores are deterministic biomechanics, calculated based on joint angles and never invented by a model. The only role of any AI is to determine how a correction is phrased when spoken aloud, not to decide whether or how badly something is wrong. This constitutes an important trust boundary in any system providing physical safety feedback.

How we built it

  • For pose tracking, an on-device pose landmarker model (using MediaPipe Tasks Vision) is executed via a GPU delegate, with a automatic CPU fallback available for devices and emulators that do not support it.
  • In the form analysis, pure geometry is used—the joint angles and ratios are calculated directly from the landmark stream and there is no learned model involved in the scoring process.
  • The correction engine is a small state machine which at any one time handles at most a single pending correction and checks it against the following rep, thereby ensuring that the feedback takes the form of a conversation rather than a series of warnings. The voice used for text-to-speech directly speaks the latest correction and interrupts the old audio so that the athlete always hears what their form currently is.
  • For the coaching feedback, we've made the language used in each correction adjustable — starting with a standard, manually written phrase table, and providing a provision for using a small local language model (we have been incorporating a compact on-device Gemma model, executed locally via the same MediaPipe Tasks family that is already used for pose estimation) so as to vary and soften the delivery, even though the core FormIssue and the score it refers to remain entirely unchanged. – The architecture is such that all the components are designed to operate entirely on the device, including the pose model, the geometry scoring, the TTS, and (if desired) the phrasing model, so that a workout can be coached without needing a network connection and with no video leaving the phone.

Challenges we ran into

  • When it comes to real-time constraints on actual hardware as opposed to emulators, GPU delegates for on-device inference generally do not function with emulator software rendering, which is why we were obliged to create an automatic CPU fallback and had to carefully test on real devices in order to obtain accurate latency figures.

Built With

  • camerax
  • kotlin
  • mediapipe
  • room
Share this project:

Updates

Submission history