Inspiration

Concussion diagnosis is almost exclusively made through information obtained from the patient, which is another issue, considering that the patients who will likely suffer from such a condition are the very ones who will be most eager to claim that they're just fine. They'll want to return to sports, go back to classes, to their daily lives. And eye movement can't be faked.

Half of concussion patients will develop some sort of oculomotor/ vestibular dysfunction; and a clinical test is available right now to detect it. The VOMS, or Vestibular/Ocular Motor Screening, consists of seven tests, and requires zero equipment.

There was another gap I kept noticing. Almost every single concussion application available right now targets returning athletes. The larger group that is not adequately served by those apps is students returning to class, and there's no equivalent to the six stage exertion test for academic workload.

And one last point influenced the entire design process of the app. Light and motion sensitivity are among the symptoms of concussion; an app meant to check for those symptoms should not include flashing animations.

What it does

The Baseline uses a VOMS-type test through the browser and compares its findings to your baseline scores.

  1. Prepare. Pre-task symptom assessment, as provocation means a deviation from your own starting point and not an absolute value.
  2. Screen. Seven tests with visual stimuli on the screen: smooth pursuits, horizontal/vertical saccades, near point of convergence, horizontal/vertical VOR, and visual motion sensitivity. You will have to assess 4 symptoms for each task.
  3. Findings. What provokes symptoms? What did the camera detect? The Return to Sport and Return to Learn phases based on Amsterdam 2023 consensus guidelines, red flags, which require going straight to the emergency department, and a printout for a clinician.

It measures things, and that is why it elevates itself from being a simple test of multiple choice answers. Near point of convergence comes up in centimeters, which is possible because the diameter of human iris is constant at 11.7 mm in all adults and can be used as a fixed reference for absolute distance measurement. Five of seven tasks give a pursuit gain, a phase lag, and vestibulo-ocular reflex gain.

The thing records your data. "Baseline" is not only a name of the tool but also a medical term because you record yourself on the day you feel fine and use all future screenings in correlation with your own and population statistics. It is much more significant than it sounds because normal variation in near point of convergence between people makes 4.5 cm divergence insignificant result for a patient whose healthy near point is 2 cm, but it may mean something else for another one. Eventually, the record becomes a recovery curve.

No information ever leaves your device, since there is no account, no upload and no server. The record is stored in your browser along with export and erase buttons.

How we built it

It uses React 19 with TypeScript, Vite and Tailwind, and it's a complete static website.

Eye tracking is powered by MediaPipe FaceLandmarker running in WebAssembly: 478 landmarks, 52 blendshapes, and a facial transformation matrix providing head pose. It's the latter which enables VOR testing to be performed at all.

Privacy in this case is an architectural choice. I have vendored the 11.7MB WASM runtime and the 3.7MB model into the repository, and they are served from the app's origin, therefore no runtime requests are being made to Google or anybody else. This is the most important aspect of any privacy-preserving on-device claim that people usually gloss over. The inference happens locally, yet the model is loaded from a CDN, which means that CDN knows exactly who is using this product and when.

The clinical logic consists of pure functions residing in src/lib/clinical/: PCSS-22 scoring, VOMS provocation, convergence geometry, Amsterdam staging, personal baseline comparison, and oculomotor mathematics. 116 tests cover that codebase, and the components are just thin wrappers around it.

Design-wise, there are two rooms. The landing area is loud; a didone typeface appears on blush-colored paper. It is grainy, within an editorial grid. Stepping into the device leads to a declared shift to a dark, quiet, low stimulus condition; no grain, no motion, low luminance, and the app verbally tells you what it’s doing and why. You can skip this phase. This is the entire thesis made tangible.

Challenges we ran into

There's an almost-falsehood in the app. On a test where zero tasks provoke symptoms and only the convergence was abnormal, Findings stated "Several tasks provoked symptoms." The logic for the rationale string was hardcoded plural and could be hit with a provoked count of zero. This was not visible in the code and entirely obvious in a screenshot. It is precisely the sort of error that this entire product is designed to prevent, where we did not measure it, we could not solve it, and we did measure it and find nothing collapse into one sentence. Now both the heading and the rationale come from the two counts separately, held in place by regression tests.

One of my tests broke, and my test was correct. I had asserted that someone who has always broken at 5.5 cm should not be told that he or she is now breaking at 5.6 cm. My first implementation changed directions when the delta was non-zero, reporting any non-zero webcam noise as deterioration to a terrified person regarding his or her brain. There is now a dead band around every comparison, and the trend refuses to give a direction with fewer than three screenings.

The clinician printout was completely black on black. This is the only one of the outputs that was never actually rendered. The rooms were painted with literal colors rather than the design tokens, which meant that the overrides in the print stylesheet never applied to them, and thus everything was rendered in nearly black text on nearly black background. Now every room interprets the design tokens, and the PDF is automatically generated via a script, so it cannot remain unseen anymore.

The camera could not be started turned out to be five distinct issues. The lack of a proper catch statement meant that a denied permission, a camera already being used by another application, a device with no camera whatsoever, and a model that fails to load all produce an equally useless sentence, for the single person who is able to solve them. Now there is a specific error for each problem, and for its solution.

Camera app testing without a camera. I created the Y4M video file based on the photo of a face using manual writing of the RGB to I420 color conversion code because ffmpeg was missing, and processed it through the Chromium’s virtual webcam. This way, all the process became headless reproducible. Also, I got my favourite output from the project: pointing at a photo, the app properly doesn’t measure anything, states the fit that caused it not to measure and says that the camera worked on 5 tasks, but none gave a valid measurement.

Accomplishments that we're proud of

More often it is the case that truth lies in what was constructed. The program will not generate any number where no number is applicable, and thus it denies on too few frames, too many blinks, a moving target that didn't move, a head that did not turn, or a fit that is too wide to be trusted. Each denial provides an explanation and cites the fit, since a blank space counts as a normal reading, and a measurement made using only six frames is worse than no measurement at all. It would be taken as gospel.

The accessibility is clinical rather than cosmetic. Low stimuli, maximum luminance, no movement, 44px targets, keyboard access, and never a finding dependent on color alone. The readers are both photophobic and motion sensitive by nature, and print and greyscale must carry the same meanings. Below it all, there are 116 tests based on logic that earns them.

What we learned

The intriguing engineering involved in this healthcare device seems to lie in what it chooses not to say. The difficult parts were the dead zones, the distinction between null and zero values, and ensuring the accuracy of four different error messages. None of that is eye tracking.

Plus the human iris, located at a nearly constant distance of 11.7 mm, is a marvelous example of a free calibration tool. With one biological constant and a webcam, we get a measuring stick.

Plus rendering is debugging method. Three of the most glaring flaws in this project are a misstatement of clinical fact, an illegible printout, and an invisible focus ring in the primary control. All three were hidden in the code but visible from the first screenshot.

What's next for Baseline

Latency of proper saccades. Target jump times are already measured, and this is one of the more reliably measured concussion biomarkers. And then tracking the trend of objective improvements over time with symptoms.

And now for the hard truth: human validation. Oculomotor math is checked against artificial signals with solutions. This validates the math, not the clinical claim. Demonstrating that these metrics reliably track concussions in actual patients requires measuring a group relative to an eye-tracker reference, and that is the task that lies between this and any clinical tool.

All of this is spelled out in full detail in docs/LIMITATIONS.md, along with what Baseline intentionally does not do.

Built With

  • accessibility
  • canvas
  • computer-vision
  • github
  • github-actions
  • localstorage
  • mediapipe
  • on-device-ml
  • playwright
  • react
  • tailwindcss
  • typescript
  • vite
  • vitest
  • webassembly
Share this project:

Updates

Submission history