Inspiration
Every fighting game asks for two hands and fast reflexes on a dozen buttons. If you can't move that way, you're locked out, no matter how sharp your reactions are. We wanted a real fighting game you can play entirely with your face — no hands, from sitting down to the final rematch.
Once your face was the controller, we asked what else your body could drive. Your heart rate and breathing, read live from your webcam, became the match's balance: stay calm and you heal and charge faster, panic and you flinch harder. An AI announcer calls out what your body is doing as it happens — staying composed under pressure becomes a skill you practice, not a metaphor.
What it does
Headiator is a two-player fighting game you can play start to finish without ever touching a keyboard or mouse — biometrics and facial gestures drive the controls, the match balance, and even the menus:
- A gesture-operated start screen: smile to sit as Player 1, raise your eyebrows for Player 2, tilt your head to browse, hold a gesture to confirm. Two laptops join the same shared room automatically, with live presence ("Player 2 is here") and reconnect handling.
- Real heart rate, breathing rate, and a derived calm/stress score come from Presage SmartSpectra, read straight from each player's own webcam feed — no wearable hardware, and no setup: the deployed site runs Presage itself, so a player just opens the page and their camera starts working.
- Calm play heals you faster; a racing heart rate and shallow breathing raise the damage you deal and take, and gate whether your special-move meter can even charge. The tenser you get, the swingier the match gets.
- Hands-free combat, powered by on-device face tracking (MediaPipe): tilt your head left/right to move, tilt up to jump, and raise your eyebrows to fire a laser. Movement is calibrated per-player and edge-triggered so a tilt reads as one clean action, not a jittery mess.
- Falling hazards and prizes raise the stakes: swords drop from the sky with a fair floor warning (dodge, jump, or block, or take real damage), and a falling coin can be caught for bonus healing and meter.
- Live AI announcer: the game spots real moments as they happen (a heart-rate spike, staying calm at low health, a comeback, a knockout). Gemini writes a line about that exact moment using the players' live heart rate and health, and ElevenLabs speaks it about a second later. If an AI service is slow or down, it falls back to pre-recorded lines, so it never goes silent. -Post-match recap: built from real, code-counted fight stats (punches and lasers thrown and landed, damage dealt and taken, swords dodged, prizes caught), drawn as readable scoreboard cards. Gemini narrates from those numbers, so the story is grounded in what actually happened.
- End of round is hands-free too: smile to play again, raise your eyebrows to head back to the start screen.
- Full match and biometric logging to a real Postgres/TimescaleDB (Tiger Data) backend, with a dashboard for reviewing history afterward.
- Two-laptop play: one laptop hosts the authoritative match simulation, the other joins as a guest, each player using their own camera and their own biometrics.
How we built it
- Frontend: plain ES modules, no build step, rendered to a single canvas with a fixed-timestep game loop and a small event-bus architecture so the fighting engine, biometrics, commentary, hazards, and UI stay decoupled.
- Biometrics: Presage's SmartSpectra Node SDK doesn't run in-browser, so it runs inside our own server instead, same origin as the deployed site — the browser streams webcam frames in over WebSocket, the server runs real on-device vitals inference, and streams heart rate / breathing / calm / stress back out in real time. Each player gets their own independent session, correlated to their own page so two players' vitals can never cross on a shared public server.
- Face controls: MediaPipe's Face Landmarker runs entirely client-side (no server round trip, no API credits) to turn head pose and eyebrow-raise blendshapes into game input — the same input path drives the keyboard, the lobby, combat, and the end-of-round menu, so "gestures as controls" is one system, not three.
- Commentary and recap: a Gemini-generated line source paired with an ElevenLabs voice announcer for live commentary, subscribed to the same event bus as the game engine; the post-match recap is built from fight stats counted in our own code from the logged events (never invented by the model), rendered as drawn scoreboard cards, with Gemini writing a short, punchy summary over real numbers.
- Hazards and prizes: a host-authoritative falling-sword and falling-coin system, timed and positioned so every warning is fair to react to on gesture controls, not just a keyboard.
- Backend/data: an Express server backed by Tiger Data (TimescaleDB on Postgres) logs every match, biometric sample, and combat event for later review on a dashboard.
- Deployment: a single Render web service running the game, dashboard, API, the two-player relay, and Presage itself, all on one port; Vultr for infrastructure we run ourselves.
- Four of us split ownership along a shared contract (documented in CONTRACT.md) — fight engine and biometrics, mechanics balancing, AI commentary, and logging/deploy — so we could build in parallel without constantly stepping on each other's code.
Challenges we ran into
- Presage's SmartSpectra SDK only ships for iOS/Android/C++/Node — there's no
browser SDK — so getting real biometric data (not mocked) into a browser game
needed a server-side component from the start, debugged against real camera
hardware and a real API key rather than just code review. Our first working
version ran that piece as a small local process per laptop, which meant a
self-signed certificate to click through (browsers block plain
ws://from anhttps://page as mixed content, even to localhost). It worked, but it was a real piece of setup friction for anyone trying the game cold. We later moved that same logic directly into our deployed server, same origin as the site itself — so today a player just opens the page and it works, with the local version kept around only for offline development. - Tuning face-gesture controls to feel clean was harder than expected: an early version let a head-tilt-up (jump) accidentally also cross the eyebrow-raise threshold and fire the laser, and a quick head bob could trigger an unwanted jump. Fixing this properly meant axis-exclusive tilt detection (whichever direction is more clearly past its threshold wins), a short hold-time before jump fires, and suppressing the laser gesture entirely while the head is actively tilted — the same rigor we then reused to drive the lobby and the end-of-round menu by gesture too.
- Getting two independent laptops, each with their own camera and biometrics feed, into the same fair match without the two simulations silently diverging — solved by making one laptop the sole authoritative simulation and the other a pure input-forwarder/renderer.
- Running two real-time WebSocket features (the two-player relay and Presage) side by side on one server surfaced a genuinely obscure bug: a popular WebSocket library's own convenience mode for attaching to a path doesn't coexist with a second instance on the same server — whichever one attaches first silently rejects the other's connections. It only showed up once both features existed together against a real server, not in any mocked test, and needed the library's documented multi-endpoint pattern to fix properly.
- Making falling hazards fair on gesture controls, not just a keyboard: a sword warning that's plenty of time to react to with W/A/S/D can be too fast to react to with a deliberate head tilt, so the warning timing and hit zones were tuned specifically around that. -Making AI fast enough for a live fight. Gemini writes the line, then ElevenLabs voices it, and together that can easily take too long to feel live. We measured every step end to end: the average was fine, but the first line of a match took 1.6 s because both connections started cold. Warming up both APIs at round start, streaming the audio, and pre-recording the 71 most common lines brought it to about a second. -- An announcer that doesn't talk over itself. The game produces dozens of events a second. We built a priority queue: one line at a time, a knockout interrupts anything, each type of moment has a cooldown, and anything more than a second or two old is dropped.
Accomplishments that we're proud of
- A complete hands-free loop, not just hands-free combat: picking your seat, playing the match, and choosing to rematch or leave are all driven by the same face gestures, tuned to be reliable rather than gimmicky.
- Real, working biometric input from a live webcam feed — not a simulated or mocked signal — actually changing match balance in real time, with zero local setup required to play the deployed site.
- Real-time AI commentary and a post-match recap that are honest: Gemini narrates numbers our own code actually counted from the match, never numbers it invented.
- A real two-laptop, two-camera multiplayer mode working end to end, each player bringing their own biometrics to the same fair, host-authoritative match.
- Falling hazards and prizes that add genuine risk/reward on top of the biometric mechanics, without undermining the "calm is an advantage" core idea.
What we learned
- Biometric signal from a webcam is a slow, noisy "vibe" signal, not a per-frame reactive one — designing mechanics that lean into that (calm as a sustained advantage, not an instant trigger) worked far better than trying to force it into a twitchy input.
- Face-gesture control needs the same rigor as any other input system: exclusivity between directions, debouncing, and edge-triggering aren't optional polish, they're the difference between "controllable" and "random." Once we got that right for combat, reusing it for menus was nearly free.
- The pragmatic fix and the right fix are sometimes two different commits: shipping a local-process workaround got real biometrics working fast, but moving that logic into the server itself later removed a whole category of setup friction for players — worth budgeting time to revisit "good enough" solutions once the core idea is proven.
- Some bugs only exist at the intersection of two features that each work fine alone (our two real-time WebSocket endpoints, on their own, both passed every test) — live, end-to-end testing against a real server caught what mocked unit tests structurally couldn't. -Measure the worst case, not the average. Our average latency looked fine; the cold first request was the real problem. -A good announcer knows when to stay quiet. Choosing what not to say mattered as much as generating lines.
What's next for Headiator
- Per-player calibration for head-tilt sensitivity (we already do this for facial gestures), so movement thresholds match each person's actual range of motion instead of one fixed constant for everyone.
- More biometric-driven mechanics beyond heal rate and meter gating — e.g. voice volume affecting punch strength.
- More hazard and prize variety now that the falling-object system is proven out, each tuned for fairness under gesture controls the same way the sword was.
- Spectator mode and match replay from the logged biometric and combat history, so a full match (inputs, biometrics, hazards, and commentary) can be watched back afterward.
Built With
- elevenlabs
- express.js
- google-gemini
- html5
- javascript
- mediapipe
- node.js
- postgresql
- presage-smartspectra
- render
- tiger-data
- timescaledb
- websocket

Log in or sign up for Devpost to join the conversation.