Inspiration

  • The idea came from Justina Miles signing Rihanna's Super Bowl halftime show. Everyone was talking about her, and you can see why: she put so much energy and expression into the performance.
  • We wanted to try translating signs into speech, and in a way that reflects the signer’s expressions too.
  • The goose is a nod to Waterloo. If we're building at Hack the North, we might as well embrace the geese.

Justina Miles signing during Rihanna's Super Bowl halftime show.

Justina Miles performing in ASL during Rihanna's 2023 Super Bowl halftime show.

What it does

  • AI harness that translates supported ASL signs into English through a talking goose that shows emotions.

How we built it

We built the consumer app with Expo, React Native, and TypeScript, connecting camera tracking, sign recognition, English speech, and our goose.

How we built it

  • Cloud experiments: We used Backboard to route structured sign-classification requests to TypeSafe's Jev, and built a separate Cerebras / GPT-OSS-120B caption endpoint with JSON and vocabulary checks. These were part of our earlier cloud prototype; the flow above uses local sign matching.
  • Recognition research: Python, OpenCV, and NumPy helped us process videos and prepare landmark sequences. We explored PyTorch models, data augmentation, calibration, and replay benchmarks, tested TensorFlow Lite/LiteRT inference, built Core ML export tooling, and used Apple Accelerate in the standalone fingerspelling experiment. The live matcher uses hand-shape normalisation, body-relative motion, gesture segmentation, and Dynamic Time Warping (DTW).
  • Reliability: We checked that recognition calculations matched in Python and Swift, and used session state machines to keep duplicate detections and outdated results from triggering speech.

Build Highlights

You sign. A goose says it out loud, in the mood of your face.

PHONE (Expo + Swift, on-device)
  camera → MediaPipe hands/body/face → sign matcher → word + mood
                                                         │ text + mood label only
BACKEND → ElevenLabs Flash v2.5 → audio + timestamps ────┘
  goose speaks, gestures timed from the timestamps

Our big design call: video never leaves the phone. In the shipped app, only text and a mood label go to the cloud, after you consent.

Built with Expo

Expo Router, Reanimated, Expo GL, Haptics, and a custom Swift module. The camera, the goose, and the voice are built to feel native, and React Native never sees a camera frame.

Expo tool What it does in Honk & Tell
Expo SDK 55, React Native, TypeScript The whole app, with a typed native bridge.
Custom Expo module (Swift) signloop-camera runs MediaPipe hand, pose, and face tracking on one camera session. Only status, word guesses, and expression labels reach JavaScript.
Expo Router /, /conversation, /settings, plus local sheets for transcript, correction, and end.
Expo GL + React Three Fiber The 3D goose, drawn from Three.js primitives, with a lighter render on software GL.
Reanimated Spring-driven buttons and sheets, layout transitions, and the goose's trip between screens.
Expo Haptics A success tap when a new phrase lands.
Expo Audio Clip playback in the goose prototype. In the app, a native MP3 player feeds its clock to the goose.

Details that make it a joy to use

  • One goose that travels. It stays mounted and slides between Home and Conversation instead of reloading. Direct entry and Reduce Motion skip the trip.
  • A camera that shows your hands. The preview is aspect-fit with matching mirrored overlays, so the small camera tile never crops away a sign.
  • A layout that doesn't jump. Camera, goose, and captions keep fixed regions (about 45, 30, and 25 percent). Framing hints wait for 300 ms of stable readiness, so they don't flicker. Captions stay on screen through tracking loss and pauses.
  • A goose that reacts to you. It watches while you sign, follows your live expression, and speaks with the mood locked to the word.
  • Good in the hand. Spring-driven buttons and sheets, and a haptic when a phrase lands.
  • Respects your settings. System text size and Reduce Motion are honored. Bigger text gives captions more room without shrinking the goose, and captions are a polite live region.
  • Honest states. Labels say whether it's preparing the voice, speaking, interrupted, or failed. Pause, sheets, and backgrounding stop capture and playback, and voice consent resets when you leave the app.
  • No ghost words. Every result carries a capture ID and its frame's timestamp, and stale ones get dropped. Pausing or flipping the camera can't replay an old word.

Built with ElevenLabs

You sign, and a goose answers out loud. Every completed sign is spoken automatically, in a mood read from the signer's face, in a custom goose voice.

  • A voice with a character. A custom ElevenLabs goose voice gives the app a personality instead of a generic narrator.
  • Emotionally expressive. Six styles: neutral, joy, sadness, anger, fear, and disgust. Each has its own speed, stability, similarity, style, and speaker boost. Sadness drags at 0.76x, and fear rushes at 1.2x.
  • Face to voice. The signer's expression locks to the word when the sign finishes, so relaxing early doesn't change the delivery. Replay and text corrections keep the original mood.
  • Autonomous. There's no Speak button. With voice on, each completed sign speaks on its own, queued in order, each with its own mood.
  • A model tradeoff we built both sides of. The backend carries v3 audio-tag mappings (fear becomes worried and nervously) and Flash v2.5 settings. We shipped Flash, since the goose waits for the whole clip and speed won.
  • Timestamps in the loop. The with-timestamps endpoint returns character-level timing, and the goose uses it to cue gestures while it talks.
  • Same take on repeat. A seed hashed from emotion and text asks for the same take, and the last 24 clips are cached.
  • Production habits. The API key stays on our backend. Voice needs your consent, redirects are refused, sizes are capped, responses are validated, and nothing retries.

Streaming PCM audio and amplitude-driven lip-sync are next.

Built on Backboard

We ran a brand-new System One model, TypeSafe's Jev, inside a live sign-recognition loop through Backboard's API.

  • Not a chatbot call. Jev returns typed answers with probabilities instead of paragraphs. A recognizer needs exactly that: a label, or an honest UNKNOWN.
  • The loop. The phone sends 1.2 s of hand landmarks (never video) to our backend, Backboard runs Jev, and we turn the answer into ranked candidates plus UNKNOWN. Captions come from Cerebras GPT-OSS-120B.
  • Built for messy real time. Replies show up late and out of order, so each one is checked against its attempt, revision, and request ID and dropped if stale. Provider work and queued windows are bounded, and hung inference runs in processes we can kill and restart. A held sign counts once.
  • It grew into a pipeline. Jev reranks candidates with context inside a streaming setup: overlapping windows, a temporal resolver, and fusion across four OpenHands WLASL checkpoints that validates the full probability distribution.
  • Careful with data. Landmarks only, keys stay server-side, the UI says when cloud analysis is on, and gateway records get best-effort cleanup.

Live recognition in the demo now runs on-device so it works offline. The Backboard path is our cloud and research route.

What's next

We'd move the clip cache into a shared store so several servers can use it, and add per-user auth and rate limits. HTTPS comes before any public release, since local development runs over plain HTTP.

Challenges we ran into

  • ASL has its own grammar, with meaning carried through hand movements and facial expressions. Translating it remains an active research problem, partly because of limited training data. We started small with a focused set of signs.
  • Knowing when a sign was finished took a few tries. Deciding too early could miss part of the movement; waiting too long made the app feel slow. We tested and revised the timing to balance the two.
  • Facial expressions were difficult to separate: an open-mouth smile could resemble fear, while small brow movements were easy to miss. We refined the facial cues and sensitivity to pick up subtler expressions.

Accomplishments that we're proud of

  • Getting sign recognition running on a phone while keeping camera footage on-device.
  • Giving our goose six expressions, each with its own animation and voice delivery.
  • Building recognition around hand shape, movement, and where a sign happens relative to the body.
  • Creating a character and visual style that feel like us. The Waterloo goose had to happen.

What we learned

  • Starting with fewer signs gave us room to understand where recognition was going wrong.
  • A fast model doesn’t automatically make a fast app. Deciding when a sign is finished can take longer than recognising it.
  • A single facial movement can fit several expressions. We had to look at combinations of cues.

What's next for Honk & Tell

  • Expand our vocabulary and test with more signers, signing speeds, and camera setups.
  • Move from individual signs toward continuous signing and full sentences.
  • Add speech-to-text so the person signing can also read the other side of the conversation.
  • Let people choose their own voice or character; although we’re pretty attached to the goose.

Built With

Share this project:

Updates

Submission history