Inspiration
Over 60 million people in India rely on sign language to communicate, yet most existing tools either ignore this entirely or only capture half the picture — the hands. Sign language isn't just hand shapes; facial expressions carry real grammatical meaning. Raised eyebrows can turn a statement into a question. Furrowed brows signal emphasis. We realized almost no accessibility tool actually captures this — so we set out to build one that does.
What it does
SignSense reads hand signs and facial expressions through a webcam in real time, and turns them into natural spoken sentences. A user signs a sequence of words, the system detects their facial expression alongside each sign, and combines both into a grammatically accurate sentence — phrased as a question, a statement, or with added emphasis, depending on what the signer's face actually indicated. The final sentence is read aloud using speech synthesis, completing the loop from sign to spoken language. No special hardware or app installation required — just a browser and a webcam.
How we built it
We used MediaPipe to extract hand and facial landmarks from live webcam frames. For hand signs, we recorded our own training data and trained a classifier that currently recognizes sixteen distinct signs. For facial grammar, we calculate relative geometry — eyebrow position compared to eye position, and head tilt — normalized against a calibrated neutral baseline, so the system works regardless of how close the signer is to the camera. Both signals feed into Google's Gemini API, which doesn't just string detected words together — it uses the facial context as a real instruction, shaping the sentence structure itself. The backend is built with FastAPI and Python, cleanly separated into routes, services, and schemas. The frontend is a React app that captures live webcam frames, calls our backend in real time, and uses the Web Speech API for the final spoken output.
Challenges we ran into
Getting facial expression detection to mean something useful, rather than just displaying a label, was the hardest technical challenge — we had to move from raw landmark coordinates (which change with camera distance) to relative geometric ratios that stay consistent regardless of how far the signer is from the camera. We also hit real integration issues: dependency conflicts between MediaPipe and newer Google API libraries, a deprecated Gemini SDK we had to migrate away from mid-build, and inconsistent training data that needed cleaning before a model would even train. On the data side, training a reliable classifier with only one signer's data meant some signs performed far better than others, which we had to work around carefully for a reliable live demo.
Accomplishments that we're proud of
Getting facial expression to genuinely change the output sentence — not just as a cosmetic badge, but as an actual instruction the language model follows — is the piece we're most proud of. We have real before-and-after proof: the same three signed words produce "You need water" when neutral, but "Do you need water?" when the signer's eyebrows are raised. That's a real, demonstrable grammatical difference driven by real facial data, and it's the core idea most sign language tools miss entirely.
What we learned
We learned how much sign language grammar actually depends on non-manual markers — facial expression, head position — something neither of us had appreciated before researching this. On the technical side, we learned to work with real-time computer vision pipelines, how to design geometry-based features that generalize across different camera setups, and how to prompt an LLM with structured context rather than just raw text. We also learned firsthand how fast the AI tooling landscape moves — we had to adapt mid-project when a model we relied on was deprecated.
What's next for SignSense — Sign Language Translator
Our biggest next step is expanding beyond single-signer training data to a multi-signer dataset, which would meaningfully improve real-world accuracy. We'd also like to move from static, isolated signs toward continuous, motion-based sign recognition, since many real signs involve movement that a single frame can't fully capture. On the output side, we're interested in building a two-way system — letting a hearing person speak or type and have it translated back into sign language, potentially through a video or avatar-based signer rather than text alone. Longer term, we're also interested in exploring physical hardware extensions, like a haptic feedback device, to bring this beyond the browser and into a fully embedded, accessible product.
Built With
- computer-vision
- fastapi
- framer-motion
- gemini-api
- google-gemini
- javascript
- joblib
- llm
- machine-learning
- mediapipe
- numpy
- opencv
- pydantic
- python
- react
- rest-api
- scikit-learn
- uvicorn
- vite
- web-speech-api
Log in or sign up for Devpost to join the conversation.