Inspiration
Navigating a road also means navigating its language. Personal experience from some of our team members showed us how unfamiliar road signs can challenge drivers learning English, especially those who drive for work. We wanted to make English practice practical and personal by turning the signs drivers encounter into lessons they can use.
What it does
ubi detects signs in dashcam footage and calls them out aloud. Once the driver presses Park, those captured signs become a personalized English practice session. Drivers can respond by voice or text, including through iMessage with our Photon integration.
How we built it
We built ubi with a Python backend using FastAPI and a web interface. Its recognition pipeline combines a pretrained US traffic-sign model with on-device text recognition for real-time responses. We manually reviewed sign predictions from real driving footage, refined the pipeline, and fine-tuned the model to improve accuracy. ElevenLabs generates spoken sign announcements, while Grok Voice powers conversational practice when the driver is parked.
Learnings and Challenges
During the hackathon, we learned to develop with AI tools such as ElevenLabs and Grok Voice and make their APIs work well together. We initially relied heavily on a cloud-based approach but ran into latency issues and quota limits. This led us to shift to on-device models, allowing us to deliver real-time responses.
We also learned more about computer vision and machine learning, including how to use segmentation principles to identify road signs based on color and shape. We fine-tuned our model by manually reviewing misclassified examples and assigning them to the correct road sign categories.
Getting clean images for the model was one of our biggest challenges, especially with small or blurry signs, nighttime footage, and bad weather.
Beyond the technical work, we learned how to build a good user experience with artificial intelligence. How much should we say out loud? Should we announce duplicate signs encountered on the road? How can we support language learning without making users feel ashamed? We worked through these questions as a team and felt we navigated them well. Other challenges included finding YouTube clips with enough road signs to experiment with. We initially used Gemini but ran into rate limits, leading us to switch to Grok.
Log in or sign up for Devpost to join the conversation.