Inspiration

Ratatouille's whole premise is "anyone can cook" if they have the right guidance standing next to them. We're broke, busy university students who mostly survive on the same five meals because cooking anything new feels like too much effort after a full day of classes. We wanted a Remy of our own: something that could stand at your shoulder, watch what you're doing, and talk you through an actual meal instead of you defaulting to instant noodles again.

What it does

Whisker acts as an AI Cooking Mentor, teaching you the recipe, watching your steps in the kitchen, and giving you tips along the way. Whisk-Ella, your masterchef cat, uses a camera and voice to assist you autonomously. The voice answers your questions and provides audible next steps, ensuring your messy hands stay off your screen. The user interface allows you to search up recipes to cook, using filters to specify dietary needs and using a search engine to find certain recipes or ingredients. An upload feature can be used to add any blog recipe or short-form video into the recipe database. You can create a personalized cookbook, favouriting recipes to save for later.

How we built it

Whisker is a web app with two Python services behind it. Everything Whisk-Ella sees, hears and says comes from one AI model, Qwen3.5-Omni.

The app. Built with Next.js and React. You search recipes with diet filters, save favourites to your cookbook, and upload a YouTube or TikTok link that Omni turns into a recipe. The cooking screen shows your webcam.

Voice. Omni can both hear and speak, so we didn't need separate speech tools. Ask a question out loud and Whisk-Ella answers in her own voice, in English, French or Spanish.

Vision. Press "Check step" and the webcam takes a photo. Omni looks at it alongside the step you're on and says whether it's done or what's missing, and Whisk-Ella reads it out.

Hands-free. A second service streams your mic to Omni live, so you can just talk, and even interrupt her. A camera coach checks the pan every few seconds, but only when the picture changes, and can move to the next step or set a timer for you.

Memory. Using Backboard, Whisk-Ella only remembers allergies or preferences if you tell her to ("remember I'm allergic to peanuts"), because storing them automatically raised privacy concerns.

Built in 36 hours. We split the project into separate parts (app, voice, live, vision) so we could work in parallel. When the OAK-D camera fell through, we swapped in a webcam without changing the rest.

Challenges we ran into

We had a hardware limitation as we were unable to get the OAK-D depth camera as originally planned. This forced us to pivot to a web camera and placing more focus on the voice assistance and AI detection. Additionally, we struggled with the limitation of a delayed camera feed. To receive feedback, the camera had to take a photo and then prompt, making the process slower.

Accomplishments that we're proud of

We're proud that we got a real end-to-end multimodal experience working, camera, voice, and language all responding to each other live, inside a 36-hour build, without ever having the "ideal" hardware we started out planning for. We're also proud of how well Whisk-Ella lines up with Huawei's OMNI Live challenge: vision, speech, and language aren't three separate bolted-on features in our app, they're one continuous loop, the camera watches your step, the voice reads it out and answers your questions, and the two stay in sync as you cook. That's exactly the kind of natural, real-time multimodal interaction the challenge is asking for, and getting it to feel responsive instead of clunky, especially after our camera pivot, is something we're genuinely proud of.

What we learned

On the AI side, we learned a lot about the quirks of working with specific recipes and the restrictions models have around food and nutrition content. For example, we learned that having LLMs store allergies often caused privacy issues. We also learned that keeping the camera and voice live and in sync is a much harder problem than it sounds, it's not just snapping a picture every second, timing, latency, and matching the right frame to the right step in the conversation all have to work together for it to feel natural rather than laggy.

What's next for Whisker

Moving forward, we'd bring back some of the features we cut for time, like the fuller filter and sort options, and layer in more detail throughout the recipe experience, richer nutrition breakdowns, more precise cost estimates, and a wider variety of cuisines. The bigger vision is turning Whisk-Ella into something social: letting students share recipes they've made or discovered, browse what students from other schools are cooking, and build a community cookbook that grows the more people use it. This way, the app gets more useful to everyone, not just the person who fed it their own ingredients.

Built With

  • api
  • vision
  • voice
Share this project:

Updates

Submission history