Inspiration
All of us struggled with at least one step in every meal prep journey. From deciding what to meal prep, where to get ingredients that hit our macros, how to cook each meal, and the overall inconvenience of having to type everything up on our phone. We wanted to create something that removed the inconvenience and thought of having an AI to talk to.
What it does
Sous is a voice-first AI sous chef. • Pick a chef: four mascot personas (Maya, Leo, Nova and Brock), each with its own ElevenLabs voice and personality. • Scan your fridge: snap a photo and Gemini lists your ingredients, then Sous suggests recipes ranked by what you already have. • Cook hands-free: Sous reads each step aloud. Say "next", "go back" or "go to step five", ask cooking questions, or say "show video" to play the recipe's video right in the app. • Shop: a grocery list of what you're missing, with nearby stores, price estimates and an in-app map. • Track nutrition: snap your plate or just say "I had a chicken wrap and a latte". Gemini estimates calories, macros and vitamins, and your diary tracks 15 nutrients against your goals. • Share: post your dish to a swipeable community feed, moderated by Gemini.
How we built it
Figma: Every screen and all four mascots (Maya, Leo, Nova, Brock) were designed before any code, including design tokens, interaction rules, and mascot states
Core Stack
Next.js 16, React 19, TypeScript, Tailwind CSS: Everything on the frontend Node.js: Server-side API routes that keep all API keys off the client Zustand: App state, with user data kept on the device Google Gemini API: The brain. Function calling, vision, nutrition estimates, and post moderation ElevenLabs: The voice. Flash v2.5 for low-latency streamed replies, and v4 for each chef's pre-recorded greeting Web Speech API: Browser speech-to-text for voice input TheMealDB + curated catalog: Recipe data YouTube: Recipe videos that play in the app, starting from your current step Claude Code: AI coding assistant. We directed and tested each feature ourselves
Key Features
AI That Acts, Not Just Talks: Gemini function calling with 14 app actions decides what the app does every turn. Each call is validated against the live app state. Fridge & Plate Vision: Snap a photo and Gemini returns your ingredients or meal as structured JSON Hands-Free Cooking: Say "next," "go to step five," "show video," or ask a cooking question Natural Voice Loop: Sous waits for a 2.5-second pause before replying, never hears itself or the video, and reopens the mic at the right moment Smart Groceries: Lists what you're missing, with nearby stores, price estimates, and an in-app map Voice Nutrition Tracking: Say "I had a chicken wrap and a latte" and Sous tracks 15 nutrients against your goals Built to Keep Working: Every AI call walks a chain of fallback models, and every feature has an offline backup
Architecture User Voice → Web Speech API → Next.js API Route → Gemini (function calling + live app state) ↓ 14 App Actions (fridge scan, recipes, cooking steps, video, groceries, map, diary log) ↓ ElevenLabs Flash v2.5 → Spoken reply to user
Photos (fridge / plate) → Gemini Vision → Structured JSON → Recipes / Nutrition Recipes ← Curated Catalog + TheMealDB API
Challenges we ran into
Standing out. Recipe apps and calorie trackers already exist. To make Sous different, we focused on three things: it's voice-first, it takes real actions in the app, and it has four chefs with their own personalities.
Voice and response time. Pauses feel awkward when you're talking out loud. We used ElevenLabs' fastest voice model and pre-recorded each chef's greeting to cut wait times. Sous also had to wait for you to finish speaking, ignore its own voice, and reopen the mic at the right time.
Real conversations. People ask follow-ups, jump between steps, or just say "yes." Sous has to understand what you mean in the moment. We give Gemini the app's current state and make it take an action every turn, not just talk.
Trustworthy information. Every recipe and grocery result links back to its real source, so Sous never makes things up.
Network issues and quotas. The Wi-Fi was unreliable, and we hit our daily API limits fast. So every AI call has backup models, and every feature works offline.
Accomplishments that we're proud of
What we learned
What's next for Sous
Video mode: A FaceTime-style experience where Sous can actually see what you're cooking, plus the ability to send photos in chat. More natural voice conversations: Fully real-time conversations where you can talk back and forth with Sous. Actually Getting it on a mobile device.
Built With
- claude
- elevenlabs
- figma
- gemini
- gemini-api
- next.js
- node.js
- react
- tailwindcss
- themealdb
- typescript
- web-speech-api
- youtube
- zustand
Log in or sign up for Devpost to join the conversation.