Inspiration
Cooking from a recipe means touching your phone with flour on your fingers, then finding your place again. Recipes are also written to be read, not cooked from: "preheat the pan" hides mid-paragraph, and by the time you reach it you're already late. Voice assistants can read steps aloud, but none can see your food, so none can answer the question cooks actually ask: is this ready? We wanted a cooking buddy for beginners who give up on new recipes because of that friction.
What it does
Hey Remy is a hands-free cooking buddy. Paste any recipe and it reorders it into short steps. Hidden prep, like heating the pan, moves into a "before you start" list, and each step gets a heads-up one step early. You can scale the servings.
In cooking mode the current step is shown in large text and read aloud:
- 👍 Thumbs-up moves to the next step
- 👎 Thumbs-down goes back
- ✋ Open palm asks Remy to look at your food
A camera on the chef's hat takes one photo, and Remy says whether the step's "ready" cue is met ("Still lumpy, keep whisking"). You never touch the screen after pressing Start.
How we built it
- Frontend: Vite, React and TypeScript
- Gestures: MediaPipe Gesture Recognizer runs in the browser on every frame, so navigation is instant and never touches the network
- Recipe parsing and food checks: the Gemini API, behind serverless functions that keep the key off the client. Each step carries a visual cue, and the vision check is judged against that cue
- Voice: ElevenLabs text-to-speech for steps and verdicts
- Camera: one Logitech webcam taped under a cap, seeing both hands and food
- Architecture: a small state machine and controller, with the camera and API behind interfaces. Most of the logic is unit tested and runs against a mock backend
Challenges we ran into
- A hat cam moves, so frames blur. A check grabs about half a second of frames and sends the sharpest.
- Hands are always in frame while cooking, so gestures trigger by accident. We require a one-second hold plus a cooldown.
- Judging "is it ready?" from a photo is fuzzy, so the model can answer "unsure" and we only offer checks on steps with an obvious visual change.
- Model output has to be trustworthy, so we validate every parsed recipe before the UI sees it.
Accomplishments that we're proud of
- It works end to end: a pasted recipe, a hands-free cook, and a spoken verdict from a real photo of the food.
- Remy reorders recipes without deleting steps or changing amounts.
- A tested, mockable architecture that let the whole team build in parallel in a short time.
What we learned
- Grounding the vision model in a per-step cue makes its answers far more consistent than asking "is this done?".
- Hands-free design is mostly about avoiding false triggers, not recognizing gestures.
- Agreeing on shared data shapes in the first hour is what let four people work at once.
What's next for Hey Remy
- A prep check: tell Remy what you have, and it flags what's missing and suggests substitutes
- Auto-checking a step in the background, speaking up only when the cue is met
- Automatic timers for steps like "bake 20 minutes"
- Importing recipes from a link, and an offline voice fallback
Built With
- claude
- elevenlabs
- gemini
- react
- typescript
Log in or sign up for Devpost to join the conversation.