Inspiration

Everyone has a graveyard of saved recipe videos that never get cooked. Recipe-extraction tools already solve "save this recipe" — the harder, unsolved problem is deciding what to actually cook tonight, together, with the people you eat with. We built Doomscroll & Dine to answer that question instead of just organizing what's already been saved.

What it does

  • Paste an Instagram link and a multimodal AI pipeline (Extractor and Verifier agents) watches the video, reads the caption, and pulls out a structured recipe — ingredients, quantities, and steps — even when a creator only speaks the steps aloud or shows ingredients on screen without ever stating an amount. The Verifier agent double-checks every field against the source (caption and video) so nothing is invented.
  • Every imported recipe is embedded into a vector database, powering Cook Together: each friend gets a personal AI agent that scores candidate recipes against their own tastes and cooking history, and a Planner agent picks the group's meal and explains why.
  • An AI Price Estimator gives an honest, clearly-labeled grocery cost estimate for each recipe — no fragile live grocery API required.
  • Step-by-step voice narration, built with ElevenLabs, reads instructions aloud one step at a time so hands stay clean while cooking.
  • Runs as an installable PWA so pasting a link and viewing or narrating recipes works on both desktop and phone, including sharing straight from Instagram's native share sheet on Android.

How we built it

  • Backend: FastAPI (Python), which orchestrates our AI pipelines, parallel agent execution, recommendation logic, and database access.
  • Frontend: React 19 + Vite + TypeScript, installable as a PWA with Android share-target support. Features include recipe browsing with difficulty ratings and ingredient counts, a Cook Together friend selector, food preference and allergy profiles, recipe detail pages with price estimation and voice narration, and a full pink/yellow design system built from scratch.
  • Database: PostgreSQL (TimescaleDB), pgvector for recipe embeddings and semantic recipe retrieval, Supabase for auth
  • AI / Multi-Agent System: Gemini (multimodal) for extraction, verification, price estimation, and recommendation reasoning -- Recommendation System: Recipe embeddings and pgvector semantic search generate relevant candidates before agent reasoning, keeping retrieval separate from LLM decision-making.
  • Voice: ElevenLabs for step-by-step narration
  • Video retrieval: yt-dlp

Challenges we ran into

  • Recipes are often split across three places — the caption, spoken narration, and pure visual demonstration — so a text-only extractor silently misses whole recipes. We made extraction genuinely multimodal (reading, watching, and listening), and had to teach the Verifier to check against all three sources instead of wrongly flagging real video-sourced content as hallucinated.
  • Handling missing quantities honestly: rather than guessing silently or leaving gaps, our agents provide their best visual estimate and clearly label it as approximate.
  • The hackathon venue WiFi blocked our database port entirely, costing us several hours of debugging before we traced it to the network. We switched to a mobile hotspot and never looked back.
  • Coordinating three people's schema changes, merge conflicts, and Gemini free-tier quota limits in real time under a 36-hour clock.

Accomplishments that we're proud of

Our extraction system uses a specialized Extractor → Verifier pipeline that can read captions, listen to narration, and inspect visual information in cooking videos while distinguishing sourced information from estimates.

For Cook Together, we built a parallel multi-agent recommendation workflow. Each person gets an independent Personal Agent that evaluates candidate meals from their perspective. Those evaluations are passed to a Planner Agent that resolves competing preferences and explains why its final recommendation works for the group.

What we learned

Multimodal AI is a genuinely different (and much harder) problem than text-only extraction. The "obvious" solution (read the caption) breaks the moment a creator relies on video and voice instead of writing everything down. We also learned that infrastructure surprises like blocked ports can cost more time than any code bug.

What's next for Doomscroll & Dine

Live grocery pricing, more narration voice options, and expanding Cook Together to handle larger friend groups.

Built With

Share this project:

Updates

Submission history