Inspiration

Every day you say something you'll want back later: the restaurant a friend swore by, the appointment you agreed to in passing, the thing you keep meaning to do. Writing it down is friction, so nobody does it, and the detail is gone by the weekend. Journals and notes apps ask for discipline at exactly the moment you have none. So we set out to remove the friction entirely: you speak for twenty seconds, and that's the whole interaction.

What it does

What happens next is the product. Recoll understands what you said and quietly files it: who was there, where it happened, what it was about. It hears the commitments hidden inside ordinary sentences and turns them into reminders that arrive when they're useful. A promise made on Sunday night surfaces at 9am on Tuesday, in your timezone, never while you're asleep. And months later you can ask a question in plain language, "what was that seafood place Carlos recommended?", and get an answer drawn from your own life, with the exact moments it came from.

  • Ask is a conversation. Follow-ups like "and when was that?" just work, and when the answer isn't in your memories, Recoll says so instead of guessing.
  • You decide how much initiative it takes, from "only what I ask" all the way to "surprise me".
  • Type when you can't talk. Written memories go through the same pipeline.
  • Quick capture. On Android, you can record from a notification or an overlay without opening the app.
  • In English and Spanish, built for Android and iOS.

How we built it

Recoll is a Flutter app on top of Firebase, with Gemini doing the listening and the understanding.

  • App: Flutter (Riverpod, GoRouter), one codebase for Android and iOS, plus a native Android foreground service so a recording survives even if the app is closed.
  • Backend: Firebase: Auth, Firestore, Cloud Functions v2 (Node 22 + TypeScript) and FCM for push notifications.
  • Understanding: Gemini 2.5 Flash transcribes and structures each memory in a single call that returns JSON: summary, date, people, places, topics and reminders.
  • Recall: each memory becomes an embedding (gemini-embedding-001) stored as a Firestore vector and searched with findNearest. Ask is a retrieval-augmented generation (RAG) pipeline built on top of that, and it keeps the context of the conversation.
  • Evaluation: a test set of 50 cases with known right answers (silence, white noise, Spanglish, long monologues, prompt injection spoken out loud, 19 borderline reminder cases, written entries), plus a separate set for Ask. It also tracks latency, thinking tokens and cost per memory.
  • Subscriptions: RevenueCat, with each user's plan mirrored and enforced on our server.
  • Process: separate dev and prod Firebase projects, GitLab CI, and code review on every merge request.

Challenges we ran into

Two problems turned out to be much harder than anything else.

The first: hand a model twenty seconds of silence, or of a café's background noise, and it will cheerfully hand you back a memory that never happened. In a journal, an invented memory is worse than no memory at all. We made the pipeline explicitly reject audio with no speech in it, and we test the opposite failure too: throwing away real speech because there's noise behind it.

The second: deciding when Recoll may schedule something on its own, and when it should only suggest it. Too eager, and the app becomes a stranger putting things in your calendar. Too timid, and it's a notebook you have to manage. We rewrote the reminder logic as an explicit decision tree (did you ask for this, or were you just thinking out loud?), and on our borderline cases the number of reminders nobody asked for dropped to zero. Then we gave the user a dial, from "only what I ask" all the way to "surprise me".

Both problems are covered by the same test set, with deliberate traps for each of these failure modes. We re-run it before shipping any change to how memories are interpreted.

Along the way we also cut a feature on purpose. Recoll used to label the emotion of every memory to draw an emotional timeline. Under GDPR that's sensitive data, so we removed it completely instead of hiding it behind a switch. And we switched AI providers mid-project, from Whisper + GPT-4o to Gemini, merging transcription and understanding into a single call. The test set is what made that switch safe.

Accomplishments that we're proud of

  • The first time you ask about something you had completely forgotten and it comes back, along with the moment it came from, it feels like magic.
  • An AI that would rather say "I don't know" than make something up.
  • An eval-first way of working: nothing that changes how memories are interpreted ships without passing the test set, traps included.
  • Privacy as a design rule, not just a policy page: explicit consent enforced on the server, everything wiped when you delete your account, and a feature we cut because it wasn't worth the risk.
  • A real product, not a demo: two platforms, two languages, subscriptions and a production backend, built by a small team.

What we learned

  • "Nothing" has to be a valid answer. Whether it's a silent recording or a question your memories can't answer, the model must be allowed, and tested, to return nothing rather than something plausible.
  • Tests with known answers beat gut feeling. More than once, a prompt tweak that looked better in a quick manual test broke something else.
  • The hardest problems in consumer AI aren't model problems. They're judgment calls: when to act, when to suggest, and when to stay quiet.
  • Limit the model's thinking budget, but don't turn thinking off: with too little, relative dates like "last Friday" start breaking.
  • Collecting less data is a feature. Everything we chose not to collect made the product easier to trust and easier to explain.

What's next for Recoll

  • Recaps: your week and your month, summarized back to you.
  • Improved reminders
  • Automatically splitting a recording that covers several things into separate memories, with one-tap undo.
  • Images in memories
  • Quick capture on iOS.
  • Moving to the Gemini 3.x models, which we're already testing against our eval.

Built With

Share this project:

Updates

Submission history