-
-
The whole flow: create a family space, record or approve a story, then repeat. Senders need no account setup beyond a one-time grant.
-
The core constraint: Gemini shapes the script, but the relative reads it. We never replace the family voice with synthetic narration.
-
The safety gate: nothing reaches the child until the Owner delivers it. Downloading or sending to Yoto is the approval. No separate step.
-
Gift credits are the revenue model: the distant relative pays, because buying a thing for a grandchild beats subscribing to another app.
-
Record once, send to many: one narration fans out to up to 50 destinations, each an independently funded delivery with its own Yoto icon.
-
The promise: a bedtime story in a grandparent's real voice, playing on the child's own Yoto player. No screen, no adult phone in the loop.
-
Three entry paths: guided questions, paste a script you already wrote, or just press record. All three land on the same teleprompter.
-
Parents post a wish, so a grandparent opens the app with something to record instead of a blank page. Each wish is a one-tap prompt.
-
Gemini reads the story, works out its subject, and gives it cover art the child recognizes on the shelf. It never blocks sending.
-
Every story a family has ever received on one shelf, newest first, filterable by voice, with a player that never leaves the page.
-
The teleprompter scrolls at the reader's own pace and dims every line but the current one, so a grandparent can look up and stay warm.
-
After a story lands, Gemini pulls out the new words the child met and suggests what to record next. Bedtime turns into breakfast talk.
-
Voice Restoration: about ten minutes of old recordings rebuild a voice for someone losing theirs. Consented, human-reviewed, revocable.
-
On a phone the whole session fills one screen. Script, countdown and microphone together, with no scrolling between them mid-take.
-
Gemini interviews the Sender out loud when they do not know what to say, then turns the answers into a first draft.
Inspiration
We are Juan and Jordan, a married couple raising two children in Spain. Juan's Venezuelan family now lives across Chile, Norway, Germany, Spain, and the United States. Jordan's family lives in the United States. Grandparents visit once a year, twice if we are lucky, and one uncle has never met our children.
When family visits, our daughter asks her grandmother to invent stories on walks to school and at bedtime. Those stories disappear as soon as they end. When the visit ends, the storytelling ends too, replaced by scheduled video calls that put a phone in a child's hands.
Spoken Letter lets a relative tell a story from anywhere and preserve it in their real voice. The result is an MP3 that works on any device. We began with Yoto because our children use it: screen-free audio on a card the child controls, without an adult phone.
Our rule is simple: AI can support the storyteller, but it must not replace the family voice.
What humans and AI do
A relative, whom we call the Sender, records or uploads a rough memory. Gemini 2.5 Flash on Vertex AI finds the story inside it and creates a child-friendly draft. If the Sender does not know what to say, a voice interview asks questions and Gemini turns the answers into a script.
The Sender edits that script, reads it aloud, optionally adds music, and sends the recording.
The parent, or Owner, receives an AI-written synopsis and decides whether the story reaches the child by downloading the MP3 or sending it to a linked Yoto account. That delivery action is the approval. AI cannot cross this boundary.
Humans provide the memories, choose the words, perform the narration, and control delivery. Gemini structures memories, drafts and revises scripts, checks content, writes parent summaries, identifies vocabulary, suggests future stories, classifies themes, selects Yoto icons, and summarizes support conversations. Speech-to-Text transcribes audio. Lyria creates optional music. Cloud Run and FFmpeg process and export the finished recording.
We log every AI job by type, model, result, and duration. At the latest count (As of August 17, 2026 at 08:15 UTC), production logs contained 653 AI calls across 132 stories and 15 job types. Judges can inspect those records, API usage, and our public evidence dashboard.
AI also helps operate the business. An agent planned our paid launch, wrote ad copy, produced creative, built Meta campaigns, applied attribution tags, and launched a $600 campaign. Scheduled agents maintain cost, analytics, and production-health ledgers. Humans choose goals and budgets, review sensitive decisions, investigate failures, and remain accountable for every outward action.
Safety and human judgment
Children never touch the platform. There is no child account, sign-in, voice recording, child-to-AI chat, public feed, or stranger discovery. An Owner can add a child's first name and optional birthday month and day, without a birth year, and can edit or delete both. Replies are limited to an Owner-tapped Like or Love.
Voice Restoration serves living adults who have lost their speaking voice. It is the only case where AI narrates in a family member's voice. Every request receives human review. The speaker must consent in their own words, an operator builds the private voice, and the speaker approves a sample. Consent can be withdrawn at any time. We do not clone children, deceased people, or anyone other than the consenting speaker.
Human growth and potential
Spoken Letter is not primarily about generating stories. It is about making another human being more present in a child's life.
A grandparent can pass down a childhood memory. An uncle can explain where the family came from. A bilingual relative can tell the same story in another language. A child can hear unfamiliar words in a familiar voice. Stories that once depended on two people being in the same room can become part of an everyday bedtime routine.
That creates opportunities for language, imagination, cultural continuity, and intergenerational learning without asking the child to use another screen or interact with an AI system.
There is growth on the Sender's side too. Many people have memories worth sharing but do not consider themselves writers or storytellers. A blank page stops them. Gemini acts as scaffolding: it asks questions, finds structure in imperfect recollections, and helps someone turn "I don't know what to say" into a story only that person could tell.
Voice Restoration extends the same idea to families facing speech loss. Losing a speaking voice should not automatically mean losing the ability to narrate a story to someone you love. Here AI does something the human can no longer physically do, while consent and human judgment remain in control.
The technology is valuable precisely because the final experience does not feel like technology. To the child, it is Grandma telling a story.
Building the business
Our users changed the product. Older relatives struggled with microphone permissions, so recording and upload became equal paths. Mobile use dominated, so recording became full-screen with the countdown, microphone, and teleprompter together. We removed a formal approval screen because downloading or sending already expressed the parent's decision. Each correction removed friction between a person and the story they wanted to tell.
One production defect placed narration in only one ear. Automated tests passed because they checked bytes but could not listen. We repaired the pipeline, corrected 30 affected stories, and added headphone listening to release verification. It taught us where human senses remain essential.
The business launched July 7. During the program window it earned $252.23 in total revenue against $305.38 in expenses. July, our first full trading month, was cash-positive by $39.98. Observed gross margin is 87.8%, with production cost between $0.032 and $0.177 per story against a $2.99 minimum price. Early sales were organic. Paid acquisition began August 16.
What comes next
Our priority is finding families whose relatives have lost their voices and serving them carefully, a few at a time. Restored narration currently reads in English, so additional languages come next. We also plan simpler delivery to other audio devices, including Tonies and Alexa.
The goal remains unchanged: use AI to help preserve what technology cannot create on its own, the voice, memories, and presence of someone a child loves.
Built With
- claudecode
- cloud-run
- cloud-storage
- codex
- elevenlabs
- ffmpeg
- firebase
- firestore
- gemini
- google-analytics
- google-cloud
- lyria
- next.js
- node.js
- playwright
- react
- sentry
- speech-to-text
- stripe
- tailwindcss
- text-to-speech
- typescript
- vercel
- vertex-ai
- yoto


Log in or sign up for Devpost to join the conversation.