Inspiration

Me and my teammate, we love books, but since we're busy with uni stuff, we wanted to have audio-books instead, so that we can listen to them while walking or doing other tasks. But the problem was that traditional text-to-speech solutions sound robotic and monotonous, and completely fail to capture the nuance of storytelling. We thought of a platform where AI could act as a Director by analyzing characters in any book and assigning them distinct, appropriate voices that match their personality, age, gender, and accent, being AWARE of the context in the book. We wanted it to feel very genuine.

What it does

  • It transforms PDFs into multi-voice audiobooks with AI-powered character casting:
  • Gemini AI identifies characters & personality traits
  • Smart Casting matches voices using weighted similarity (gender, age, pitch, accent)
  • Real-time Streaming with synchronized transcript highlighting
  • Prefetching and chunking so that you don't have to wait 20 minutes for it to generate the whole book.

How we built it

AI Pipeline: Python, FastAPI, Gemini Flash Voice Casting: Custom weighted Euclidean distance algorithm Audio: ElevenLabs Turbo v2.5 streaming Mobile: React Native, Expo, TypeScript Web (More of a dashboard for testing, but we have cross platform plans): Vite, React, TailwindCSS Database: MongoDB

Challenges we ran into

  • Voice Consistency: Ensuring that the same character has the same voice throughout
  • Real-time Sync: Transcript highlighting with streamed audio (and ensuring that the transcript exactly matches to avoid ghost dialogue!)
  • Character recognition: standardizing names ("The king" vs his real name)

Accomplishments that we're proud of

  • Multi-dimensional voice matching (gender, pitch, roughness, accent)
  • Seamless audio streaming with live transcript sync
  • How the audiobook actually sounds (Doesn't feel AI at all), largely in thanks to elevenlabs

What we learned

  • Gemini structured JSON output with Pydantic schemas
  • Designing similarity algorithms for voice casting
  • Cross-platform media players with synchronized text

What's next for audiolore

  • User auth
  • Offline downloads
  • Global library, with proceeds going towards authors
  • Cloud sync with google firestore
  • Voice cloning
  • Complete voice navigation
  • The Librarian: A mode where you can interrupt the story, ask, "Wait, who is this character again?" and the AI answers in context.

Built With

Share this project:

Updates