Inspiration

While building content with Google Flow, I hit a frustrating wall — once a video is generated, you're completely locked into the voice it came with. No way to swap the narrator, change the tone, or try a different voice without regenerating everything from scratch.

That one limitation inspired VoiceSwap.AI. I wanted to prove that intelligent voice replacement — one that actually preserves the soul of the original delivery — is possible today, using tools Google already provides.

What it does

VoiceSwap.AI lets you upload any video with a single speaker and replace the voice with one of 6 Google Cloud TTS Chirp3-HD voices — while preserving the original emotion, tone, pacing and timing.

The key insight: Gemini 2.5 Flash doesn't just transcribe — it directs.

Instead of dumping plain text into a TTS engine, Gemini watches the video, understands how the person is speaking — the urgency, the warmth, the pacing — and generates a complete Voice Direction JSON with emotion markers, pace controls and emphasis cues. That intelligence drives the synthesis. The voice changes. The soul stays.

How I built it

Pipeline (end to end):

  1. User uploads MP4/MOV → stored in a UUID session directory on Cloud Run
  2. Video sent to Gemini 2.5 Flash via multimodal API → returns transcript + voice direction JSON
  3. User selects a Chirp3-HD voice from the UI
  4. Voice direction is used to construct a synthesis request → Google Cloud TTS v1beta1 returns synthesized MP3
  5. ffmpeg time-stretches the audio using the atempo filter to exactly match the original video duration
  6. Stretched audio merged with muted original video → final MP4 ready to download

Stack:

  • Frontend: Next.js 15 + Tailwind CSS + Framer Motion → deployed on Vercel
  • Backend: FastAPI + Python 3.11 → deployed on Google Cloud Run
  • AI: Gemini 2.5 Flash via Google GenAI SDK
  • Voice: Google Cloud TTS Chirp3-HD (v1beta1)
  • Processing: ffmpeg (atempo filter for audio sync)
  • Infra: Google Cloud Artifact Registry + Cloud Run

Challenges I ran into

1. Chirp3-HD doesn't accept SSML Chirp3-HD voices require the v1beta1 endpoint and reject SSML input entirely. I had to build a fallback that strips SSML to clean text while preserving the performance intent through punctuation and the speaking_rate parameter.

2. Audio sync is non-trivial TTS produces audio of unpredictable length. Simply cutting with -shortest in ffmpeg produced jarring results. The fix was ffmpeg's atempo filter — time-stretching the synthesized audio to exactly match the original video duration while preserving pitch.

3. Making Gemini the brain, not just a transcriber The hardest design challenge was prompting Gemini to go beyond transcription and act as a genuine creative director — producing structured, actionable voice direction from video alone. Getting consistent, well-structured JSON output across different video types took significant prompt iteration.

Accomplishments that I'm proud of

  • Built a fully working end-to-end voice replacement pipeline in under a week, solo
  • Gemini 2.5 Flash acting as a genuine AI Voice Director — not just a transcription tool
  • Perfect audio/video sync using ffmpeg atempo stretching
  • 100% Google ecosystem — Gemini + Cloud TTS + Cloud Run + Artifact Registry
  • Clean, production-grade UI that actually looks like a real product

What I learned

  • Gemini 2.5 Flash's multimodal capabilities go far deeper than transcription — it genuinely understands emotional nuance in speech
  • Chirp3-HD voices are significantly more expressive than Neural2 but come with important constraints
  • Stateless session design on Cloud Run makes horizontal scaling trivial
  • Audio/video synchronization is a harder problem than it looks

What's next for VoiceSwap.AI

  • Native Google Flow integration — pre-select a voice persona before generating, or swap post-generation without losing visuals
  • Multi-voice export — generate the same video in all 6 voices simultaneously for A/B testing
  • Lip sync — match the new voice to the speaker's lip movements
  • Multi-speaker support — detect and replace individual speakers in multi-person videos
  • SSML support — switch to Journey or Neural2 voices for full prosody control when Chirp3-HD limitations are a constraint

Built With

  • docker
  • fastapi
  • ffmpeg
  • google-cloud-artifact-registry
  • google-cloud-run
  • google-cloud-tts
  • google-gemini-2.5-flash
  • google-genai-sdk
  • next.js
  • python
  • tailwind
  • typescript
  • vercel
Share this project:

Updates

Submission history