Inspiration
While building content with Google Flow, I hit a frustrating wall — once a video is generated, you're completely locked into the voice it came with. No way to swap the narrator, change the tone, or try a different voice without regenerating everything from scratch.
That one limitation inspired VoiceSwap.AI. I wanted to prove that intelligent voice replacement — one that actually preserves the soul of the original delivery — is possible today, using tools Google already provides.
What it does
VoiceSwap.AI lets you upload any video with a single speaker and replace the voice with one of 6 Google Cloud TTS Chirp3-HD voices — while preserving the original emotion, tone, pacing and timing.
The key insight: Gemini 2.5 Flash doesn't just transcribe — it directs.
Instead of dumping plain text into a TTS engine, Gemini watches the video, understands how the person is speaking — the urgency, the warmth, the pacing — and generates a complete Voice Direction JSON with emotion markers, pace controls and emphasis cues. That intelligence drives the synthesis. The voice changes. The soul stays.
How I built it
Pipeline (end to end):
- User uploads MP4/MOV → stored in a UUID session directory on Cloud Run
- Video sent to Gemini 2.5 Flash via multimodal API → returns transcript + voice direction JSON
- User selects a Chirp3-HD voice from the UI
- Voice direction is used to construct a synthesis request → Google Cloud TTS v1beta1 returns synthesized MP3
- ffmpeg time-stretches the audio using the
atempofilter to exactly match the original video duration - Stretched audio merged with muted original video → final MP4 ready to download
Stack:
- Frontend: Next.js 15 + Tailwind CSS + Framer Motion → deployed on Vercel
- Backend: FastAPI + Python 3.11 → deployed on Google Cloud Run
- AI: Gemini 2.5 Flash via Google GenAI SDK
- Voice: Google Cloud TTS Chirp3-HD (v1beta1)
- Processing: ffmpeg (atempo filter for audio sync)
- Infra: Google Cloud Artifact Registry + Cloud Run
Challenges I ran into
1. Chirp3-HD doesn't accept SSML
Chirp3-HD voices require the v1beta1 endpoint and reject SSML input entirely. I had to build a fallback that strips SSML to clean text while preserving the performance intent through punctuation and the speaking_rate parameter.
2. Audio sync is non-trivial
TTS produces audio of unpredictable length. Simply cutting with -shortest in ffmpeg produced jarring results. The fix was ffmpeg's atempo filter — time-stretching the synthesized audio to exactly match the original video duration while preserving pitch.
3. Making Gemini the brain, not just a transcriber The hardest design challenge was prompting Gemini to go beyond transcription and act as a genuine creative director — producing structured, actionable voice direction from video alone. Getting consistent, well-structured JSON output across different video types took significant prompt iteration.
Accomplishments that I'm proud of
- Built a fully working end-to-end voice replacement pipeline in under a week, solo
- Gemini 2.5 Flash acting as a genuine AI Voice Director — not just a transcription tool
- Perfect audio/video sync using ffmpeg atempo stretching
- 100% Google ecosystem — Gemini + Cloud TTS + Cloud Run + Artifact Registry
- Clean, production-grade UI that actually looks like a real product
What I learned
- Gemini 2.5 Flash's multimodal capabilities go far deeper than transcription — it genuinely understands emotional nuance in speech
- Chirp3-HD voices are significantly more expressive than Neural2 but come with important constraints
- Stateless session design on Cloud Run makes horizontal scaling trivial
- Audio/video synchronization is a harder problem than it looks
What's next for VoiceSwap.AI
- Native Google Flow integration — pre-select a voice persona before generating, or swap post-generation without losing visuals
- Multi-voice export — generate the same video in all 6 voices simultaneously for A/B testing
- Lip sync — match the new voice to the speaker's lip movements
- Multi-speaker support — detect and replace individual speakers in multi-person videos
- SSML support — switch to Journey or Neural2 voices for full prosody control when Chirp3-HD limitations are a constraint
Built With
- docker
- fastapi
- ffmpeg
- google-cloud-artifact-registry
- google-cloud-run
- google-cloud-tts
- google-gemini-2.5-flash
- google-genai-sdk
- next.js
- python
- tailwind
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.