Inspiration
Translating a video is easy now. Publishing it safely is not.
A synthetic voice still needs to be authorized. The translated speech has to fit the timing of the original video. And if something goes wrong, the team needs a clear record of what happened, what was approved, and which output was finally published.
We built Toluva because we wanted video localization to feel like a reliable production workflow, not just another one-click AI demo.
What it does
Toluva turns a short English source video into a governed German edition.
Before anything is generated, the user confirms their upload rights and acknowledges that the localized track uses a disclosed synthetic stock voice. Toluva then:
transcribes the source video into timed segments
checks the transcript before spending voice-generation credits
translates the approved text into German
checks whether the requested voice, language, and purpose are authorized
generates German speech
measures how well each generated segment fits its original time slot
stops for human approval when a correction is required
generates captions and composes the final localized video
records every important file, decision, attempt, and approval in Backblaze B2
The result is a localized video with German audio and captions, together with an inspectable history showing how it was produced.
How we built it
The product interface is built with Next.js, React, TypeScript, Tailwind CSS, and Vinext, and is hosted publicly with OpenAI Sites.
The long-running media pipeline is handled by a Python worker running separately on a VPS. Backblaze B2 acts as both the durable job queue and the system of record, so jobs can survive browser refreshes, worker restarts, and interrupted processing.
Genblaze makes the individual stages inspectable instead of hiding the workflow behind one dubbing API. The verified pipeline uses:
Faster Whisper for timed English transcription
Argos Translate for English-to-German translation
ElevenLabs Flash v2.5 for synthetic speech
FFmpeg for audio assembly, captions, and final video composition
Backblaze B2 for source files, transcripts, translations, speech attempts, captions, approvals, manifests, and final outputs
Every completed generative stage has a Genblaze manifest, and important stored assets are independently checked against their recorded SHA-256 hashes.
Challenges we faced
The hardest part was preventing translated speech from breaking the timing of the original video.
German speech can be longer or shorter than the English source slot. Toluva measures the real generated audio instead of estimating from the text. When the result does not fit safely, the job stops and requests a specific, hash-bound correction. A worker restart can then continue from the last verified checkpoint without repeating completed provider calls.
We also had to handle imperfect transcription. During testing, Whisper added a suspicious trailing fragment to one recording. Instead of silently translating it, Toluva blocked the job before ElevenLabs was called and allowed an immutable human correction to resume the same job.
Another challenge was keeping Backblaze B2 central to the workflow without wasting storage transactions. We moved to bounded queue scans, append-only job events, reusable checkpoints, and a single worker so the system remains inspectable without accidentally duplicating paid work.
What we learned
We learned that the difficult part of generative media is not simply calling a model. It is deciding when that model is allowed to run, measuring whether its output is usable, recovering safely when something fails, and preserving enough evidence for another person to understand the result.
Genblaze and Backblaze B2 helped us turn those concerns into visible parts of the product rather than hidden implementation details.
We also learned that a smaller, verified workflow is more credible than claiming broad support that has not been tested. That is why the submitted production lane focuses on English-to-German localization.
What’s next
Next, we want to add more languages only after testing each voice and model through the complete authorization and timing workflow.
We also plan to add collaborative review, stronger terminology management, downloadable evidence bundles, provider fallbacks, and an atomic queue that can safely support multiple workers.
Built With
- ai
- assemblyai
- b2
- backblaze
- css
- docker
- elevenlabs
- fastapi
- ffmpeg
- genblaze
- generative
- gpt-4.1
- localization
- next.js
- openai
- postgresql
- python
- react
- redis
- speech-to-text
- tailwind
- text-to-speech
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.