Inspiration

Translating a video is easy now. Publishing it safely is not.

A synthetic voice still needs to be authorized. The translated speech has to fit the timing of the original video. And if something goes wrong, the team needs a clear record of what happened, what was approved, and which output was finally published.

We built Toluva because we wanted video localization to feel like a reliable production workflow, not just another one-click AI demo.

What it does

Toluva turns a short English source video into a governed German edition.

Before anything is generated, the user confirms their upload rights and acknowledges that the localized track uses a disclosed synthetic stock voice. Toluva then:

  • transcribes the source video into timed segments

  • checks the transcript before spending voice-generation credits

  • translates the approved text into German

  • checks whether the requested voice, language, and purpose are authorized

  • generates German speech

  • measures how well each generated segment fits its original time slot

  • stops for human approval when a correction is required

  • generates captions and composes the final localized video

  • records every important file, decision, attempt, and approval in Backblaze B2

The result is a localized video with German audio and captions, together with an inspectable history showing how it was produced.

How we built it

The product interface is built with Next.js, React, TypeScript, Tailwind CSS, and Vinext, and is hosted publicly with OpenAI Sites.

The long-running media pipeline is handled by a Python worker running separately on a VPS. Backblaze B2 acts as both the durable job queue and the system of record, so jobs can survive browser refreshes, worker restarts, and interrupted processing.

Genblaze makes the individual stages inspectable instead of hiding the workflow behind one dubbing API. The verified pipeline uses:

Faster Whisper for timed English transcription

Argos Translate for English-to-German translation

ElevenLabs Flash v2.5 for synthetic speech

FFmpeg for audio assembly, captions, and final video composition

Backblaze B2 for source files, transcripts, translations, speech attempts, captions, approvals, manifests, and final outputs

Every completed generative stage has a Genblaze manifest, and important stored assets are independently checked against their recorded SHA-256 hashes.

Challenges we faced

The hardest part was preventing translated speech from breaking the timing of the original video.

German speech can be longer or shorter than the English source slot. Toluva measures the real generated audio instead of estimating from the text. When the result does not fit safely, the job stops and requests a specific, hash-bound correction. A worker restart can then continue from the last verified checkpoint without repeating completed provider calls.

We also had to handle imperfect transcription. During testing, Whisper added a suspicious trailing fragment to one recording. Instead of silently translating it, Toluva blocked the job before ElevenLabs was called and allowed an immutable human correction to resume the same job.

Another challenge was keeping Backblaze B2 central to the workflow without wasting storage transactions. We moved to bounded queue scans, append-only job events, reusable checkpoints, and a single worker so the system remains inspectable without accidentally duplicating paid work.

What we learned

We learned that the difficult part of generative media is not simply calling a model. It is deciding when that model is allowed to run, measuring whether its output is usable, recovering safely when something fails, and preserving enough evidence for another person to understand the result.

Genblaze and Backblaze B2 helped us turn those concerns into visible parts of the product rather than hidden implementation details.

We also learned that a smaller, verified workflow is more credible than claiming broad support that has not been tested. That is why the submitted production lane focuses on English-to-German localization.

What’s next

Next, we want to add more languages only after testing each voice and model through the complete authorization and timing workflow.

We also plan to add collaborative review, stronger terminology management, downloadable evidence bundles, provider fallbacks, and an atomic queue that can safely support multiple workers.

Built With

Share this project:

Updates

Submission history