Inspiration
Text-to-speech has been around for a long time, and modern systems can produce remarkably natural voices. The problem is that most of them still sound like one continuous, polished speaker. They are good voices, but often a little plain-bagel.
I wanted computers to sound more like computers: a voice emerging from a swarm of recordings, broadcasts, noise, and data. Each word could come from somewhere different, yet together they would form one recognizable transmission.
Maybe I just like Bumblebee’s fragmented radio voice.
Unfortunately, I do not have Hasbro licensing money, so this project is called FrankenVoice, by Frankenmint.
FrankenVoice explores whether a machine voice can be assembled rather than spoken—one independent fragment at a time.
What it does
FrankenVoice is a cloud-deployed composite speech application that builds new sentences from a growing corpus of individual word recordings.
A user can import a YouTube recording they have permission to use. FrankenVoice extracts the audio, transcribes it with word-level timestamps, cuts the recording into reusable word clips, and adds those clips to a persistent shared corpus.
When the user enters a new sentence, FrankenVoice independently selects a clip for each word and stitches the fragments into a new WAV file. The same sentence can sound different on repeated generations because different available recordings may be selected.
For example, the grammatically valid sentence:
Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo.
could be generated using eight different recordings of the word “buffalo” rather than replaying the same clip eight times.
The application includes three primary interfaces:
- Autopilot uses Qwen to inspect the user’s goal and corpus coverage, create a constrained workflow, estimate external actions, and wait for human approval before importing media or using paid cloud enrichment.
- Composer allows the user to directly generate a sentence and adjust variation, source diversity, pauses, filtering, speed, and glitch effects.
- Reader prepares longer Markdown or conversational text and generates it in manageable chunks.
Qwen does not generate the completed sentence as one continuous TTS response. Qwen Cloud is used for workflow planning, optional timestamped transcription, and isolated missing-word enrichment. FrankenVoice’s own fragment engine assembles the final sentence word by word.
FrankenVoice also exposes a direct generation API and an OpenAI-compatible speech endpoint, making it possible to connect the composite voice to another application or AI agent.
The public deployment is available at:
https://frankenvoice.frankenmint.com
How I built it
The project began as a smaller local experiment: ingest audio, transcribe it, cut out individual words, and determine whether those fragments could be assembled into understandable speech.
I chose a TypeScript and React frontend with a Python FastAPI backend. FFmpeg and yt-dlp handle media ingestion, Faster Whisper performs local word-level transcription for imported YouTube sources, SQLite stores source and fragment metadata, and the audio pipeline selects, adjusts, filters, and joins the resulting WAV clips.
As the project evolved toward the hackathon’s Autopilot Agent track, I added Qwen Cloud as the planning and enrichment layer.
The agent receives:
- an ambiguous creative objective;
- the final text to generate;
- optional source URLs;
- and the current shared-corpus coverage.
Qwen returns a constrained JSON tool plan. FrankenVoice persists that plan and displays a human checkpoint before allowing source downloads or paid cloud enrichment. Once approved, the executor imports permitted sources, checks coverage again, adds missing isolated words when necessary, generates the composite speech, and saves the final WAV, event log, and execution report.
For production deployment, I containerized the frontend and backend with Docker Compose and deployed them on Alibaba Cloud ECS. The public path is:
frankenvoice.frankenmint.com
→ Cloudflare
→ Alibaba Cloud ECS host Nginx
→ React/Nginx container
→ FastAPI Autopilot container
→ Alibaba Cloud Model Studio / DashScope
The backend stores its SQLite corpus, processed sources, fragment files, and Autopilot run state in a persistent Docker volume. The browser accesses the frontend and API through one public origin.
I used AI development tools throughout the project for planning, code review, debugging, and implementation support. I treated their output as proposed engineering work rather than accepted truth: I reviewed changes, tested behavior, corrected false assumptions, and maintained the final architecture and deployment myself.
Challenges I ran into
One of my earliest mistakes was allowing the implementation, the written specification, and the code being reviewed by different models to drift out of sync. A model would evaluate an older snapshot while I was already changing the active repository, which caused duplicated work and recommendations that no longer applied.
I also learned that prototyping with an LLM and extending an existing system with one are very different tasks. Generating an isolated feature is easy. Correctly integrating that feature into an evolving codebase requires careful attention to state, interfaces, removed behavior, error handling, and dependencies.
I still had to review the code line by line and frequently determine what should remain, what should be replaced, and whether a proposed feature was actually connected to the rest of the application.
The media pipeline created another major challenge. YouTube began rejecting unauthenticated requests and requiring proof that the importer was not a bot. I had to implement authenticated cookie handling without exposing those credentials in the repository.
Modern YouTube extraction also required a supported JavaScript runtime and challenge-solver package, so the production backend needed Deno and the full yt-dlp EJS dependencies.
Large source files exposed the resource limits of the ECS instance. A long recording could consume enough CPU and memory during transcription to make the server unresponsive. I added a hard ten-minute source limit, disabled playlist expansion, and rejected livestreams before expensive audio processing begins.
The most difficult product problem was audio cohesion. Individual words cut from different recordings have different microphones, room noise, pitch, timing, loudness, and emotional delivery. Even when the sentence is understandable, transitions can sound abrupt.
I also discovered that Qwen-generated vocabulary clips were initially being selected more often than my original recordings because their stored confidence score was higher. I changed the selection rules so that original source recordings are preferred whenever they exist, Qwen-derived clips are used only for missing vocabulary, and eSpeak remains the final fallback.
The prototype now proves the architecture, but smooth fragment-level audio remains a serious engineering problem rather than a solved one.
Accomplishments that I’m proud of
I am proud that I completed and publicly deployed an end-to-end agentic audio system as a solo developer within the hackathon time frame.
FrankenVoice can now:
- ingest approved YouTube recordings;
- transcribe them into timestamped words;
- cut and persist reusable audio fragments;
- measure target-sentence corpus coverage;
- use Qwen to generate a constrained workflow plan;
- stop at a human approval checkpoint;
- execute approved downloads and cloud enrichment;
- persist its run state and event history;
- assemble a final sentence from independent clips;
- expose the result through both its web interface and API;
- and run publicly on Alibaba Cloud ECS.
I am especially proud that the final sentence is genuinely assembled by FrankenVoice. The project does not quietly replace its central idea with a whole-sentence cloud TTS response when things become difficult.
Qwen expands the toolbox, but the fragment engine remains responsible for the final voice.
I am also proud of the project’s visual identity. The repository’s stereogram contains the FrankenVoice name and robot image hidden within the noise, which mirrors the central idea of the project: look through the noise and find the voice inside it.
What I learned
The clearest lesson is that TTS is hard, and composite TTS is hard in a different way.
A conventional neural TTS system tries to produce one coherent performance. FrankenVoice instead attempts to make hundreds or thousands of unrelated performances feel like one changing speaker.
That introduces problems involving timing, silence, loudness, pitch, source continuity, and pronunciation before the system even begins dealing with vocabulary coverage.
I learned that the size of the corpus matters, but corpus quality and composition matter just as much. A large number of random words does not guarantee a good sentence.
The system benefits from multiple clean variants, consistent recording conditions, accurate word boundaries, and enough coverage to avoid fallback voices.
I also learned that agent reliability comes from the harness surrounding the model. The useful part is not simply asking Qwen what to do.
The useful part is constraining its available tools, validating its structured response, persisting the plan, separating free local actions from paid or external actions, requiring human approval, and recording the outcome.
Finally, I learned that AI coding assistance does not eliminate software engineering judgment. It can dramatically accelerate investigation and implementation, but context management, integration, verification, and product decisions still belong to the developer.
What’s next for FrankenVoice
The immediate priority is improving audio cohesion.
Planned improvements include:
- per-fragment loudness normalization;
- automatic silence trimming;
- short fades and crossfades between clips;
- stronger preference for consecutive words from the same recording;
- improved pitch and duration matching;
- better handling of contractions and alternate pronunciations;
- visible per-word provenance;
- per-word preview, reroll, and locking;
- waveform editing;
- and larger rights-cleared source libraries.
I would also like to turn voice configurations into reusable preset objects. A project could maintain a library of different composite voices—robot transmissions, damaged radios, fantasy creatures, game characters, announcers, or user-created voices—and select between them dynamically.
That opens several interesting possibilities:
- local LLM narration;
- personalized assistive text-to-speech;
- dynamic non-player-character voices;
- interactive fiction and world-building;
- game dialogue assembled from licensed character libraries;
- custom voices built from a user’s own recordings;
- and automated narration that intentionally sounds synthetic rather than attempting to imitate a perfectly natural speaker.
The audio is still rough, but the central idea works: as the corpus grows, a voice begins to emerge from the fragments.
And yes, the stereogram is still cool too.
Log in or sign up for Devpost to join the conversation.