-
-
Landing
-
Sign Up
-
Studio Dashboard
-
The Crew at Work
-
Create a new Drama
-
Script Editor page with script quality score
-
Casting Bible
-
Character Relationship Graph
-
Storyboard Generation
-
Story Relationship Graph & Camera Plan
-
Video Generation
-
Agent Crew modal with Workflow Graph
-
Export Editor
-
Video Editing with Live Cost Update
-
User Usage & Analytics
Inspiration
AI short dramas are everywhere now, on Instagram Reels, TikTok, and Facebook. Strange fruit stories, brainrot micro clips, and even actual studio productions quietly using AI to create entire scenes. I kept asking myself one question: could an AI agent go beyond generating individual assets and handle the complete production workflow, from an initial idea to a finished drama episode, with no human manually connecting every step?
After exploring the existing AI drama workflows, I realized the biggest challenge was not content generation itself. Most pipelines still rely on separate tools for scriptwriting, image generation, voice creation, video generation, and editing, with humans responsible for maintaining consistency between each stage. This creates real problems: characters lose their identity across shots, generated scenes misunderstand the intended action, stories lose continuity, and video generation costs escalate quickly without careful planning.
Writing a script is no longer the hardest part. Modern AI models can already create compelling stories. The real challenge is everything after the script: keeping one face consistent across forty shots, ensuring the right action is generated, and preventing the budget from spiraling. That gap is exactly why there is still no "type one line or import whole script, get a finished drama" button.
Rexgent was built to close it: an autonomous AI showrunner that runs the entire production pipeline while optimizing quality, consistency, and cost. We built it around solving those three problems, hallucination, consistency, and cost, so creativity can finally ride on top of a reliable production pipeline.
What it does
Rexgent is an autonomous AI showrunner. You type a one line premise or import your own screenplay, pick an episode count, video format and visual style, set a spend cap then pick your output language (English or Chinese) and finally choose Full Auto (it runs the whole production untouched) or guided (you review and approve each stage). The output language is your choice, independent of your input. Every staging rule is CJK-aware, so a Chinese production boards, stages, and renders with the same guarantees as an English one.
A LangGraph state machine runs the pipeline:
Develops, then Writes.
A one-sentence premise is a situation, not a story, so a story development step (Qwen-Max) first invents the dramatic spine: the central conflict, a charged relationship, the buried secret, and the mid-story turn that drives the cliffhanger. The screenwriter writes cliffhanger-structured episodes from that brief, then judges its own draft on 8 axes (including "would the first 3 seconds stop a scrolling viewer?") and rewrites weak ones using the judge's own critique. You type one sentence; the agent does the story development work. Already have a screenplay? Import it instead: a structurer parses the file (PDF, DOCX, TXT, Markdown, or Fountain) into scenes and dialogue, merges scenes that share one location, and your script rides the exact same production pipeline from casting onward, in English or Chinese.Casts the Production Bible.
Every character gets identity and per-scene costume plates image-edited from one locked face, rendered in the drama's chosen visual style so an anime style drama gets anime style plates. Pets and creatures are cast too, with their species and real-world scale pinned. You can upload a real photo for any character or let the agent invent a face. Each finished plate is probed against the video model's own reference-input check at casting time.Storyboards with an AI Director.
An AI Director plans the coverage before a single prompt is written: it gives each shot a purpose, varies the shot size, lens and camera move so a scene is not forty identical mediums, and breaks up dialogue with reaction shots and inserts instead of pointing the camera at whoever happens to be talking. Every shot then carries each character's screen position, facing, eye direction and posture, with the 180° rule enforced by a stage map, and actions that change scene states split into single-state shots. A deterministic staging engine then threads world state across the scene. An LLM cast audit double-checks every shot's roster, removing characters who are talked about but not visibly present. You can open any scene as an interactive top-down camera plan, a story map of who appears where, and a scene-by-scene flow. A set dresser agent tracks prop state, so a vase shattered stays shattered in corresponding shots.Fits the Plan to your Money.
Each shot is classified by an identity role before it renders: an establishing shot locks every face from the full plate stack, a character's entrance adds the newcomer's references, an angle change re-locks identity instead of dragging the old composition into the new framing, and a quiet continuation carries the previous frame inside its reference stack. Anything with an identity or a line renders on HappyHorse, the character model that holds faces and speaks lines natively, territory it earned by beating Wan in scored head-to-head runs; the few shots with nothing to lock, silent same-angle continuations and people-free scenery, render on Wan 2.7, extending straight from the previous frame so the room survives the cut. Every shot is scored for importance, and anything over the cap defers, with the agent naming the smallest cap that fits. Nothing spends silently: every paid button opens an itemized, per-model receipt with a live total before you confirm.Renders with Receipts, and Lets you Fix takes.
Each clip is conditioned on an explicit reference stack (identity, costume, previous frame, location, style) and validated by real ArcFace face-matching plus a vision check on outfit and background drift. Weak clips are flagged, never silently shipped, and a clip that comes back catastrophically wrong, a black frame or a faceless render, is caught and regenerated automatically. A prompt refused by the platform's content moderation is rewritten by a sanitizer model that keeps the shot's meaning while removing the trigger, then retried, so one flagged word never kills a render. In the editor you can flag any take, say what's wrong in plain words, and regenerate just that one shot.Voices, Cuts, and Exports itself.
On a talking shot, HappyHorse speaks the scripted line as it renders, so the voice, the mouth, and the emotion come from a single pass with nothing to keep in sync. I also built a second audio mode: a TTS overlay that gives each character one fixed voice from the Qwen-TTS catalog and repaces it word by word onto the on-screen mouth using ASR word timestamps. In my side-by-side tests the native voice still wins, it matches the mouth exactly and sounds more real, so native ships as the default and the overlay stays an option for when a perfectly stable voice matters more. Each line lands on the exact shot that speaks it, music ducks under speech, captions burn in, endings hold and fade.
Throughout, a live crew dock shows every stage expanding into its real tools next to a cost meter counting every token. A Usage & Analytics page turns the ledger into proof: routing efficiency against the all-premium counterfactual, per-model receipts, and reliability health.
How I built it
AI Production Pipeline
Rexgent is built as an autonomous AI showrunner powered by 17 specialized models on Qwen Cloud, each routed to the task it performs best across the production workflow.
Story & Reasoning:
Qwen-Max develops story structure and writes episodes, Qwen-Plus critiques and evaluates scripts, Qwen-Flash handles lightweight deterministic tasks, while Qwen3-VL-Plus scores visual continuity and narrates each clip's final frame so the next shot knows exactly where the motion left off and Qwen-VL-Max reads uploaded reference photos.Visual Production Bible:
Wan-T2I and Qwen-Image-Edit-Max generate the foundation of each production, including character identity references, costume plates, and world assets.Video Generation:
The two video models are split by what each shot must hold. HappyHorse 1.1 carries every shot with an identity or a line: establishing shots, entrances, angle changes and all dialogue render through its R2V reference conditioning, which locks each face and outfit from the production bible and speaks each line natively while it lip-syncs its own mouth. It earned that territory by beating Wan in scored head-to-head runs on real dramas. Wan 2.7 renders the shots with nothing to lock: silent same-angle continuations and people-free scenery, extending straight from the previous frame so the room survives the cut. HappyHorse-Video-Edit powers the fix-a-take editing loop, patching a flagged clip from user feedback.Native Audio, with a TTS Overlay Mode:
Dialogue is voiced by the character model itself as each talking shot renders, keeping the mouth and the voice generated together instead of stitched afterward. Because the native voice can rarely drift between clips, I also built a full voice stack as an optional overlay mode: casting writes each character a voice brief from their age and personality (Qwen-Voice-Design), renders it as a bespoke timbre (Qwen3-TTS-VD), can clone an uploaded voice (Qwen-Voice-Enrollment + Qwen3-TTS-VC-Realtime), or falls back to the official 45-voice preset catalog (Qwen3-TTS-Flash).The chosen line is repaced word by word onto the native mouth using Fun-ASR and Qwen3-ASR-Flash word timestamps, with the clip's own music kept as a quiet bed only when ASR proves no hallucinated dialogue hides in it. Side-by-side tests still favor the native pass on realism and exact sync, so it is the default and the overlay remains a switchable mode.
Technology Stack
Backend:
FastAPI + Celery workers + Socket.IO for pipeline orchestration, generation jobs, and real-time progress updates.Frontend:
Next.js 14 dashboard with React, Tailwind CSS, React Query, and Zustand, where D3 powers the interactive story maps and camera plans, and a Socket.IO client streams live generation progress into the crew dock.Data & Memory:
PostgreSQL + pgvector stores character face embeddings computed by InsightFace (ArcFace), Redis manages task queues and events, and Neo4j maintains story and production knowledge graphs.Media Pipeline:
FFmpeg handles video composition and export, while Alibaba Cloud OSS stores generated assets including character references, clips, and final episodes. The system is deployed using Docker Compose on Alibaba Cloud ECS.Agent Infrastructure:
LangGraph orchestrates the write → judge → extract characters → storyboard → cast plates → budget → render → export state machine, with a self-correction loop where a rejected draft is rewritten from the judge's own critique and a rewrite that scores lower is discarded for the original. A two-pass Director Engine plans cinematic coverage before the prompt compiler builds each shot, and 7 custom MCP tools built on the official Model Context Protocol SDK do the specialized work:- Narrative judge
- Plot gap detector
- Ending engine
- Token optimizer
- Scene prompt craft
- Set dresser
- Consistency guard
These tools allow the AI showrunner to plan, generate, validate, and refine an entire drama production pipeline autonomously.
Challenges I ran into
Building Under a $40 Cloud Credit Budget:
Our entire budget was only $40 of Qwen Cloud credit, which is quite little for a project like this, and every test run to check the quality of a generated drama spent real money. So I built budget optimization into the core: a shot-scoring token optimizer that fits the plan to a cap, routing each shot to the model that fits its job, deferring the least important shots when the plan runs over the cap, and a cost ledger tracking every call.Hallucinations in Video Generation:
Some drift is expected, but the number of distinct failures in video generation was the real surprise, from faces changing between shots to crowds cheering in tragic moments to people appearing out of nowhere. I identified and categorized each failure mode, then designed a dedicated defense strategy for each one.Inconsistent Audio and Dialogue:
Early on, separate text-to-speech never matched what was happening on screen: the mouth and the voice disagreed, and fixed-length shots could not hold a line. The fix was letting the character model speak the scripted line itself and lip-sync its own mouth as the shot renders, so the voice, the timing, and the emotion all come from a single pass. That pass has one honest weakness, the same character's voice can occasionally drift slightly between clips, so I engineered a TTS overlay as a second mode: one fixed voice per character, placed at the measured speech onset and repaced word by word onto the on-screen mouth using ASR word timestamps. When I tested both on real dramas, the native voice still sounded more real and matched the mouth exactly, so it stays the default, and the overlay is the switch I flip if voice stability ever matters more than realism.Maintaining Character and Scene Consistency Across Shots:
Even with locked faces and location plates, scenes still drifted: rooms changed between angles, props appeared or disappeared, and characters teleported across barriers or replayed actions they had already done. We solved this by chaining shots through previous-frame references, anchoring each scene to its opening wide shot, tracking prop states with a set dresser agent, and building a deterministic staging engine that threads positions, held objects, and absences from shot to shot before any prompt is written. Character identity is maintained through reference stacks and ArcFace verification.Making the Pipeline Truly Bilingual:
Supporting Chinese was not a translation job, it was an engineering one. Word-boundary regex silently never matches between CJK characters, so cast detection, staging rules, and location matching that worked perfectly in English all failed in Chinese without throwing a single error. Every text rule in the pipeline was rebuilt with CJK-aware boundaries, Chinese spatial grammar, Chinese numerals for ages, and Chinese film vocabulary in the prompts, so a Chinese drama now gets the exact same continuity guarantees as an English one.
Accomplishments that I'm proud of
Fully Autonomous Production:
A complete run from one premise to finished, voiced, and captioned episodes, with the agent capable of handling the entire workflow autonomously while still allowing human review and approval at any stage.Story Development from One Sentence:
A dedicated development stage invents the central conflict, the charged relationship, and the mid-story turn before a word of script is written, so a thin premise becomes a real story instead of padding.Making AI Video Generation Reliable:
Every predictable failure mode of video generation received a dedicated defense. Each clip is backed by reference plates, face/outfit/background validation scores, and generation evidence, while the shared placement pipeline keeps editing and export perfectly aligned.Model Roles Earned by Measurement:
I planned a two-model video split, Wan for visuals and HappyHorse for characters, then ran them head to head on real dramas. HappyHorse held faces better on every identity job, so it took every shot with a face or a line, while Wan kept what it genuinely wins: silent continuations, people-free scenery, and the image side of the production bible. In Rexgent a model keeps a job only while it keeps winning it.Provable Cost Efficiency:
The analytics page turns the ledger into proof, showing the share of language work that ran on cheaper model tiers, the Wan and HappyHorse split across the shots, and the exact dollars each routing decision saved against an all-premium pipeline.One Pipeline, Two Languages, 22 Looks:
The same premise renders as photoreal cinema, Ghibli softness, pixel art, or claymation, in English or Chinese, in vertical or widescreen, with every continuity guarantee intact. The 22 styles on the landing page are not mockups: each sample is a real clip rendered by the pipeline from the same scene, same character reference, same prompt, only the style changed.End-to-End AI Production Orchestration
An AI Director plus seven custom MCP tools coordinate story evaluation, scene planning, cost optimization, set tracking, and consistency checks, allowing multiple AI models to operate as a single production team.
What I learned
Specific Instructions Create Better Generations:
AI video generation struggles when prompts are too generic or lack production details. Giving the model concrete camera positions, character descriptions, scene context, and exact actions gives it far stronger guidance and produces more consistent, controllable results.Test Your Assumptions About Model Strengths:
I assumed one video model was best at faces and another at motion and scenery, and built routing around an even split. Head-to-head runs on real dramas redrew the map: the character model won far more territory than expected, taking every shot with a face or a line, and the visual model kept only the shots with nothing to lock. The routing idea still pays off, on video and on the language side alike, but it was the measurement, not the assumption, that drew the boundary.Building AI at Scale Requires Cost Awareness:
Every generation with AI consumes tokens and cloud credit, so cost cannot be an afterthought. Careful planning, choosing the right model for each task, and budget-aware decisions are what make an AI product practical instead of just impressive.Breaking Complex Tasks into Smaller Steps Works Better:
AI video generation struggles with complex actions and long sequences. Splitting each scene into smaller, well-defined shots made the results far more consistent and much easier to control.A Great Shot Does Not Make a Great Story:
A single AI clip can look impressive, but a drama depends on continuity across dozens of shots. I learnt that consistency requires engineering. For example, locking characters to reference images, carrying previous frames forward, and tracking scene state so characters, environments, and props remain consistent throughout the story.Integration of Different AI Models in Qwen Cloud:
Connecting Qwen Cloud's text, image, and video models, each with different APIs and execution flows, into one pipeline required careful coordination to make sure one failed call did not stop the entire production.
What's next for Rexgent
Scaling from Episodes to Entire Series
Extending Rexgent's memory and planning capabilities to support longer narratives, where characters, relationships, and world states remain consistent across an entire season.A Community Platform for Drama Templates
Turning Rexgent into a platform where creators publish their own drama templates and visual styles that anyone can browse, reuse, and remix to kickstart a new story, turning a personal tool into a shared creative community.Publishing Directly to Social Platforms
Connecting finished episodes straight to TikTok, Reels, and YouTube Shorts, completing the journey from a single idea to a posted video without ever leaving the app.Support for More Languages:
Expanding beyond English and Chinese so creators can generate dramas, dialogue, and voices in their own language, opening Rexgent to a truly global audience.
Built With
- alibaba-cloud
- celery
- docker
- fastapi
- ffmpeg
- insightface
- langgraph
- mcp
- neo4j
- nextjs
- pgvector
- postgresql
- python
- qwen
- react
- redis
- socket-io
- tailwindcss
- typescript

Log in or sign up for Devpost to join the conversation.