Inspiration
Most AI video tools are very good at generating content. But generation is only one part of production.
In a real production workflow, things can go wrong after generation: an audio stream can be missing, captions can fail, a render can have technical issues, or an output can simply not be ready to ship.
That led us to a simple question:
What if an AI agent wasn't just responsible for creating the video, but responsible for getting the video all the way to a verified, production-ready result?
That is the idea behind Taf's Pilot.
We wanted to build an agent that could take responsibility for the production lifecycle while still keeping the human creator in control of important creative decisions.
What it does
Taf's Pilot is an autonomous executive video production agent that manages a video from creative brief to final delivery.
Its production lifecycle is:
Understand → Plan → Ask → Produce → Inspect → Correct → Deliver
The agent:
- Understands the creative brief
- Develops a production plan
- Proposes creative directions and asks the creator for approval
- Coordinates voiceover, visuals, captions, and video assembly
- Inspects the generated media
- Detects production defects
- Determines what needs to be corrected
- Re-renders the affected production
- Runs quality control again
- Presents the verified result for final approval
The key behavior is the Inspect → Correct → Verify loop.
In our demonstration, the agent detects that a required narration stream is missing from one of the scenes. Instead of simply reporting the problem and stopping, it identifies the required recovery action, regenerates the affected voiceover, reassembles the video, and runs validation again.
The revised cut passes QC and is then presented to the creator for final release approval.
This is the difference we are exploring between an AI that can generate and an agent that can take responsibility for an outcome.
How we built it
Taf's Pilot is built in Python using the Strands Agents SDK, with a tool-oriented architecture designed around the production lifecycle.
The system combines agent orchestration with a real media-processing pipeline.
Key components include:
- Strands Agents SDK for agent orchestration
- Python for the production system and tools
- FFmpeg for video/audio processing and assembly
- FFprobe for technical media inspection
- Pydantic for structured data and validation
- Uvicorn / Starlette for the web console
- Voice generation with cloud integration and a local Windows fallback
- Automated tests for the production pipeline
- A web console that exposes the agent's state, decisions, QC results, revision history, and telemetry
The architecture separates the production pipeline from the presentation layer, allowing the core production tools and quality-control logic to be tested independently.
The agent's workflow is represented as a state machine:
Understand → Plan → Ask → Produce → Inspect → Correct → Deliver
The Ask stage provides a human-in-the-loop creative gate, while the Inspect/Correct stages provide the autonomous recovery loop.
Challenges we ran into
The biggest challenge was moving beyond a simple "prompt → generate → done" AI workflow.
Video production is a chain of dependent operations. A successful render does not necessarily mean that the resulting video is actually production-ready.
We had to think about questions such as:
- How does the agent know whether its output is valid?
- How can it detect a failure that isn't obvious visually?
- How should it determine what needs to be regenerated?
- How can it verify that a correction actually fixed the problem?
- How do we keep the human involved where creative judgment matters without requiring them to manually manage the entire production?
We also dealt with practical media-pipeline issues involving audio, subtitles, safe-zone placement, rendering, and browser playback.
These challenges pushed us toward treating quality control as part of the agent's responsibility, rather than as a final manual step.
Accomplishments that we're proud of
We are most proud of turning the idea of an autonomous production agent into an observable end-to-end system.
The demonstration doesn't stop when the first video is generated.
It shows the agent:
producing → inspecting → detecting a defect → deciding on a correction → reassembling → validating again → delivering
We are also proud of building real technical inspection into the workflow rather than relying only on an AI model to say that something "looks good."
The production pipeline uses FFprobe-based inspection and automated validation, giving the agent concrete evidence about the generated media.
We also built a human-in-the-loop creative checkpoint, so autonomy doesn't mean removing the creator from the process.
Finally, we built and tested the underlying production pipeline with an automated test suite, giving us confidence that the workflow isn't just a visual prototype.
What we learned
The biggest lesson was that autonomy is not the same as generation.
Giving an AI agent tools is only part of the problem.
For an agent to be genuinely useful in production, it needs a goal, a workflow, feedback about the state of its work, the ability to recognize when something has gone wrong, and the ability to take corrective action.
We also learned that human-in-the-loop and autonomy aren't opposites.
The creator can remain responsible for creative decisions while the agent takes responsibility for repetitive production work, technical inspection, recovery, and verification.
That led us to a principle that shaped the project:
Don't just generate. Inspect. Correct. Verify. Deliver.
What's next for Taf's Pilot
The current version focuses on proving the autonomous production loop.
The next step is to make that loop increasingly production-grade.
We want Taf's Pilot to evolve toward:
- More sophisticated creative planning and audience-aware decisions
- More visual generation and editing capabilities
- Deeper audio and video quality analysis
- More granular self-healing instead of reprocessing larger sections than necessary
- Persistent production memory across projects
- Integration with real publishing and content-management workflows
- Richer production telemetry and audit trails
- More cloud-native execution using AWS services
- The ability to manage multiple videos and production jobs concurrently
Ultimately, the vision is for Taf's Pilot to become more than an AI video generator.
We want it to behave like a production teammate one that can understand an objective, execute the work, inspect its own output, recover from failures, and only hand the result back when it has evidence that the work is ready.
Built With
- agents
- ai
- amazon
- automation
- autonomous
- ffmpeg
- ffprobe
- generative
- human-in-the-loop
- processing
- pydantic
- python
- sdk
- services
- starlette
- strands
- uvicorn
- video
- web


Log in or sign up for Devpost to join the conversation.