Inspiration
Post-production is creative storytelling, yet professional video editors spend 80% of their time on repetitive mechanical chores: sifting through hundreds of raw takes, matching aspect ratios, resolving overlapping timecodes, looking up shifting social platform delivery specs, and tweaking loudness curves.
For the Agentic Cinema Hackathon, we set out to build an autonomous, blockbuster-grade post-production studio that treats editing as a unified multi-agent orchestration challenge. We wanted an editor to simply declare their narrative vision, while an autonomous AI swarm handles the ingestion, multimodal vision analysis, web research, timeline assembly, and export directly inside an industry-standard NLE: Blackmagic DaVinci Resolve.
What It Does
OVIO is an autonomous 8-agent cinema production platform that takes raw video rushes, photos, and natural language editing prompts and produces broadcast-ready DaVinci Resolve timelines.
Key capabilities:
- Multimodal Ingestion: Samples keyframes and uses Gemini 2.5 vision to understand scene dynamics, pacing, camera motion, and subject focus.
- Real-World Delivery Grounding: Queries the Parallel Search API and Model Context Protocol (MCP) to ingest exact platform specifications (Instagram Reels, TikTok, YouTube Shorts, DCI 4K Cinema) and audio loudness rules (-14 LUFS, -24 LUFS).
- Declarative Edit DSL: Compiles director goals into an atomic, timecode-validated JSON Edit DSL covering clips, cuts, cross dissolves, audio normalization, and delivery presets.
- DaVinci Resolve Worker Bridge: Executes timeline operations directly inside DaVinci Resolve via a high-speed in-process worker on port 7799, bypassing external socket restrictions.
- Self-Healing QA Critic: Evaluates rendered media against target constraints. If discrepancies are detected, Agent 8 isolates the failure, patches the DSL, and re-executes automatically.
- Google Cloud Archiving: Backs up finished timelines and render artifacts to Google Cloud Storage.
How We Built It
- Google Cloud Platform:
- Google Gemini 2.5 Multimodal Models: Powered via the google-genai SDK for visual asset reasoning, scene tagging, and edit planning.
- Google Cloud Agent Builder: Manages agent state transitions and multi-turn workflow continuity.
Google Cloud Storage: Automated archival of rendered video masters to gs://ovio-cinema-renders/.
Partner Track Integration (Parallel):
We integrated Parallel as our core intelligence partner using the Parallel Search API and Model Context Protocol (MCP).
Agent 4 (Parallel Research Agent) queries Parallel for live platform safe zones, maximum bitrates, frame rates, and loudness targets, ensuring the edit planner never hallucinates video delivery standards.
Autonomous Multi-Agent Swarm:
Agent 1: Director Agent (Intent and pacing interpretation)
Agent 2: Multimodal Ingestion (Frame sampling and Gemini vision analysis)
Agent 4: Parallel Researcher (Real-world media delivery standards)
Agent 3: Edit Planner (Declarative Edit DSL assembly)
Agent 5: Plan Validator (Timecode collisions and syntax checks)
Agent 6: Resolve Executor (In-process worker dispatch)
Agent 7: QA Critic (Quality and loudness evaluation)
Agent 8: Correction Agent (Autonomous self-healing repair loop)
Local Host Integration:
In-process Python worker (worker/ovio_worker.py) running natively inside DaVinci Resolve, exposing a non-blocking REST bridge for instantaneous UI responsiveness.
Challenges We Ran Into
On Windows systems, DaVinci Resolve Free edition restricts external socket IPC, causing standard external scripting calls to fail silently. Rather than forcing users onto paid hardware dongles, we engineered an in-process worker bridge that runs directly inside DaVinci Resolve's Fusion scripting environment. The worker hosts a lightweight HTTP daemon on port 7799 that interacts directly with the live timeline, giving free and studio users equal access to full autonomous timeline control.
Accomplishments That We're Proud Of
- 10 out of 10 automated end-to-end tests passing with zero failures across photo montages, video highlights, mixed media, 9:16 social reels, and audio normalization.
- Live integration with Parallel Search API and MCP delivering accurate media constraints.
- Fully autonomous self-healing loop: Agent 8 successfully diagnoses execution failures and re-plans without user intervention.
- Zero UI freezing during heavy render and timeline generation passes.
What We Learned
We discovered that combining multimodal LLMs with live domain search (Parallel) eliminates hallucinated technical settings in creative tools. Grounding the agent in real delivery specs before timeline generation is what transforms a toy demo into a studio-ready production system.
What's Next for OVIO
- Multimodal audio stem separation and automated ducking against dialogue tracks.
- Direct generation of Blackmagic Fusion node graphs for custom kinetic title motion and VFX.
- Distributed render dispatch across Google Cloud GPU spot instances.
Links and Resources
- GitHub Repository: https://github.com/Shambhavi500/Ovio
- YouTube Demonstration: https://youtu.be/B8040l4kvBw
Built With
- davinci-resolve
- fastapi
- gemini-2.5
- google-cloud
- google-genai
- mcp
- next.js
- parallel-api
- python
- tailwindcss
- typescript
- websockets
Log in or sign up for Devpost to join the conversation.