Inspiration
A tiny film crew can own three good cameras and still be missing one crucial person: someone who can watch every angle, follow the script, and decide what should be on screen.
That is especially painful for student filmmakers, indie short-film teams, and scripted creators working with 2–4 cameras but no dedicated switcher or editor on set. Modern multicam tools can capture and synchronize several angles, but the creative switching decision still becomes another job for the crew and another pass in post.
Clappy asks a different question: what if the screenplay could become that missing crew member?
What it does
Clappy is an AI director for small multicamera productions. It follows the scene while it is being performed, remembers what the production has already shown, and turns several camera feeds into a live first cut without destroying any originals.
A filmmaker can:
- Load a scene and map the cameras. One angle might cover Tom, another Bella, and another the wide.
- Arm and roll the whole shoot. Clappy coordinates camera state and keeps each source recording separately.
- Follow the dialogue. Google Speech-to-Text recognizes the performance and aligns it to the scene so the director agent knows where the actors are in the script.
- Direct the live program feed. Before making a creative shot choice, the agent reads the production history through the official ClickHouse MCP server, then Gemini proposes the next camera based on the script, recent shot history, camera availability, and any active director instruction.
- Take direction. Commands such as “hold Bella,” “go wide,” “roll,” and “cut” are handled through a separate director voice/control path so creative intent can override automation.
- Keep every original and make new cuts later. Clappy compiles the live decisions into an OpenTimelineIO timeline and renders an actual MP4 with FFmpeg. The filmmaker can then ask for another version, such as “make Bella dominate this scene,” without destroying the original edit.
The result is not a chatbot attached to an editor. It is an agent participating in the production workflow itself.
The demo is honest
The current hosted hackathon build uses three labeled virtual iPhone camera sources, not a claimed native three-phone shoot.
The phones in the demo are virtual. The media pipeline is not. Three real media publishers stream through MediaMTX, the browser decodes the WebRTC feeds, each camera is recorded independently, source identities are verified, and OTIO + FFmpeg produce playable multicamera edits from the preserved recordings.
That gave us a way to build and test the hard orchestration layer deeply before pretending the native-device layer was finished.
Why ClickHouse is load-bearing
ClickHouse is not a logging sidecar in Clappy. It is the agent's production memory.
During a take, Clappy records production events such as:
- scene and take state
- camera health and availability
- transcript/script alignment
- program-camera decisions
- manual holds and overrides
- director commands
- agent proposals and outcomes
- edit history
Every live autonomous shot-selection request performs a mandatory history read through the official mcp-clickhouse server before Gemini makes its choice. Alternate edits and production recall also use that memory through Google ADK.
The ClickHouse access path is scoped and read-only for the agent. Cross-production reads, base-table access, and writes through the agent view are denied, and event redelivery is deduplicated.
Remove ClickHouse and Clappy loses the memory that lets it reason about what has already happened in the shoot rather than treating every line as an isolated prompt.
How we built it
Google intelligence
- Google Agent Development Kit (ADK)
- Gemini on Vertex AI
- Google Cloud Speech-to-Text for live dialogue/director intent
- Google Cloud Text-to-Speech for owned rehearsal dialogue
- Google image generation for the disclosed cinematic rehearsal plate
Production memory
- ClickHouse
- official ClickHouse MCP server
Media + editing
- MediaMTX for media transport
- WebRTC for the live program monitor
- FFmpeg for publishers, recordings, and renders
- OpenTimelineIO for edit representation
Application
- React + Vite + TypeScript studio
- FastAPI + Python coordinator
- SQLite/WAL for transactional local state and recovery
- HTTPS-hosted Google Cloud deployment
The creative agent decides what should happen. Deterministic software performs the recording, validation, switching, timeline compilation, and rendering.
What actually works today
The latest hosted walkthrough completed a 29.3-second scene with:
- five recognized dialogue lines
- three accepted Google + ClickHouse-backed live directing proposals
- an actual live camera boundary from the wide to Bella
- manual Bella hold/override behavior
- three independently preserved and verified recordings
- a READY original MP4
- a second READY 100% Bella alternate edit generated from the same preserved take
- both versions retained for A/B review
Voice-control paths for roll, hold Bella, wide, and cut have also passed against actual Google recognition and intent handling.
Clappy also fails visibly rather than fabricating success. If a selected camera dies, the program falls back to a healthy source and rejects new cuts to the dead camera. If the entire rig disappears, the take stops as PARTIAL and preserves what was captured. Model/tool failures do not become canned AI responses.
The current repository records 40 passing Python tests plus a passing production web build, alongside Chrome and WebKit/iPhone-emulation media tests and end-to-end Google/ClickHouse smoke tests.
Challenges we ran into
Live AI has to be fast enough to still be relevant
A correct camera decision that arrives after the moment has passed is still wrong. Early hosted calls were too slow, so Clappy now bounds creative context, uses a lighter Gemini model for live choices, keeps the mandatory ClickHouse history read focused, and rejects stale proposals instead of applying them late.
Multicam timing is less simple than “all videos started together”
Browsers, media clocks, transport, and source timestamps do not give us one magically perfect clock. We added immutable frame-coded source derivatives, browser timing observations, decoded frame evidence, and explicit uncertainty instead of claiming frame-accurate synchronization we have not yet proven.
The agent should not own the physics
We learned to keep creative reasoning and deterministic execution separate. Gemini can decide that Bella should remain on screen; plain software should verify that Bella's camera exists, enforce a manual hold, record the decision, and compile the timeline.
What we learned
The most useful agentic systems are not the ones where the LLM does everything.
For Clappy, the strong division is:
Gemini reasons. ClickHouse remembers. Deterministic media software executes.
That separation made the system easier to test, easier to audit, and much harder to fake accidentally.
We also learned that sponsor integration becomes far more convincing when it changes the architecture. ClickHouse is useful here because a director is inherently temporal: what should happen next depends on what the audience has already seen.
What's next
The next step is the product Clappy was designed for from the beginning: native phone camera clients running the same production protocol.
That work includes:
- real 2–4 phone capture instead of virtual camera sources
- stronger live clock calibration and frame-accurate conforming
- physical microphone validation for hands-free direction
- richer script/shot-plan understanding
- native transfer of full-quality phone originals after the shoot
The long-term goal stays intentionally simple:
Three phones. One script. No switcher operator.
Built With
- clickhouse
- fastapi
- ffmpeg
- gemini
- google-adk
- google-cloud
- google-cloud-speech-to-text
- google-cloud-text-to-speech
- mcp-clickhouse
- mediamtx
- model-context-protocol
- opentimelineio
- python
- react
- sqlite
- typescript
- vertex-ai
- vite
- webrtc

Log in or sign up for Devpost to join the conversation.