Martini Shot
Inspiration
Post-production is often described as a creative process, but the hardest operational problems are coordination problems. A supervisor must keep track of incoming media, broken files, missing audio, scene context, sound levels, visual defects, dubs, captions, delivery requirements, optional finishing ideas, approvals, retries, budgets, and handoffs across many clips and stations.
That work is usually spread across media tools, job queues, spreadsheets, cloud logs, dashboards, and approval records. A supervisor may know that a job failed without knowing whether the cause was the footage, the factory, a dependency, a budget policy, or an upstream handoff. They may also see a technically possible improvement without knowing whether it is worth the cost or safe to introduce into an approved cut.
We wanted to build a system that could take responsibility for the operational loop without taking accountability away from the human supervisor. The inspiration for Martini Shot is the idea that a production team should be able to upload real footage, set the budget, walk away, and return to an explained and governed run.
The product is designed for real uploaded production clips. Generative video is not the starting point or the requirement. The core value is the ability to inspect, clean up, organize, prioritize, deliver, and explain real media safely. Optional generative finishing can extend or alter an identified shot, but it sits inside that larger production workflow.
What it does
Martini Shot is an observability-native post-production supervisor for film and television teams. It turns a batch of uploaded clips into a governed chain of media jobs.
A supervisor uploads clips in order and sets a budget. Martini Shot then:
- Validates each uploaded file and quarantines files that are corrupt, incomplete, undecodable, or missing required audio.
- Uses Gemini to understand what is actually in the footage, including the scene and spoken words. Silence is represented honestly rather than filled with invented dialogue.
- Runs mandatory loudness and pickup checks for every uploaded clip before optional finishing work begins.
- Asks specialist agents to inspect the updated clips for delivery, dubbing, extending, corrections, relighting, coverage, and camera language.
- Estimates the cost of each proposed action.
- Uses a Google ADK orchestrator to rank the complete proposal set by impact, cost, dependencies, upload order, scene context, and handoff notes.
- Dispatches only the work that fits the remaining budget and clears its dependencies.
- Keeps generated media in alternate lanes rather than silently overwriting an approved cut.
- Applies draft-first and visual-quality gates where appropriate.
- Adds only passed artifacts to the final references while keeping failed, paused, incomplete, and human-review artifacts out of the cut.
The product includes a Timeline, governed worklist, Agent Notes, Suggestions, Approvals, Decisions, Changes, Studio, Reports, playable clip references, and Analytics/Run Pulse.
The core workflow works with real clips even when no generative operation is selected. The backend’s final-reference assembly always keeps the original references and adds only artifacts from jobs whose authoritative status is passed. The frontend’s Final cut strip chooses the latest successful After for each original clip in upload order; if there is no successful transformation, the original remains the playable reference.
Studio and the Stage 1a routes provide optional, approval-tracked directed edits on an identified shot. The current implementation records whether an edit used the primary Omni path or a Veo fallback. Omni’s refusal to process some public-domain or public-commons source media is a limitation of that optional generative edit request. It does not prevent Martini Shot from ingesting, measuring, dubbing, cleaning up, delivering, or supervising real uploaded production clips.
A separate Post Supervisor path investigates failed, quarantined, and human-review jobs. It uses the available job evidence, specialist investigators, a Verification Agent, a Gemini synthesis step, bounded autonomy, and budget controls. The current production trigger is wired for failed, quarantined, and needs-human terminal states.
How we built it
We built Martini Shot as an asynchronous production system rather than a synchronous chatbot. The browser starts a run through a FastAPI API, and long-running work is processed by a Firestore-backed lease queue.
The user-facing flow begins in FinishBar. The frontend accepts multiple media files in the order chosen by the user, displays each ingest state, allows reordering before the run, converts the dollar budget into integer micro-units, waits for file checks to reach terminal states, and sends only passed ingest job IDs to POST /api/v1/projects/{project_id}/finish.
The backend then follows a controlled production order:
- Collect original references.
collect_original_refs()reads passed ingest jobs from Firestore and preserves their creation order. If the user supplies explicit ingest IDs, the code rejects duplicate, missing, or invalid IDs. - Create scene context.
clip_context_after_ingest()downloads the original from Google Cloud Storage, runs the ingest-understanding path, and attachesingested,spoken_words, andscenefields to the shot context. - Create mandatory cleanup work.
mandatory_cleanup_items()creates a loudness item and a pickups item for every shot. Loudness is ordered across the uploaded clips, and each pickup item is blocked on its shot’s loudness item. - Dispatch cleanup.
dispatch_next()enqueues only jobs whose dependencies are clear. The worklist is persisted and published through the project event hub while the run continues. - Run specialist attendance. Once cleanup is terminal,
run_proposal_phase()inspects each updated clip in upload order. The current proposal roster is delivery, dub, extend, corrections, relight, coverage, and camera language. - Price proposals. Candidate notes with
needs_workstatus and a proposal are passed to the spend-pricing agent. If pricing fails, the code retains station defaults as advice rather than claiming a fabricated live price. - Rank the complete proposal bag. The Google ADK orchestrator receives the notes, remaining micro-budget, shot order, scene context, and handoff spine. It returns an order, dependencies, dropped work, and a reason. A bounded fallback planner handles an ADK ranking failure.
- Dispatch governed work. Ranked notes become worklist items. Delivery is pinned last, two generative edits cannot run concurrently on the same shot, and items that do not fit the remaining budget stay waiting or become throttled when the budget is exhausted.
- Execute explicit stations. The dispatcher has real branches for ingest, loudness, delivery, spend, pickups, dub, extend, corrections, relight, coverage, and camera language.
- Assemble safe references.
assemble_final()keeps all original references and appends only passed job artifacts. A failed, quarantined, throttled, incomplete, or human-review artifact cannot enter the backend’s final-reference list.
Agent architecture
The finishing path is a bounded multi-agent cycle:
Specialist inspection → structured proposal → spend pricing → global orchestration → validated worklist → asynchronous station execution → quality and approval gates.
build_finishing_team() constructs a Google ADK ParallelAgent with one LlmAgent for each proposal station. Each agent receives a station-specific inspection prompt and a FunctionTool that calls the real inspection path. The prompts distinguish defects from improvements and explicitly allow the agent to leave a shot alone.
The ADK runner is invoked through Runner.run_async(). Its events are parsed into validated InspectNote objects. A separate ADK orchestrator then receives the complete candidate bag and returns a validated rank plan. Empty notes remain attendance holes, ok or leave-it notes do not become jobs, and malformed model output goes through bounded fallback handling.
The Post Supervisor is a separate loop. The worker terminal hook calls maybe_deliberate() for the currently wired failure states. The cycle is idempotent per job, loads the real specialist map, invokes the real Verification Agent and Post Supervisor synthesizer, and runs through the budget loop without allowing a deliberation failure to change the worker’s own job outcome.
Real media operations
The core media path uses the installed FFmpeg and FFprobe binaries directly. The media adapter probes duration, frame rate, codec, audio presence, dimensions, format, and bitrate. It checks decodability with FFmpeg error detection, extracts frames, decodes audio to mono 16 kHz LINEAR16 WAV for listen-capable agents, measures integrated loudness and true peak through the ebur128 filter, and performs audio fitting, segment splitting, reassembly, and concatenation.
The optional bounded Omni edit path uses FFmpeg to split a source that exceeds the model’s accepted duration, calls the real Vertex interaction for each segment, and reassembles the returned videos. If Omni refuses or fails and the station is configured for fallback, the Veo path can be used. The job records the actual render model, fallback flag, and error so the result is not misrepresented.
Storage and state
GCSMedia wraps a real Google Cloud Storage client for uploads, downloads, existence checks, deletion, and signed media URLs. The signing path supports service-account credentials, impersonated credentials, and IAM signBlob for runtime credentials that do not contain a private key.
FirestoreStore wraps a real Firestore client. It stores projects, jobs, shots, alternates, approvals, worklists, and deliberations. Transactional updates support the lease queue, intake pause flags, and approval machine without stale full-document overwrites.
Grafana MCP and Run Pulse
Grafana is part of the runtime control loop. GrafanaMcpConnector supports a hosted Streamable HTTP connection to Grafana Cloud MCP and a headless mcp-grafana binary over stdio. It discovers datasource UIDs and resolves MCP tools for dashboards, PromQL, Loki, Tempo traces, annotations, and incidents.
register_grafana_tools() exposes read tools for Prometheus, Loki, dashboards, and traces. Annotation and incident tools are act-class tools that check the autonomy mode before dispatch. In a proposal-only mode, they return a structured proposal instead of writing to Grafana.
Run Pulse is deliberately deterministic and has no LLM on its path. It queries global and project failure rates, station p50 durations, annotations, optional Grafana deep links, and Tempo trace enrichment. It calculates this-run burn from Firestore job costs, derives remaining work from the persisted worklist, estimates ETA from Grafana station history, and returns factory health, burn, ETA, intervention (“wheel”) items, dashboards, and job rows. The frontend polls Run Pulse every 15 seconds while active work remains.
Frontend control room
The React frontend uses typed endpoint functions for projects, ingest, jobs, worklists, approvals, alternates, reports, directed edits, scripts, and Run Pulse. It consumes project events and updates the React Query cache for jobs, worklists, approvals, and deliberations.
The Timeline groups work by original clip and shows Before, After, Final cut, What changed, Agent Notes, status, attempts, and cost. The Final cut strip can play the latest successful After for each original clip. Analytics loads the same selected project as Timeline and combines live jobs, worklist state, Firestore costs, and Grafana-backed health and duration context.
Challenges we ran into
Enforcing the house order across asynchronous work
The hard part was not creating a list of stations. It was making the list behave correctly when jobs run asynchronously, fail, retry, or finish out of order. The worklist stores explicit blocked_by IDs. Loudness precedes pickups, optional work waits for cleanup, delivery is pinned last, and dispatch considers both dependencies and remaining budget. Independent clips can continue while two generative edits on the same shot are prevented from running at the same time.
Making real clips the primary path
The system had to work on ordinary production media, not only on synthetic or generated examples. That meant validating real containers, decoding real audio, checking for missing streams, measuring actual loudness, carrying scene context between stations, validating handoffs, and preserving the original clip when no repair or transformation was needed. Generative operations are optional decisions inside that workflow, not prerequisites for the product to be useful.
Handling real media and model constraints
Real media introduces corruption, missing audio, variable frame rates, decodability failures, loudness ambiguity, long-running renders, transport timeouts, quota failures, and incompatible encodes. Optional generative operations add another constraint: a provider may refuse a source even when the source is public or public-domain. In this project, Omni refused some public-commons media. The correct behavior is to record that refusal, use an allowed fallback where configured, or leave the original clip intact—not to treat the whole post-production run as invalid.
The code responds with FFmpeg probes and decode gates, real measurements, segment splitting, timeout-aware retries, explicit fallback metadata, and separate evaluation accounting for Omni and Veo.
Keeping model judgment inside safe contracts
The agents must return typed inspection notes, station decisions, and rank plans. Empty notes are not permission to invent work. Leave-it notes are not converted into jobs. The orchestrator is instructed not to invent stations, costs, commands, or dialogue. Parsing failures use bounded fallback logic instead of allowing malformed model output to reach the worker.
Preserving the final cut
The backend’s assemble_final() function retains originals and appends only passed artifacts. The frontend adds a user-facing final-cut selection layer, but promotion is approval-tracked. This prevents a preview, failed render, paused job, or human-review artifact from quietly becoming the production reference.
Making Grafana actionable without making it authoritative for media state
Grafana supplies operational evidence, while Firestore remains the source of truth for jobs and worklists. Run Pulse treats absent Grafana samples as unknown rather than inventing a healthy result. Cost comes from Firestore cost_micros; health, duration, annotations, and trace enrichment come from Grafana. This split lets the control room remain honest when telemetry is incomplete.
Designing graceful failure
The system has deliberate fallback paths: billed Python inspection if ADK attendance fails, a fallback rank planner if the ADK rank response is invalid, Veo after a failed Omni operation where configured, retryable queue failure before terminal failure, and structured Grafana proposals when autonomy does not permit writes. Each fallback is recorded in code or job state.
Accomplishments that we're proud of
We are proud that Martini Shot has a real end-to-end execution spine. A browser action creates real ingest jobs, passed ingest references become shot context, cleanup work is persisted, specialists inspect actual uploaded media, proposals are priced and ranked, workers run explicit stations, job state is reconciled, and the frontend updates from project events and typed API reads.
We are proud of the separation between model judgment and deterministic enforcement. Agents can identify defects and propose improvements, but code decides whether the source decoded, whether the handoff is valid, whether the job is affordable, whether a dependency is clear, whether a draft passed, and whether an artifact is allowed into final references.
We are proud of treating real clips as first-class inputs. Ingest, scene understanding, loudness, pickup QC, delivery validation, dubbing, handoffs, budget governance, failure investigation, and Grafana-backed operations remain useful without a generated video edit. The optional Omni/Veo path adds finishing capability without defining the product.
We are proud of the Grafana implementation because it is executable. The code opens a real MCP session, discovers datasource UIDs, queries PromQL, Loki, and Tempo, exposes tools through the supervisor registry, checks autonomy before writes, writes job annotations from the worker path, and assembles Run Pulse from Grafana plus Firestore rather than rendering a canned dashboard.
We are proud of the frontend because it makes the state legible. The Timeline exposes Before, After, Final cut, What changed, Agent Notes, status, attempts, and cost. The Final cut strip can play the latest successful After for each original clip. Analytics combines this-run jobs with factory telemetry and explains burn, failure, ETA, and intervention context.
What we learned
We learned that reliable agentic media software is a systems problem. The model call is only one stage. The surrounding system must provide inputs, context, contracts, budgets, dependencies, retries, leases, evidence, quality gates, and safe state transitions.
We learned that deterministic enforcement should not be replaced by model confidence. A model can identify a defect or suggest an improvement, but code must decide whether the media is valid, the handoff is complete, the job is affordable, the dependency is clear, and the artifact can enter the final references.
We learned that observability must be designed into the workflow from the beginning. A dashboard added after the fact can show that a job ran, but it cannot reconstruct the evidence, cost, dependency, and approval context needed to govern an autonomous decision. Grafana MCP makes telemetry actionable by allowing the supervisor to investigate and record interventions through the same evidence layer.
We learned that bounded autonomy is more useful than unrestricted autonomy. Budgets, lease ownership, retry limits, draft-first rendering, continuity locks, approval transitions, and final-reference rules turn creative agents into controlled production participants.
We learned that model limitations should be isolated rather than allowed to contaminate the rest of the product. An Omni refusal on a public-commons clip is a failure of one optional edit request. It is not a failure of ingest, loudness, dubbing, delivery, worklist governance, or operational supervision. The original clip remains valid input and can continue through the real-clip workflow.
We learned that abstention is a quality behavior. A camera-language agent should leave a still life alone. A delivery or relight agent should not invent a defect. A missing or unavailable capability should produce an empty or abstaining result rather than an all-clear claim.
Most importantly, we learned that the strongest product story is not “AI can make video.” It is that a post-production supervisor can delegate a messy, expensive operation without delegating accountability.
What's next for Martini Shot
The immediate next step is to deploy and rehearse the exact current tree. The production path needs a hosted verification of real clip upload, ingest, cleanup, attendance, pricing, ranking, dispatch, one passed artifact or clean original path, one budget-waiting item, and Run Pulse’s Grafana reads and annotations. The final demo should be recorded from that running path rather than from precomputed evidence.
The next reliability step is to complete the release gate. The repository’s declared target is 90 percent service-layer coverage, and the current project state still records that the full coverage gate is below target. The detached deliberation path also needs a dedicated test for failed-task cleanup during shutdown.
The next observability step is to wire more durable event sources into the production supervisor trigger. The current team.py trigger map covers failed, quarantined, and needs-human jobs. Stuck leases, crash recovery, QC breaches, spend breaches, runaway retries, daily-budget breaches, and dubbing breaches have routing vocabulary but still require durable event emission before they can reliably start a production deliberation cycle.
The next product step is to deepen the operator experience. The current code already provides the Final cut strip, Agent Notes, Changes, Decisions, Studio, and Run Pulse. Future improvements should make evidence links from each worklist row more direct, show the dependency and budget reason beside every decision, expose Grafana trace and annotation context without leaving the run, and make the final-reference player more polished.
The next creative step is not to add an unbounded list of effects. It is to extend the governed pattern: every new operation should have a real station branch, an input contract, cost accounting, draft or quality gates, approval behavior, fallback metadata, Grafana telemetry, and final-reference rules. Script-aware revision and selective regeneration are promising directions, but they should only be presented as complete once those execution and evidence paths exist.
The long-term goal is unchanged: make post-production easier to delegate while keeping every important decision explainable, budgeted, reversible, and accountable.
Code references
backend/api/finish.py— finish endpoint, background scheduling, scene context, cleanup, proposal phase, worklist, and Run Pulse routes.backend/supervisor/finishing_loop.py— cleanup dependencies, delivery ordering, budget dispatch, generative-job constraints, and final-reference assembly.backend/supervisor/adk_finishing.py— ADK specialist construction, billed inspection tools, runner invocation, and global ranking.backend/jobs/worker.py— lease claims, handoff validation, station execution, authoritative terminal state, event publication, and annotations.backend/stations/run.py— explicit station dispatcher.backend/core/gcs.py— real Cloud Storage operations and signed media URLs.backend/core/firestore.py— Firestore state and transactional updates.backend/core/media.py— FFmpeg/FFprobe probing, decoding, loudness, splitting, and reassembly.backend/supervisor/mcp.py— hosted and OSS Grafana MCP transports and tool resolution.backend/supervisor/tools.py— Grafana MCP supervisor tools and autonomy checks.backend/supervisor/run_pulse.py— deterministic Grafana plus Firestore health, burn, ETA, trace, and intervention snapshot.backend/supervisor/team.py— production deliberation triggers and investigator, verifier, and synthesizer wiring.frontend/src/components/FinishBar.tsx— upload order, ingest polling, budget input, and finish action.frontend/src/components/TimelineBoard.tsx— Before/After, Agent Notes, Final cut, cost, attempts, and playable playlist UI.frontend/src/pages/AnalyticsRoute.tsx— worklist and Run Pulse integration.frontend/src/api/endpoints.ts— typed frontend-to-backend contracts.
Built With
- chirp-3-hd
- cloud-firestore
- cloud-iam/iam-credentials
- cloud-run
- cloud-storage
- fastapi
- ffmpeg/ffprobe
- gemini
- gemini-omni
- google-agent-development-kit-(adk)
- google-cloud-text-to-speech
- google-gen-ai-sdk
- grafana-cloud
- grafana-mcp
- grafana/mcp-grafana
- loki
- opentelemetry
- prometheus/mimir
- python
- react
- tempo
- typescript
- veo
- vertex-ai
Log in or sign up for Devpost to join the conversation.