-
-
Pool B memory incident detected with high-confidence evidence and bounded remediation awaiting human approval.
-
Gemini proposes concurrency 4→2, while execution remains blocked behind explicit human authorization.
-
CineOps verifies fresh post-remediation telemetry before allowing the incident to reach RESOLVED.
-
Canonical incident reaches RESOLVED v7 with concurrency 2 and zero subsequent memory-budget breaches.
Inspiration
Modern film, VFX, and media-production pipelines depend on large fleets of render and transcode workers operating under extreme memory and compute pressure. When a worker pool exceeds its memory budget, jobs can fail, capacity collapses, and production throughput can degrade quickly.
Today, incident responders face two conflicting problems:
- Diagnosis is slow. Engineers must correlate logs, metrics, thresholds, and workload behavior under time pressure.
- Unconstrained AI is unsafe. An LLM may be useful for investigation, but it should never have unchecked authority to mutate production infrastructure.
CineOps was built around one principle:
Use AI for reasoning — never as the final authority.
What it does
CineOps is an evidence-grounded agentic incident-response system for media-production infrastructure.
It combines real observability, structured Gemini diagnosis, deterministic safety controls, explicit human approval, live remediation, and post-execution recovery verification.
The lifecycle is durable and versioned:
ALERTED v0 → INVESTIGATING v1 → EVIDENCE_READY v2 → DIAGNOSED v3 → AWAITING_APPROVAL v4 → REMEDIATING v5 → VERIFYING v6 → RESOLVED v7
A typical CineOps incident works like this:
Detect.
A media simulator emits operational telemetry and exposes memory-budget failures from active worker pools.
Gather evidence.
CineOps queries Grafana Cloud through the official Grafana MCP integration. Loki provides structured incident events while Prometheus provides the corresponding metrics.
Diagnose.
Google ADK invokes Gemini 3.6 Flash under a strict structured-output contract. Gemini analyzes only the certified evidence and produces a typed diagnosis candidate identifying the correlated pool and failure mechanism.
Ground.
The deterministic backend independently cross-validates Gemini's diagnosis against the telemetry. Gemini does not control lifecycle state.
Plan.
CineOps creates a bounded remediation plan. In our canonical incident, Pool B concurrency is reduced from 4 workers to 2.
Authorize.
The plan cannot execute until a human operator explicitly approves it. Approval is bound to both the exact lifecycle state version and a SHA-256 plan hash, rejecting stale or conflicting authorization.
Execute.
Only after approval does the deterministic backend apply the remediation to the same running simulator environment.
Verify.
CineOps does not assume that a successful API call means the incident is fixed. It queries fresh post-remediation telemetry and verifies causal recovery before allowing the lifecycle to reach RESOLVED.
The authority model
The core architectural distinction in CineOps is the separation of reasoning from authority:
Grafana provides evidence → Gemini investigates and proposes → Human approves → Deterministic backend executes → Fresh telemetry proves recovery
Gemini deliberately cannot:
- approve a remediation;
- advance lifecycle versions;
- execute infrastructure mutations;
- or declare an incident resolved.
Those responsibilities remain outside the LLM.
This creates an operational boundary suitable for systems where AI assistance is valuable but uncontrolled autonomy is unacceptable.
How we built it
CineOps uses Google Agent Development Kit (ADK) with Gemini 3.6 Flash for schema-bound incident diagnosis.
Observability is provided by Grafana Cloud, with Loki LogQL and Prometheus PromQL accessed through the official read-only Grafana Model Context Protocol (MCP) integration.
The deterministic control plane is built with Python 3.12, FastAPI, Pydantic, SQLite durable state, Uvicorn, and HTTPX.
The operator-facing Control Room is built with Next.js 15, React 19, and Lucide React.
OpenTelemetry instrumentation connects the simulated production workload to the evidence pipeline.
Canonical proof
CineOps includes an authoritative end-to-end production-path proof for incident:
inc-fullstack-real-015
The proof demonstrates:
- exactly one fault trigger;
- exactly one incident creation;
- exactly one human approval;
- the complete 8-state lifecycle from
ALERTED v0toRESOLVED v7; - Gemini diagnosis grounded in Loki and Prometheus evidence;
- live Pool B remediation from concurrency 4 → 2;
- fresh post-remediation telemetry;
- recovery memory stabilized at 700 MiB ≤ 1024 MiB;
- and zero subsequent qualifying memory-budget violations.
Recovery evidence is required to occur after verification begins, preventing stale pre-remediation telemetry from being used as false proof of success.
Challenges we ran into
One major challenge was keeping LLM output useful without allowing it to destabilize the deterministic state machine. We solved this with strongly typed structured output, strict schema validation, and independent telemetry grounding.
Another challenge was distinguishing cloud-provider failures, network behavior, and observability latency from actual CineOps defects while preserving deterministic tests.
We also had to guarantee same-runtime remediation continuity: the process being changed must be the same live workload that generated the incident.
Finally, recovery verification required causal ordering. A metric that existed before remediation cannot be accepted as evidence that remediation succeeded, so CineOps tracks strict sequence and lifecycle boundaries for fresh telemetry.
Accomplishments that we're proud of
The complete engineering verification suite currently includes:
950 Python tests + 54 Control Room tests = 1,004 passing tests
Additionally:
- 0 Ruff errors;
- 0 Mypy errors across 152 source files;
- clean Next.js production build;
- no tracked secrets in the repository;
- deterministic lifecycle and approval invariants;
- and an end-to-end canonical recovery proof.
The result is not only an AI diagnosis prototype. CineOps demonstrates a complete operational loop from evidence to safely authorized action to verified recovery.
What we learned
The most important lesson was that agentic AI for mission-critical infrastructure should not mean unrestricted autonomy.
A stronger design treats the LLM as a specialized reasoning component inside a deterministic operational system.
Telemetry grounding, cryptographic approval locks, durable lifecycle state, explicit human authorization, and causal recovery verification provide a much stronger foundation for trustworthy AI-assisted operations.
What's next for CineOps
The next step is connecting the remediation layer to production orchestrators such as Kubernetes, AWS Deadline Cloud, and render-farm schedulers.
We also plan to expand the remediation catalog to include memory scaling, workload redistribution, dynamic chunk splitting, and additional GPU and media-pipeline failure patterns.
Longer term, CineOps can become a multi-tenant operations platform with enterprise SSO, role-based approval policies, and broader observability correlation across production infrastructure.
AI-assisted speed. Deterministic safety. Verified recovery.

Log in or sign up for Devpost to join the conversation.