-
-
API Dashboard: Production Guardian API is live with PostgreSQL, Grafana Cloud, and MCP integration operational.
-
Incident Investigation: AI agents analyze infrastructure failures and translate them into real production impact and remediation actions.
-
Live Telemetry: Real-time synthetic production telemetry is streamed to Grafana Cloud to simulate infrastructure incidents.
-
Telemetry & MCP: The simulator, Grafana Cloud, and AI agents work together for real-time monitoring and investigation.
-
Incident Management: Centralized incident tracking from detection and investigation through resolution.
-
Production Overview: NIGHTFALL’s production context is combined with infrastructure health to identify operational risks.
-
Real-time telemetry trends reveal infrastructure degradation and enable targeted incident simulation for AI-driven diagnosis.
Inspiration
A film can lose millions because of something as simple as a server running out of storage.
On a film set, a technical failure isn't just an IT problem. A storage bottleneck can stop footage from being ingested. A network failure can delay transfers. A camera failure can mean lost frames. A rendering bottleneck can push an entire post-production schedule behind.
Yet traditional monitoring systems usually stop at:
"Storage utilization is 96%."
That isn't the question a producer needs answered.
The real question is:
"What does this failure mean for today's shoot—and what should we do before it becomes a production delay?"
That gap between infrastructure telemetry and production decisions inspired us to build Production Guardian.
We wanted to build a system that doesn't just detect that something is broken, but understands why it is happening, what it means for the production, and what should happen next.
Production Guardian connects technical infrastructure data with production context and uses AI agents to turn raw telemetry into actionable production decisions.
What it does
Production Guardian is an AI-powered autonomous operations system for film and media production.
It monitors production infrastructure, detects incidents, investigates them through Grafana Cloud MCP, correlates evidence across storage, networking, camera, and editing systems, diagnoses the likely root cause using Google Gemini, predicts the impact on the production, recommends remediation, and verifies recovery.
For our demonstration, we created a fictional production called NIGHTFALL, currently on Production Day 47, with Scene 42 representing a critical workflow.
The system follows a complete incident lifecycle:
Detect → Investigate → Correlate → Diagnose → Predict Impact → Recommend → Approve → Remediate → Verify
For example, a storage incident can produce a chain such as:
$$ Storage\ Utilization \uparrow \rightarrow Write\ Latency \uparrow \rightarrow Ingest\ Throughput \downarrow $$
Instead of simply reporting that storage utilization is high, Production Guardian can determine that the issue is likely to affect the production workflow and estimate the resulting editorial delay.
In our demo, the system identifies storage saturation on INGEST-01 with 94% confidence and predicts a 47-minute editorial delay for Scene 42.
The system then presents a remediation recommendation, waits for human approval, simulates the remediation, and checks Grafana again to verify that the incident has recovered.
How we built it
We built Production Guardian around a multi-agent architecture using Google Gemini and Google ADK.
At the center is an OrchestratorAgent that coordinates three specialized agents.
InvestigationAgent
The InvestigationAgent investigates the infrastructure incident using Grafana MCP.
It can:
- Query metrics
- Compare metrics against a 24-hour baseline
- Query logs through Loki
- Retrieve active alerts from Alertmanager
This gives the AI access to actual observability data instead of relying only on predefined information.
ProductionImpactAgent
The ProductionImpactAgent connects infrastructure problems to production context.
It retrieves:
- Production context
- Scene context
- Relevant production parameters
It then uses deterministic calculations to estimate the impact of the incident on the production workflow.
RemediationAgent
The RemediationAgent uses Gemini to generate remediation recommendations.
It then:
- Presents the recommendation for human approval.
- Simulates the remediation.
- Checks the infrastructure again through Grafana.
- Verifies whether the incident has recovered.
For observability, we use Grafana Cloud with Prometheus and Loki, connected through the official mcp-grafana integration.
We also built our own telemetry simulator so we could reproduce realistic production incidents without affecting real infrastructure.
The simulator currently supports five scenarios:
- Storage Saturation
- Network Degradation
- Camera Failure
- Render Bottleneck
- Media Integrity issues
The backend is built with Python, FastAPI, Pydantic, PostgreSQL, and SQLAlchemy, while the frontend uses Next.js, TypeScript, Tailwind CSS, and Recharts.
Challenges we ran into
Our biggest challenge was solving the gap between infrastructure language and production language.
Infrastructure systems talk about:
- CPU utilization
- Storage utilization
- Packet loss
- Latency
- GPU usage
- Write latency
- Throughput
But production teams think about:
- Footage
- Scenes
- Editorial deadlines
- Production schedules
- Shooting days
- Post-production
Simply giving an AI model infrastructure metrics wasn't enough. We needed to provide it with the production context required to understand why those metrics matter.
Another major challenge was creating realistic incident scenarios.
We didn't want an incident to simply be:
storage_utilization = 95%
We wanted the telemetry to behave like a real cascading problem.
For example:
$$ Storage\ Utilization \uparrow \rightarrow Write\ Latency \uparrow \rightarrow Ingest\ Throughput \downarrow $$
At the same time, the agent needs to determine whether other systems are behaving normally so it can distinguish storage saturation from network, camera, or other failures.
We also faced the challenge of making AI-driven remediation safe.
We deliberately designed the system so that the AI does not directly modify real production infrastructure. Remediation actions are simulated and clearly labeled, with human approval required before the simulation proceeds.
Accomplishments that we're proud of
We are proud that Production Guardian goes beyond being another AI chatbot placed on top of monitoring dashboards.
The agent can actually investigate observability data through Grafana MCP, correlate multiple signals, reason about the root cause, connect the problem to production context, and verify the system after remediation.
We built a complete end-to-end incident workflow rather than demonstrating isolated AI capabilities.
During the demo:
- We open the NIGHTFALL production dashboard.
- We trigger a Storage Saturation incident.
- Telemetry begins degrading in real time.
- The data is pushed into Grafana.
- We ask the AI agent to investigate.
- The agent investigates through Grafana MCP.
- Gemini identifies the likely root cause.
- The system calculates the production impact.
- The user can see what happens if nothing is done.
- The user reviews and approves remediation.
- The remediation is simulated.
- Telemetry recovers.
- Gemini verifies the recovery.
The result is a system that can move from:
"Something is wrong."
to:
"Here is what is wrong, here is how it affects your production, and here is what you can safely do about it."
That is the accomplishment we are most proud of.
What we learned
One of our biggest learnings was that context is what makes AI operationally useful.
A metric by itself has limited meaning.
For example:
Storage utilization: 94%
is useful information, but it doesn't tell a producer whether today's footage will reach editorial on time.
Once that metric is combined with production context, scene information, workflow dependencies, and deadlines, the system can reason about the actual business consequence.
We also learned the importance of combining AI reasoning with deterministic logic.
We use Gemini for tasks where reasoning is valuable, such as investigation, diagnosis, and generating recommendations.
For production-impact calculations, we use deterministic calculations instead. This gives us the flexibility of AI while keeping important numerical decisions predictable.
Another major learning was working with MCP.
Connecting an AI agent directly to observability tooling changed the way we thought about agentic systems. Instead of giving an agent a static dataset and asking it to explain the data, the agent can actively investigate the environment using tools.
We also learned that autonomous systems need guardrails and human oversight. Even when an AI can recommend an action, allowing it to immediately modify production infrastructure isn't always appropriate. Our approval-based and simulated remediation workflow reflects that principle.
What's next for Production Guardian
Production Guardian is currently a hackathon demonstration using synthetic telemetry, simulated remediation, and demo-calibrated production-impact parameters.
The next step is to move beyond the simulation and connect the system to real production environments.
We want to integrate Production Guardian with:
- Real media servers
- Camera systems
- Production infrastructure
- Post-production workflow systems
We also want to add predictive alerting, allowing Production Guardian to identify patterns that could become production-impacting incidents before they actually happen.
Other areas we plan to explore include:
- Multi-production support
- Historical incident analysis
- Predictive failure detection
- Post-production workflow integration
- More advanced remediation and approval workflows
Our long-term vision is to make Production Guardian the layer between technical infrastructure and production decision-making.
We don't want production teams to have to interpret dozens of dashboards to understand whether a technical issue threatens their schedule.
We want them to receive a clear answer:
What happened? Why did it happen? What will it affect? What happens if we do nothing? And what is the safest action to take?
Production Guardian turns production telemetry into production decisions.
Built With
- ai
- alembic
- docker
- fastapi
- google-adk
- google-cloud-run
- google-gemini
- grafana-cloud
- grafana-mcp
- loki
- mcp
- next.js
- postgresql
- prometheus
- pydantic
- python
- recharts
- tailwind-css
- typescript

Log in or sign up for Devpost to join the conversation.