Inspiration
Every day, VFX studios, animation houses, and AI engineering teams lose hundreds of compute hours to silent background failures—GPU memory leaks, corrupted texture file locks, runaway LLM tool loops, and stalled render nodes. Engineers are constantly interrupted by low-level telemetry noise, pulling focus away from creative production. We built CineFlow IRM to give teams an autonomous SRE & Technical Director agent that silently handles routine infrastructure incidents in the background, only surfacing when a high-impact decision requires human authorization.
What it does
CineFlow IRM is an autonomous background AI agent built with the Strands Agents SDK and integrated with Grafana telemetry:
- Autonomous Background Monitoring: Continuously tracks render farm nodes (Maya, Houdini, Nuke) and multi-agent AI pipelines.
- Root Cause Investigation: When a telemetry alert fires, the agent queries Loki logs and Tempo distributed traces via Grafana MCP tools to pinpoint exact root causes (e.g. CUDA memory leaks vs 0-byte file locks vs runaway agent token burn).
- Silent Background Remediation: Automatically isolates failing nodes, quarantines corrupted assets, kills runaway loops, and re-allocates render frames without human intervention.
- Human Decision Gate: Only pings human operators when an operational decision requires sign-off, delivering an executive Markdown Incident Postmortem with MTTR metrics.
How we built it
- Orchestration: Built with the Strands Agents SDK for reliable tool use, reasoning, and decision boundaries.
- AI Core & Tools: Powered by Google Gemini and Grafana Cloud MCP Server (
query_loki_logs,get_tempo_trace,resolve_incident,post_postmortem). - Backend & Simulator: Python FastAPI backend with built-in GPU/CPU telemetry and fault injectors.
- Frontend Dashboard: React, Vite, Tailwind CSS, and Recharts for live status tracking.
Challenges we ran into
Correlating unstructured GPU driver logs with distributed OpenTelemetry trace spans in real time, and tuning the agent's confidence threshold so it resolves routine errors silently while reliably escalating complex decisions to human operators.
Accomplishments that we're proud of
- Achieving an estimated 87% reduction in Mean Time to Resolution (MTTR) for common render node and AI pipeline failures.
- Building a complete end-to-end telemetry simulator with real-time fault injection and automated incident postmortem generation.
What we learned
Building autonomous agents that run quietly in the background requires strict decision gating ("AI Recommends, Code Authorizes") so humans retain control without suffering alert fatigue.
What's next for CineFlow IRM
- Deploying with AWS AgentCore for cloud production scale.
- Adding predictive telemetry forecasting to resolve GPU memory leaks before frames crash.
Log in or sign up for Devpost to join the conversation.