Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for Premiere Playback investigator

Premiere Playback Investigator Premiere Playback Investigator was inspired by the operational pressure of a major content premiere. When thousands of viewers begin streaming at once, playback failures can escalate quickly, and engineers need to identify the cause before the audience experience—and the launch—suffers.

I built the project to demonstrate how an AI investigation agent can help an operations team move from a vague alert such as “playback is failing” to an evidence-based explanation of what is breaking, where it is breaking, and what responders should investigate next. The project focuses on premiere-night playback 5xx errors, where fast correlation across telemetry is critical.

How I built it The project combines Gemini with the Google Agent Development Kit (ADK) and Grafana Cloud MCP in a Go-based implementation. The agent is intentionally scoped to read-only Grafana access (grafana:read), allowing it to investigate observability data without making changes to production infrastructure.

The investigator can use Grafana Cloud telemetry to examine signals relevant to playback incidents—such as error patterns, affected services, and abnormal behaviour—then turn the findings into a clearer investigation narrative for the operator. I also included a TTS demo trailer builder so the project could communicate the investigation scenario and workflow effectively during the hackathon presentation.

What I learned Building this project showed me that an effective operations agent is not just a chatbot connected to monitoring tools. It needs constrained permissions, a clear investigation goal, reliable access to evidence, and outputs that help an engineer verify conclusions rather than blindly accept them.

I also learned how valuable the Model Context Protocol can be for connecting an agent to a live observability environment in a structured way. It provides a practical path for an AI agent to investigate metrics, logs, traces, and dashboards while keeping the system’s operational boundaries explicit.

Challenges I faced One challenge was defining a focused use case instead of creating a broad “AI observability assistant.” I narrowed the project to premiere-night playback 5xx errors so the agent has a concrete incident scenario and clear evidence to investigate.

Another challenge was balancing autonomy with safety. Because incident-response systems can affect critical services, I designed the agent around read-only investigation rather than automated remediation. This preserves human control over production actions while still reducing the time required to understand an incident.

The project was published as an initial public snapshot for the Grafana hackathon under the MIT license.

Built With

  • 5xx
  • adk
  • agents
  • ai
  • analysis
  • analytics
  • cloud
  • devops
  • distributed
  • gemini
  • go
  • google
  • grafana
  • http
  • llm
  • log
  • metrics
  • monitoring
  • observability
  • playback
  • production
  • sre
  • tracing
  • tts
Share this project:

Updates

Submission history