Current readiness — September 10, 2026

This is an unfinished contest candidate, not a completed Agentic Cinema entry. The hosted page is a synthetic replay. Live credentialed Gemini/Google Cloud plus Grafana runtime, corrected factual-evidence handling, and final contest submission remain outstanding. The public 1:58 demo is available on Vimeo. OpenAI Codex assisted preparation and review; the organizer's AI-tool restriction needs clarification before compliance can be asserted.

Judge quick start

Open the hosted judge experience and run all three canonical incidents. The fastest path is to compare the two conclusive cases with the intentionally ambiguous review-delay case. The key behavior to inspect is not just whether the agent can produce an answer, but whether it refuses to overstate a root cause when the evidence ledger does not support one.

For technical review, the public repository contains the Google Agent Development Kit + Gemini-compatible runtime path, official Grafana MCP integration, synthetic Grafana/Loki environment, tests, evaluation receipts, and reproduction instructions. The hosted experience is intentionally labeled Verified synthetic replay so it does not imply a live credentialed Grafana Cloud or Gemini session.

Why Grafana is essential to the design

Grafana is not a decorative dashboard layer here; it is the agent's evidence boundary. Metrics, logs, dashboards, and alerts are queried through the official Grafana MCP path, successful read operations are recorded in an evidence ledger, and the final response is checked against that ledger. This makes observability data part of the reasoning contract rather than something shown after the fact.

Inspiration

A film or episodic review session can be lost because a dailies pipeline quietly failed hours earlier. Post-production teams may have metrics in one place, logs in another, alerts elsewhere, and only minutes to assemble a defensible diagnosis before editors and producers are waiting.

The problem is not generating a plausible explanation. It is proving which observations support it—and refusing to manufacture certainty when the evidence is incomplete.

What it does

Dailies Guardian is an evidence-first incident-response agent for film and episodic dailies workflows. It helps an operator investigate delayed ingest, transcode, and editorial-review packages without giving the agent permission to mutate production systems.

The production runtime is designed to:

accept a narrowly scoped incident question for a fictional production; use a Gemini agent orchestrated with Google Agent Development Kit; query the official Grafana MCP server in write-disabled mode; correlate metrics, logs, dashboards, and alerts; associate citations with successful tool calls and absolute UTC windows; and return a concise brief separating observed facts, bounded inference, unknowns, ownership, and a reversible next action.

The current checks reject missing, failed, or mismatched tool/window citations. They do not yet prove that every factual claim is supported by the returned values; an unsupported observed-fact sentence was accepted during local review. One of the three canonical incidents is intentionally inconclusive so the agent must abstain rather than invent a root cause.

Why this is useful

Post-production failures are expensive because they consume scarce review time and force specialists to reconstruct context under pressure. Dailies Guardian narrows that work into an auditable first response: what was observed, what it may mean, what remains unknown, who should own the next check, and what can be done without making the incident worse.

It is deliberately not an autonomous remediation system. The product's operating rule is: observe, recommend, escalate.

How it is built

Gemini and Google Cloud

The agent is implemented with Google Agent Development Kit and Gemini-compatible Google GenAI runtime configuration. Its system instruction requires a specific evidence sequence, calibrated language, explicit unknowns, and reversible recommendations.

Grafana partner integration

The repository launches the official grafana/mcp-grafana server v1.0.0 with writes disabled. A source-pinned disposable Grafana 12.1.0 and Loki 3.5.1 environment provisions synthetic data sources and a three-panel dailies dashboard. An ephemeral Viewer token was used to exercise real MCP operations including data-source discovery, dashboard search and retrieval, and Loki queries. The application exposes only five allowlisted read operations.

Evidence ledger

The runtime records eligible tool calls and successful responses. Prometheus and Loki observations are valid only when the call includes a well-formed absolute UTC interval. The final brief is checked against this ledger; a fabricated, failed, relative-window, or mismatched citation is rejected.

Synthetic incident suite

Three deterministic cases model:

checksum retries producing an ingest backlog; transcode workers reaching their configured limit while queue depth and P95 duration rise; and an ambiguous review delay where the available evidence does not prove lateness or causation.

All production names, services, people, and telemetry are fictional. No employer, customer, household, personal recording, credential, or private-network data is used.

Hosted judge experience

The public hosted experience is a transparent replay of the checked-in, hash-bound synthetic evidence suite. Judges can select each incident, inspect the synthetic trend and evidence counts, replay the verified brief, and review the architecture and safety boundaries.

It is intentionally labeled Verified synthetic replay. It does not pretend to be a live Grafana Cloud or credentialed Gemini session. The open-source repository contains the complete Gemini + Google ADK + official Grafana MCP runtime path, fixtures, tests, container configuration, and reproduction instructions.

Validation

The public repository includes:

64 policy, privacy, fixture, integration, export, evaluation, release, API-contract, and configuration tests; a no-operation MCP discovery check against the official Grafana server; a disposable-stack MCP verification using an ephemeral Viewer token; a deterministic 100-point evaluation rubric; three canonical fixture receipts scoring fixed example briefs, not live model responses; the receipt source attribution also needs correction before it can be used as runtime evidence; a non-root container smoke test; a credential-free CI workflow that rebuilds the allowlisted public release and verifies its manifest hashes; and desktop and mobile checks for the judge-facing interface.

The scoring contract weights evidence integrity most heavily: evidence contract 40 points, case-specific evidence 30, uncertainty calibration 20, and reversible action/escalation 10. These fixed-prose scores are contract examples, not measured model accuracy or judge scores.

Challenges

The hardest design problem was preventing a fluent model response from outrunning its evidence. Merely instructing an agent to cite tools is not enough. The application needed a runtime contract that observes actual function calls and responses, validates time windows, and checks citation membership in the successful-call ledger. Matching factual claim values to actual returned evidence remains unfinished.

The second challenge was creating a realistic demonstration without using confidential production data. Deterministic fictional fixtures allow the integration, uncertainty behavior, and failure modes to be tested and reproduced publicly.

What I learned

Observability can be more than a dashboard layer; it can act as a decision-evidence plane for an agent. The most trustworthy result is sometimes not a diagnosis but a well-formed abstention that explains exactly which evidence is missing and requests the next bounded query.

What's next

The next production step is a credentialed Google Cloud deployment connected to a dedicated Grafana Cloud sandbox, followed by evaluation with post-production operators using synthetic or expressly authorized data. Remediation would remain human-approved and outside the agent's write surface.

Built With

Share this project:

Updates