What it does

GATECHECK is a night-shift render wrangler for a VFX studio. A render farm renders frames of a film overnight, and some come back wrong while being marked successful: the renderer could not find a texture and silently substituted a fallback, or the scene was saved with the wrong colour transform, or the sample count was clobbered by a bad submit, or the file was truncated mid-write. Every one of those exits 0. The queue is green, the shift report says the night went fine, and days later a lighting supervisor asks why the crates are grey. GATECHECK finds those frames, explains what happened, prices the remediation in dollars, and asks a human before it touches anything.

The problem it solves

"Render wrangler" is a real, staffed, night-shift job. The written duties in current postings are to distil render logs for triage, to work out whether a failure came from the scene, the assets, the software, the licensing, the hardware or the pipeline configuration, and to hand the shift over in the morning. One posting notes that a 1% error rate costs hours of work. Existing monitoring covers the failures that announce themselves — a crash, a queue backing up, a node falling over. It cannot see a frame that finished successfully and is wrong, because nothing in the pipeline is looking at the pixels. Meanwhile the money is not where people assume: on published render-farm unit economics, roughly three quarters of the cost of a wasted frame is renderer licence time rather than compute, so retrying a frame that cannot possibly succeed spends the expensive part twice.

How it works

A deterministic verifier scores every rendered frame with no model involved at all: histogram divergence against the shot's own neighbours, collapse of the unique-colour count, edge density, a noise estimate, block variance in the tail of the file, and a self-calibrating baseline fitted with a robust trend estimator. It writes a verdict as a Prometheus metric and a Loki line. A second, independent lane compares what the render sidecar says actually happened against the shot's declared render spec, which catches the defect the first lane cannot see by construction — a whole shot rendered with the wrong colour transform, where no frame is an outlier because they are all equally wrong. The statistics decide where the problem is. The agent, an ADK graph workflow on Gemini 3.8 Flash, reads all of that through Grafana, correlates it with the farm's own queue metrics and the renderer's logs, explains the root cause, prices the options and recommends exactly one action. Then it stops at a graph interrupt and waits for a person.

The part we think is new

The distrust of the agent is not implemented in its prompt. Grafana Agent Observability supports guards: rules stored in the tenant that run inline on the request path and can deny a tool call outright. They are inert until an application calls the hooks endpoint itself and obeys the answer, so GATECHECK calls it before every action that touches the farm. One guard ships with the project and denies a retry whenever the defect class is one that rerunning cannot fix. The same tool with the same arguments is denied for a missing texture and allowed for an out-of-memory kill, and the decision comes from a rule an operator can edit without reading a line of the agent's source. The vendor SDK defaults to failing open so a transport error never blocks a model call; for something that can spend money on a render farm that default is wrong, and we override it. Second, the agent has to show it has been right lately: every evaluation run is published to the same Grafana tenant as an offline experiment, one trial per scenario, graded deterministically against an answer key the agent has no tool that can reach, and it reads that report card back before proposing anything. Run before any evaluation exists, it diagnosed the fault correctly, priced the retry at 85% success and +$10.19 expected value, and then refused itself with "there is not enough evidence that this agent is currently reliable". Third, no language model in the system is ever handed a tool that can act: the write tools exist on the same MCP server, but they are called by a deterministic node after a human approves a specific validated action, and the server re-checks it anyway.

How we built it

mcp-gatecheck is a single Go binary that imports github.com/grafana/mcp-grafana as a library and registers Grafana's own tool implementations unmodified alongside eight render-farm tools of ours — 72 tools on one endpoint — and publishes its own MCP App, a ui:// HTML resource that renders the approval card with both frames, the signals table and the priced options. Every write runs six checks server-side, in a frozen order, after a human has approved: a signed approval token bound to exactly this action on exactly this set of frames; the Grafana guard; the shot's remaining retry budget; a cap on frames per action; a node allowlist; and whether the defect class is one a rerun can fix, where the class is established by the server from the frame's own evidence rather than from the caller. The first denial stops the sequence and the checks it never reached are reported as skipped rather than implied to have passed. Minting an approval is deliberately not an MCP tool: the endpoint takes a different credential and the server refuses to start if the two match, because a process holding both can approve itself. Underneath it all is a real OpenCue farm in Docker — Cuebot, two render nodes, a Postgres, the REST gateway — rendering real Blender frames of a scene built by script, with faults injected into the render itself so they leave exactly the evidence a real failure leaves.

Challenges we ran into

The faults had to be real, and that was harder than expected. Blender loads images lazily, so repointing a texture datablock produced the magenta fallback and complete silence in the log — and forcing the path resolve only works before the repoint, because the failed lookup is cached and a second attempt says nothing. A stale frame had to copy its predecessor's exact bytes, because Cycles is not bit-reproducible and re-rendering the same frame twice produces two files that look identical and hash differently. On the farm side, RQD's user creation collided with an existing uid on the base image and every frame launch failed silently with the work stuck in WAITING, and a container reading the host's /proc/loadavg made Cuebot compute negative idle cores for the second node and quietly stop booking it, so one node did all the work while the farm looked healthy. The detector had its own version of the same lesson twice over: thresholds calibrated on a corpus we generated ourselves turned out to be meaningless on real renders — one comparison could never fire, and an absolute scale floor that was 29% of the signal level on synthetic frames was 83% of it on a real dim interior, flattening the measurement into noise.

What we learned

Both false positives the detector ever produced came from real renders rather than from our own corpus, and both were at a shot boundary where the comparison window is one-sided. Neither was fixed by moving a threshold. The second one is the more interesting: half of that window was itself defective, which dragged the trend fit above the clean level, and the frame missed a proportional cut by 0.003 while sitting well inside the clean population. Requiring the drop to be statistically significant as well as proportional makes the rule cautious exactly where its baseline is weakest, which is the property you actually want. More generally, the honest version of "we tested it" turned out to be worth more than the score: a detector graded on the corpus its own author designed is the easiest thing in the world to make perfect.

What's next

The farm is a demonstration substrate, not the product. The action surface behind the MCP server speaks OpenCue's REST gateway; another queue manager is another client behind the same interface, and the policy engine, the cost model and the agent do not change. The obvious next step is a Deadline adapter and a real studio's render spec, plus per-shot success priors learned from that studio's own history rather than the published table we ship.

Built With

  • agent-observability
  • blender
  • cloud-run
  • cycles
  • docker
  • gemini-3.8-flash
  • go
  • google-adk
  • grafana-alloy
  • grafana-cloud
  • grafana-mcp
  • loki
  • mcp-apps
  • model-context-protocol
  • numpy
  • opencue
  • opentelemetry
  • pillow
  • postgresql
  • prometheus
  • python
  • tempo
  • vertex-ai
Share this project:

Updates

Submission history