Inspiration
AI agents are becoming increasingly capable at reading operational signals, reasoning about risk, and proposing actions. In production media workflows, however, intelligence alone should not become execution authority.
BOSAI Studio Control Plane was built around one thesis:
Intelligence is not authority.
The project explores how an AI-powered studio workflow can separate observation, reasoning, authorization, execution, verification, and proof so that a capable model can help recover a production pipeline without silently gaining permission to mutate it.
What it does
BOSAI is a governed control plane for autonomous studio operations.
The demonstration scenario is a final-trailer delivery workflow. A delivery SLA is at risk. Grafana provides operational telemetry. Gemini, through Google ADK, reads the evidence and proposes a recovery action such as restarting a transcode worker.
The proposal does not execute anything.
At proposal time there is:
- no permit;
- no mutation;
- no authority claim.
BOSAI evaluates the proposal deterministically. If policy allows the action, BOSAI issues a single-use permit. If a safety invariant is violated, the action is denied. Execution can only occur through the authorized path, and replay of a consumed permit is rejected.
The canonical workflow is:
OBSERVE → REASON → PROPOSE → AUTHORIZE → EXECUTE → VERIFY → PROVE
The public judge experience exposes three deterministic outcomes:
- governed transcode-worker restart → AUTHORIZED → execute → verify;
- unsafe QC bypass → DENIED → no permit / no mutation;
- consumed-permit replay → DENIED → replay blocked.
The browser experience is explicitly a recorded evidence replay, not a fake live backend. It does not call private cloud services, consume a real permit, or mutate a pipeline. Real runtime evidence is preserved separately in the public repository.
Hosted judge experience: https://arthur9293.github.io/bosai-studio-control-plane/
Demo video: https://youtu.be/bNyq_NPUYco
Open-source repository: https://github.com/Arthur9293/bosai-studio-control-plane
IBM Partner Track
BOSAI is submitted for the IBM Partner Track.
IBM Bob was used as part of the development process in both Plan and Agent modes. Bob inspected the contest repository, produced a bounded compliance-gap analysis, helped refine the judge-facing README and presentation layer, and contributed human-reviewed edits to the judge surface and its tests.
IBM Bob is deliberately not presented as runtime authority. It cannot issue BOSAI permits, execute mutations, or bypass policy. The repository contains a dedicated public evidence record of the Bob session, inspected files, accepted contributions, rejected/corrected suggestions, and regression evidence.
This distinction is important to the architecture:
- IBM Bob helped build and refine the contest project;
- Gemini / Google ADK provides proposal-only runtime intelligence;
- BOSAI owns deterministic authorization;
- Google Cloud IAM enforces runtime service boundaries.
How we built it
The architecture is intentionally separated:
- Grafana observes — operational telemetry is available through a read-only Grafana MCP path;
- Gemini proposes — Google ADK + Gemini consume evidence and produce a constrained proposal envelope;
- BOSAI authorizes — deterministic policy evaluates the proposal and can issue a single-use permit;
- Firestore persists — durable authority, permit, execution, and audit state is retained;
- Cloud Run / IAM enforces — runtime service-to-service boundaries constrain which component can invoke the execution path;
- Verifier proves — postconditions are verified from authoritative state and run-bound evidence rather than model prose.
The Gemini implementation uses Google ADK in Vertex AI mode and is intentionally proposal-only. The model sees a minimal read-only operational tool surface and receives no direct mutation tools.
The runtime boundary was designed so the control plane may invoke the authority executor, and the authority executor may invoke the simulated media pipeline, while a direct control-plane-to-pipeline mutation is not a valid edge.
Data sources and workload boundary
The project uses a synthetic media-production workflow, not a customer production workload.
The evidence path includes:
- synthetic media-delivery operational events and telemetry;
- Grafana Cloud / Loki readback through the official Grafana MCP path;
- Firestore durable authority and audit state;
- run-bound verification evidence generated by the contest implementation.
No customer dataset, customer workload, secret value, or public identity token is required by the hosted judge experience.
Contest provenance and clean-room boundary
The Agentic Cinema contest began on July 27, 2026. This repository was created on August 9, 2026 during the contest period. Its initial commit contained only the MIT license and a short repository description.
The broader BOSAI governance thesis predates the contest, but the BOSAI Studio Control Plane implementation submitted here was built inside this contest repository. No source file from another BOSAI application repository was imported into this project, no pre-existing BOSAI deployment is required to run it, and the media workflow, tests, runtime topology, evidence registers, and judge experience were implemented in this repository during the contest period.
The chronological Git history and evidence registers are public so judges can inspect that build sequence directly.
Challenges we ran into
The hardest engineering problem was not generating an AI recommendation; it was proving that the recommendation could not silently become authority.
That required separate controls for:
- proposal versus authorization;
- permit issuance versus permit consumption;
- direct execution versus controlled execution;
- fresh evidence versus stale evidence;
- safe recovery versus trajectory-level safety invariants;
- successful execution versus verified postconditions.
A second challenge was judge accessibility. The real runtime evidence depends on private cloud credentials that should not be exposed publicly. We therefore separated the credential-free interactive replay from the real contest-period runtime evidence, and made that boundary explicit rather than pretending the hosted page was a live privileged control endpoint.
Accomplishments
The project demonstrates both positive and negative authority paths rather than only a happy-path AI demo.
Key accomplishments include:
- Gemini proposal-only integration through Google ADK;
- deterministic BOSAI authorization with single-use permits;
- replay denial and QC-bypass denial;
- durable Firestore authority state;
- Cloud Run / IAM positive and negative invocation boundaries;
- real Grafana Cloud telemetry and read-only MCP evidence;
- postcondition verification based on evidence rather than model claims;
- IBM Bob development-process evidence for the IBM track;
- a public, credential-free interactive judge surface linked to the underlying proof trail;
- a public MIT-licensed repository with reproducible tests and documentation.
What we learned
Three findings shaped the final product:
1. Agent intelligence and execution authority should be separate systems. A strong model can propose the right action and still be the wrong component to authorize it.
2. Negative-path evidence is product evidence. Showing that unsafe QC bypass and consumed-permit replay are denied is as important as showing a successful restart.
3. Verification must be tied to authoritative, current evidence. A model saying “the recovery succeeded” is not sufficient; stale, missing, or mismatched evidence should fail closed.
We also learned that judge experience matters for governance products. A secure system should be inspectable without forcing judges to receive privileged cloud credentials, so the public interaction must clearly distinguish deterministic replay from real runtime proof.
What’s next
The next step is to extend the same governed pattern to additional studio operations and multi-step recovery workflows while preserving the same invariant: AI can reason and propose, but authority remains explicit, bounded, auditable, and verifiable.
Log in or sign up for Devpost to join the conversation.