A shoot day isn't recovered until it's recovered for everyone.

All-Access is a Gemini-powered multi-agent recovery system for film sets, built with Google ADK on Gemini Enterprise Agent Platform and IBM Bob in the development workflow. It turns a disruption into recovery options, proves each option against hard production and access constraints, requires human approval, and reconciles downstream execution before the day can be declared ready.


Inspiration

A film set changes all day. Weather closes an exterior, a performer is delayed, a generator fails, a location withdraws access at two hours' notice. Weather disruption alone is estimated to cost a production up to $500,000 a day, and 85% of productions hit scheduling problems. A first AD absorbs that by hand, on the phone, from memory.

The thing that gets dropped when a day goes sideways is never the camera. It is the arrangement that took three weeks to approve and exists for one person: the step-free route, the interpreter booked for every called minute, the refrigerated storage at base, the chaperone released at 17:30 to collect a dependant.

That is not a hypothetical. The industry has invented a job to hold it together — the production accessibility coordinator, whose published duties include adapting access plans "as necessary with changing production and accommodation needs." That is a person manually re-validating a set of approved arrangements against a schedule that keeps moving, under time pressure, often not in the room when the schedule changes, and frequently treated as a luxury line item.

Meanwhile the commercial reality is unforgiving in exactly the way software is good at: sign language interpreting agencies enforce 48-hour cancellation windows, bill the full block inside them, and charge rush premiums under five business days' notice. An 18:00 schedule change cannot extend a 09:00 interpreter booking, and no amount of goodwill changes that.

I wanted to build the system where losing one of those arrangements is structurally impossible rather than merely discouraged.


What it does

All-Access turns each disruption into a governed event and runs it through a closed loop:

  1. Ingest. The change arrives as a typed, schema-validated event on a schema-governed, Confluent-compatible event backbone — 23 data contracts, 49 event types, hash-chained, idempotent, dead-lettered on malformed input.
  2. Scope. A bitemporal production digital twin — 208 entities, 699 relationships — computes the blast radius, scoped to the day being planned and ranked by how directly the disruption reaches each consequence.
  3. Assess. Eleven expert agents run concurrently on one thread pool and return typed findings with evidence. An agent with no evidence returns ABSTAINED. No agent can decide feasibility.
  4. Plan. A deterministic engine — 34 constraints, 28 predicates — generates structurally diverse recovery plans, simulates each under uncertainty, and publishes nothing it cannot prove feasible. Plans it cannot build come back as a minimal conflict set: the smallest group of rules that cannot be satisfied together, each naming the document it came from and the person who owns it.
  5. Approve. The workflow blocks inside the coordinator and goes no further until a named authority chooses a plan and signs for every role that plan requires. It blocks for a person: the approval workspace opens on a sign-in, and until a session exists there is no control on the screen that routes a plan. The role is read off that session and so is the signer's identity. There is no field in the request that names either, and an anonymous caller gets 401 at every endpoint that changes anything. A session holds exactly the authority the production directory gives its holder, so a plan needing two signatures needs two people. Approvals are HMAC-signed, single-use, expiring, and bound to both the plan hash and the constraint-set hash; change either and the approval stops applying. Every approval also records how it was obtained: human when a person signed at the workspace, judge for an evaluation session, stand_in for an unattended run such as the benchmark. The data contract requires that field, so an approval that does not declare its channel never reaches the topic, and neither a stand-in nor an evaluation approval may carry a crew member's identifier..
  6. Execute. Typed commands, saga-coordinated, to call sheets, transport dispatch, department queues and crew mobile.
  7. Reconcile. Intended state is compared against observed state. Nobody can call the day ready until they match. There is no override.

The web product turns that loop into 13 role-oriented views—from disruption intake and evidence-backed assessment through plan comparison, human approval, execution, and reconciliation—including an infeasible-plan explorer that shows exactly why a seemingly obvious recovery cannot be published.


The gap this covers

Movie Magic Scheduling and StudioBinder are established scheduling and production-management tools; All-Access addresses a different layer: determining whether a proposed recovery remains feasible under hard operational and access constraints before it can be published.

Movie Magic StudioBinder Phone + spreadsheet All-Access
Breakdown, stripboard, scheduling ● ● ○ ○
Call sheet generation and distribution ◐ ● ◐ ●
Delivery tracking ○ ● ○ ●
Resource conflict detection ● ◐ ◐ ●
Cross-department consequence of a change ○ ○ ◐ ●
Feasibility proved before an option is offered ○ ○ ○ ●
Access arrangements as hard constraints ○ ○ ○ ●
Named reason when no option exists ○ ○ ◐ ●
Reconciliation of intended vs. observed state ○ ○ ○ ●

Four differences carry it:

Conflict detection is not feasibility. Movie Magic's red flags catch a double-booked actor. They warn; they do not stop a schedule being published, and they have no concept of a rule that may not be waived.

Distribution is not arrival. StudioBinder tracks whether a call sheet was delivered and confirmed — genuinely ahead of the field, and still about the document. All-Access reconciles the state it intended against the state downstream systems report, and blocks readiness on the difference.

All-Access’s key distinction is that an approved access arrangement participates directly in feasibility as a hard constraint rather than remaining passive schedule metadata.

Finally, "No" without a reason is not an answer. A conflict marker is not an explanation a location manager can act on.


Inclusion, as architecture rather than as a feature

This is why it is called All-Access — the phrase is the backstage pass and the promise at once: everyone who is called is on the set, and the system is built so a rushed plan cannot quietly revoke that.

The six approved access arrangements in the corpus — a step-free route, interpretation covering every called minute, captioned safety briefings, refrigerated storage at base, a quiet rest space, a caregiving release time — are hard constraints with no soft weight. The objective function is not permitted to trade one against a saved hour, because it has no coefficient with which to do so. A plan that drops one is not a worse plan. It is not a plan.

It fails closed, out loud. Take the hero storm and also remove the only step-free vehicle, and the system returns 0 feasible, 5 rejected — four citing "no step-free vehicle is available to satisfy ACC-001", and one citing the boatshed, which is free, obvious, and exactly what a conventional scheduler would offer:

Ormsvik boatshed has no step-free route from arrival to the working position and no surveyed step-free alternate position.

That is a sentence a location manager can act on. "Accessibility concern" is not. The spatial model works in millimetres and gradients — clear widths, threshold heights, surface, steps — so the refusal always carries the measurement that produced it.

It never records why anyone needs anything. There is no field in the schema for a diagnosis, a condition, or a category of impairment. This is a direct response to a documented failure mode: fear of disclosure suppresses the accommodation requests that get made at all, and performers have supplied their own accessible facilities rather than ask — Daryl Mitchell brought his own wheelchair-accessible trailer. A system that stores a reason creates a disclosure the person did not choose. The transport coordinator is told "step-free transport; accessible vehicle VEH-ACC-1 with powered lift" and nothing else, because dispatching a lift does not require knowing why. A redaction layer enforces audience clearance and per-arrangement visibility, and a tripwire counts any attempt to serialise a prohibited field. Across 52,775 events: zero.

Inclusion here is not a screen. It is a constraint class, a redaction boundary, and a refusal.


The numbers

1,000 disruptions, every figure recomputable from committed artifacts in bench/results/.

Hard-constraint violations in 1,629 published plans 0.000
Access preservation across 3,630 approved arrangements 1.000
Rejected plans carrying a minimal conflict set 1.000
Fabricated constraints across 18,420 findings 0.000
Prohibited personal fields across 52,775 events 0
Replay reproduces live state exactly · hash chain intact 1.000 · 1.000
Incomplete executions caught before closure 75 of 75
Disruptions ending in a named minimal conflict set 0.395

75 of 75 is the argument in one number. The harness creates executions where every command completes and a critical downstream step is still outstanding — a department that never accepted, an update that reached only part of its audience. Reconciliation caught every one. A system that closed the day on command acknowledgment would have declared all 75 ready: 12.4% of the 605 executions in the corpus.

In 39.5% of benchmark disruptions, the system ends with a named minimal conflict set instead of publishing a plan it cannot prove feasible — turning ‘no feasible plan’ into an actionable operational reason.

What the pieces are worth

Same 200 scenarios, one capability removed at a time. Remove the independent feasibility recheck — which is what a plan authored by a language model or a spreadsheet actually is — and:

Full No independent validation
Hard violations per published plan 0.000 2.412
Access preserved 1.000 0.780
False closure 0.100 0.255
Named conflict set instead of a wrong answer 0.400 0.000

Without the independent feasibility recheck, published plans average 2.412 hard-rule violations, 22% of approved plans drop an approved access arrangement, and the named-conflict-set rate falls to zero. The full system prevents those failures by validating independently before publication.


How AI meaningfully improves the workflow

All-Access is a Gemini-powered multi-agent system built with Google ADK and deployed on Vertex AI Agent Engine: eleven specialist agents assess a production digital twin while a deterministic constraint engine retains authority over feasibility.

It improves a workflow. Recovering a disrupted shoot day is a manual job today — a first AD on the phone, from memory, re-checking a moving schedule against a set of approved arrangements. All-Access turns each disruption into a governed event and runs the whole loop — scope, assess, plan, prove, approve, execute, reconcile — in a few seconds, and shows its work at every step.

It enhances decision-making. Every recovery plan carries a feasibility proof, and every refusal carries a named minimal conflict set — the smallest group of rules that cannot be satisfied together, each naming the document it came from and the person who owns it. A call made under pressure becomes an auditable, explainable one: across 1,000 benchmark disruptions, 39.5% end with a named minimal conflict set rather than a plan the engine cannot prove feasible — a reason a location manager can actually act on.r.

It elevates the experience of the people it serves. The customer is the production and its crew, and the promise is that everyone who is called can work. Across 1,000 disruptions and 3,630 approved arrangements, access preservation is 1.000 — not one was ever traded away for a faster day.

It drives a measurable operational outcome. Weather disruption alone costs a production up to $500,000 a day. Reconciliation caught 75 of 75 incomplete executions before anyone could call the day ready — the 12.4% of runs a command-acknowledgment system would have closed early — and published 0 hard-constraint violations across 1,629 plans. Remove the independent feasibility recheck and every published plan breaks roughly two hard rules; the boundary is worth exactly that.

The architectural boundary is deliberate: Gemini assesses evidence and explains consequences, but deterministic code alone decides feasibility. The model improves the decision workflow without being allowed to waive the rules that make a plan safe to publish. That separation gives the agent useful autonomy while keeping the production decision deterministic, testable, and auditable.


How I built it

Google Cloud / Gemini Enterprise Agent Platform. Gemini (gemini-3.7-flash) via google-genai is the reasoning plane for the eleven specialist agents. The hosted agent is implemented with Google ADK as a google.adk.agents.Agent and deployed on Google Cloud's managed Agent Runtime. The hosted agent is a google.adk.agents.Agent with seven read-only function tools over the product's own read model (agents/adk_tools.py); the deployment refuses to run if the implemented tool surface and the approved allowlist ever disagree, and a denylist blocks approve_plan, issue_command, declare_ready and override_constraint explicitly. Deployment targets Vertex AI Agent Engine, Cloud Run, Artifact Registry and Secret Manager, with Terraform in infra/.

The reasoning plane is selectable: the public URL runs the deterministic plane, which cannot be exhausted by an open endpoint. The judging link is the same build with an evaluation session, and every disruption started from it runs the eleven specialist agents on gemini-3.7-flash on Vertex AI.

IBM Bob was used as part of the development process. In the committed Bob session, working in Connector Modernization mode with scoped engineering permissions and only the legacy adapter plus brief, Bob rebuilt the call-sheet connector from scratch without opening the reference implementation. The resulting typed, idempotent, event-driven callsheet_bob.py resolves all eight documented defects with failing-first tests (16/16 green; full suite 275).

Confluent is the event backbone: ConfluentEventBus and ConfluentSchemaRegistry register 23 subjects, validate every payload before append, enforce BACKWARD compatibility and dead-letter failures. It is built for Confluent Cloud — the schema-registry client speaks the Confluent Schema Registry REST API, and its compatibility enforcement is unit-tested (an added optional field is accepted; a new required field or a narrowed type is rejected). Moving from the hosted demo to a live cluster is a single environment switch — AA_STREAM_MODE=confluent plus the cluster credentials — with no code change; the deployment ships with a byte-compatible local drop-in bus so judges can run the entire event-driven loop with no credentials at all.

What I learned

Agentic does not have to mean unconstrained. The strongest architecture was to give Gemini room to synthesize evidence and explain consequences while keeping feasibility behind deterministic predicates.

Accessibility fails when it is treated as metadata. Once an approved access arrangement participates directly in feasibility, the optimizer has no mechanism to quietly trade it for time or convenience.

Command success is not operational success. Catching 75 of 75 incomplete executions showed why the workflow must reconcile intended state against observed state rather than treating successful command delivery as completion.

Letting a language model into a production decision system

The model plane is allowed to write, and then its writing is checked by deterministic code.

agents/core.py::ungrounded_claims reads every response for identifiers (C-ACC-002, LOC-BOATSHED) and quantities (45 mm, 18:30, 68 kph) that do not appear in the facts the model was given. A response carrying an invented constraint id or an invented threshold is discarded in favour of the deterministic text, and the rejection is recorded and shown in the product.

It cannot establish that the prose is true — no regular expression can. It establishes that every checkable token came from the input, which is exactly the class of detail a fluent paragraph carries convincingly and a reader cannot verify. That is a control that does not ask the model to behave; it checks whether it did.

The load-bearing boundary sits between the two planes. engine.publish takes its verdict from engine.validate, which runs constraints.registry.evaluate, which runs 28 predicates. No model output reaches it. That is the call graph, not a policy — and the ablation table above is what it is worth.

The boundary is measured. bench/reasoning_plane.py runs the same disruptions on both reasoning planes and diffs the results. Across 40 disruptions on gemini-3.7-flash, all 40 produced identical feasibility outcomes across the Gemini and deterministic control planes, while Gemini produced grounded explanatory language in 106 of 440 findings. That separation is intentional: Gemini changes what a human receives as evidence-backed assessment and explanation; it cannot change what the constraint engine is permitted to publish. The benchmark fails if either guarantee is lost..

Reasoning tokens are the real cost of a Gemini 3.x plane. The model spends them before it answers and they draw on the same output budget: 50,572 of them against 8,178 tokens of finished sentence, close to six to one. The output budget is therefore sized for the thinking rather than for the answer, the current Flash model is addressed on the global endpoint it is served from, and a response that reaches the budget is treated as a failed call and replaced by the deterministic text. A partial sentence about a safety finding is worth less than the template it would have replaced.

Eleven Gemini assessment calls in series would add too much latency to a time-critical recovery workflow, so the specialist round runs concurrently: 17.2 s falls to 9.9 s per disruption on the Gemini plane over the same eight disruptions.. The event log is identical either way. Results come back in the order the agents were declared, events are emitted in that order from a single thread, and the reasoning ledger is sorted back into it, so AA_ASSESSMENT_WORKERS=1 is a speed switch rather than a second system. tests/test_parallel_assessment.py asserts both halves: that the agents overlap in time, and that the overlap changes nothing about what they produce.


Built With

Share this project:

Updates

Submission history