Inspiration

Turnaround is the number of hours between finishing one day's work and being called for the next. There are legal minimums, and they exist because people died driving home.

Rico Priem, a grip in IATSE Local 80, wrapped a fourteen-hour overnight on 9-1-1 at around four in the morning on 11 May 2024. At 4:27 a.m. his vehicle left the northbound 57 and flipped. Twenty-seven minutes between wrap and the crash. The CHP recorded it as leaving the road for unknown reasons, and I am not going to claim more than that.

Brent Hershman, a second assistant cameraman on Pleasantville, finished a nineteen-hour day on 6 March 1997 and fell asleep on the Century Freeway. He was

  1. He left a widow and daughters aged 8 and 3. His lawyer said he had worked some 70 hours in five days.

During the 2021 AMPTP negotiation, cinematographers wrote to IATSE about "the hazards of unsafe working hours... Most notable are the numerous car accidents our colleagues have suffered in recent years, including the weekend before we entered these negotiations." The deal signed on 16 October 2021 set a ten-hour daily turnaround.

Underneath the union agreements, California sets a floor in statute. Wage Order 12-2001 §3(F): "No employee shall be required to report to work unless ten (10) hours have elapsed since the termination of the previous day's employment." That one binds a production with no union agreement at all.

The objection, and what survives it

The obvious objection is that turnaround checking already exists, and it does. Obelisk's Grace refuses call-sheet approval on short turnaround. Filmustage flags violations before the sheet goes out. Wrapbook bought Cinapse for it.

Every one of them checks on demand. You hand it a sheet, it tells you about that sheet. Nobody watches the rebuild — the 9 p.m. edit that turns a legal day into an illegal one while everybody is looking at the weather.

Before writing the checker I put the question to IBM Bob: what are the general shapes of "damage lands somewhere you were not looking" in a call-sheet rebuild? It named four — the Phantom Crew, the First-Call Trap, the Borrowed-Time Domino, the Wrap-Time Cascade — and judged the Phantom Crew least obvious, because catching it "requires you to consult a document that the rebuild process never touches." That transcript is in the repo, and the two failures the demo turns on are its first two shapes.


What it does

Callsheet holds one production day — HOLLOW BAY, day 34 of 48 — in two states: as planned at six o'clock, and as reissued at nine after rain took the pier and a night interior came forward.

As planned: zero violations across seven people. As reissued: four.

  • Marcus Reyes gets 11.08h rest against SAG's 12h. The sheet prints 12.58h, because it measures to the crew call at 09:15. SAG rest runs to the first call of any kind, and his hair and makeup call is 07:45. The right number is on the sheet; it is in the wrong column.
  • Vero Aslan, the same shape: 11.33h.
  • Ava Sandoval is 14. She is at the place of employment for 9.92h against a 9h ceiling for her age band, under 8 CCR 11760.
  • The pre-rig crew rig from 22:00 to 03:00 and are called back at 08:45 — 5.75h rest against the statutory 10h. Their call does not exist on the planned day, and they are not on this call sheet. Nobody edited them because there was nothing to edit.

Then the agent looks for a lawful version of tomorrow. It can only propose a repair it has run back through the schedule, so nothing it says is a guess. On the recorded run it tried eleven repairs; two came back clean. It chose to drop scene 78 over scene 42 and said why: 78 is the scene the rebuild pulled forward, so dropping it reverses the edit and the work returns to tomorrow. 42 was always meant for today, so dropping it throws work away.

The card ends with what goes on the record if they publish anyway. A production is allowed to take a penalty. This just makes it a decision instead of a surprise.


How I built it

The rules are code, not a model. Every rule in callsheet/rules.py is a Python function carrying its own citation — Wage Order 12 §3(F), the SAG-AFTRA Codified Basic Agreement, 8 CCR 11760 with all six minor age bands. If something tells you it is legal to call a fourteen-year-old, you need to be able to check its working. "The language model said so" is not an audit. 33 tests cover them.

The model does judgement, not arithmetic. Gemini 3.5 Flash on Vertex AI, orchestrated with Google's Agent Development Kit as a three-step agent — read → repair → card. The repair step calls try_fix(), which lays the day out again and re-runs every rule; the agent only ever reports repairs the deterministic engine has already confirmed.

IBM Bob shaped the design. Bob runs as an Agent Client Protocol server over stdio, driven by a small JSON-RPC client in bob/drive.py. Every exchange is logged to bob/transcripts/.

Thinking budget, split. ThinkingConfig(thinking_level="low") on the tool steps and the default on the judgement steps. Low everywhere made a run much faster and the answer measurably worse.

Deployed on Cloud Run. Front end is plain JavaScript, no framework.


Challenges I ran into

The agent picked the wrong right answer. Both drop 78 + cancel pre-rig and drop 42 + cancel pre-rig test clean, and it kept choosing 42. The fix was not a better prompt — it was that the cost of a repair was not in the data. _costs() now prices a pulled-forward block as a reversal and anything else as lost work, and the agent reasons from that.

The gap sign was inverted for ceiling rules. A 14-year-old 55 minutes over her permitted hours was reported as 0.92h under her limit, because the same subtraction serves a floor rule and a ceiling rule. Split into gap (magnitude) and over (direction).

ADK leaks the MCP child process. Runner.close() is the only path to MCPToolset.close() — there is no __del__, no atexit, no weakref.finalize fallback. Every run loop is now wrapped in try/finally: await asyncio.shield(runner.close()), shielded so cancellation cannot skip it. Verified: one orphan before, zero after.

Bob would not run headless at first. bob run refuses without BOB_API_KEY, and bob chat needs a human at a terminal. bob acp speaks the Agent Client Protocol over stdio and accepts the stored SSO session. The prompt reply carries only a stopReason; the text arrives beforehand as session/update notifications, so the client has to pump both.

Research killed my first premise. I had assumed nobody checked turnaround automatically. Three products do. What I found instead was the actual gap — they all check on demand, and the dangerous moment is the rebuild. I also had the wrong job title: the 2nd AD builds the call sheet, not the 1st.


What I learned

Nearly every real defect was in my design, not my code. The agent choosing scene 42 was correct behaviour on incomplete data. The inverted gap was one subtraction serving two kinds of rule.

And a rules engine is worth more than a clever model here. The model's job is to decide what matters and explain it. The arithmetic that decides whether a fourteen-year-old can be called tomorrow should be something a lawyer can read.

What's next

Read a real call sheet PDF rather than a modelled one; wire the rest of the minor provisions (school hours, studio teacher ratios); watch a live scheduling system and fire on the edit rather than on a page load.

Built With

  • cloud-run
  • fastapi
  • gemini
  • ibm-bob
  • vertex-ai
Share this project:

Updates

Submission history