Inspiration for BluePages : Git for Scripts

On a film set, the script changes constantly. Not once a week: often overnight, during production, while the crew sleeps. Every revision quietly invalidates the element lists that a dozen departments work from. Props has a list. Wardrobe has a list. Transport, locations, casting, stunts, art. All of them derived from a version of the script that is now out of date.

Somebody has to work out what changed and tell each department what it means for them. That somebody is the 1st Assistant Director, and they do it by reading both drafts side by side, by hand, at night, because the revision arrives after the shooting day wraps and is needed before call time the next morning.

It takes hours. It is repetitive. It is also judgment-heavy, which is why nobody has automated it. And when something gets missed, a department shows up without the thing they needed, and a lost shoot day costs tens of thousands of dollars.

That is the exact shape of work worth handing to an agent: it recurs on a schedule nobody controls, it is mechanical at the bottom and judgment at the top, and the human's real contribution is the decision at the end, not the four hours of comparison before it.

What it does

A revised draft lands in the watched folder. Bluepages wakes on its own, reads both drafts, works out what actually changed, reasons about what those changes mean, and routes the consequences to each department in that department's own vocabulary. The AD reviews one screen and approves. Only then does anything go out.

The obvious version of this is a text diff. We built that in the first week, and it taught us the real problem: a mechanical diff is not just insufficient, it is actively worse than nothing. A department head who opens a wall of string changes stops reading them by the second revision.

Here is what a diff sees, versus what the production needs to know:

The diff sees The real question Who pays
"brass letter opener" gone from sc. 3, present in sc. 7 Same object relocated, or a cut and a separate new buy? Props: a continuity note, or a purchase order
JANITOR becomes CUSTODIAN Renamed role, or a genuinely new one to cast? Casting: a day player is worth thousands
"hands her the envelope" becomes "slides it across the table" Same prop, same cast, but a different camera setup The AD, not Props, and Props should not be told
INT. DINER - DAY becomes INT. DINER - NIGHT Not a prop change at all Scheduling, and possibly a location re-quote

Four words in a scene heading can cost a day. A full page of rewritten dialogue can cost nothing. No amount of string comparison separates those two, and that separation is the entire product.

Beyond routing, the agent sources a real priced product when a revision introduces an element a department does not have, does the budget arithmetic against that department's remaining line, and records what it decided and why.

How we built it

A draft lands in the watched bucket. The upload is the entire interface: nobody opens an app, nobody clicks run.

  1. Parse. .fdx through lxml is the primary path, not a shortcut. It is what productions actually use, and it gives typed elements, stable scene numbers and native revision marks. PDF through pdfplumber is tier two, driven by position, because screenplay PDFs encode structure as indentation: scene headings at 1.5", dialogue at 2.5", character cues at 3.7".

  2. Align on scene number. Scene numbers are stable across drafts by industry convention, which is exactly why a cut scene is marked OMITTED rather than deleted: deleting scene 34 would renumber everything after it.

  3. Diff mechanically. difflib within each matched pair produces change spans and relocation candidates. Note the word candidate. The mechanical layer never claims an element moved, because asserting that would bake an answer in below the layer capable of reasoning about it.

  4. Extract. A bulk-tier model reads each changed scene for props, cast, vehicles, wardrobe and clearance risks.

  5. Reason. A judgment-tier model takes each change with its scene text and decides what it means and who needs to know.

  6. Route to twelve department agents in parallel, each rewriting the same finding into that department's own language.

  7. Stop. Everything persists unapproved.

That last step is load-bearing. An agent that wakes on its own and emails twelve department heads unasked is precisely the thing that gets switched off on day two. The approval gate is the same code path whether a human or an upload started the run, and the send step only ever sees what approve touched.

On Strands. Strands' Agent drives every model call. What mattered to us was putting the agent underneath our cost guards rather than beside them. AWS has no hard spending cap and billing lags by hours, so an unbounded agent loop is not a bug you find the next morning, it is a bug you find on an invoice.

  • max_tokens is BedrockModel config, so no code path can omit it.
  • The per-run call ceiling is a BeforeInvocationEvent hook.
  • Agents are constructed per call: stateless, no tools, no history. Each call judges one scene on its own evidence, and a reused agent would accumulate conversation the next scene has no business seeing.

We also chose not to use a multi-agent primitive for the department fan-out. Swarm and GraphBuilder are for dependent hand-off work; department reports are N independent, stateless rewrites of an already-routed finding. A thread pool is the honest shape for that, and Strands still drives each individual call.

Challenges we ran into

Picking the wrong Strands hook would have silently defeated the cost ceiling. We hung the per-run limit on BeforeInvocationEvent rather than BeforeModelCallEvent, because the model-call event does not fire for structured_output. Getting that backwards would have left every structured call unbounded, which is the exact failure the ceiling exists to prevent. It would not have thrown an error. It would have shown up on a bill.

There is no public corpus of the same screenplay at draft N and N+1. Nothing to train against, nothing to evaluate against. We had to author the revision pairs ourselves, by hand, with a labelled answer key.

Confident wrong findings turned out to be worse than missing ones. We added a pass that rejects any finding naming a scene the diff never flagged. Then we noticed some of those rejections were legitimate: the model had spotted a real consequence in a neighbouring scene. Rather than silently lose them, a cheap verification pass now re-reads the actual scene text and asks whether the line supports the hunch. It either earns its place with a quote, or it goes.

Deployment taught us things the test suite could not. Supabase's direct connection is IPv6-only and our container host had no IPv6 egress, so every database call failed with "network is unreachable". The fix was the transaction-mode pooler, which then surfaced a second problem: a pooler multiplexes backend sessions, and psycopg's automatic prepared statements collide across them. One connection parameter, findable only by putting it on the internet.

Resisting the fake feature. A revision that introduces a new element genuinely implies an acquisition, and it was tempting to let the agent "place the order" for the demo. We did not. Bluepages has no vendor account and no payment rail, so a record claiming a purchase happened would be a fabrication.

Accomplishments that we're proud of

We built a ground truth and then held ourselves to it. The answer key records what changed, which department each change belongs to, and a must_not_say list per change: the specific wrong conclusions that would be expensive on a real production. Scoring runs three ways, for recall, for judgment, and for forbidden phrases. Semantic output with no ground truth to check against is not verified, it is just plausible.

Cost is an architectural property, not a hope. Extraction and reasoning scope to the scenes a revision actually touched, so cost tracks the size of the revision rather than the length of the screenplay:

small pair feature-length pair
scenes 9 120 to 121
scenes changed 6 4
model calls ~8 11
cost ~$0.04 ~$0.06
wall clock a few seconds 31s
answer key 8/8, 0 forbidden 8/8, 0 forbidden

The feature-length row is a single measured run, reproduced by a live test that asserts the call count directly rather than leaving it as a claim. A regression that stops scoping extraction to the changed scenes fails that test before it fails the AWS bill.

544 tests, 12,500 lines of Python across 47 modules, a 5,800-line React console, twelve department agents, and a live deployment, with a complete headless path so every screen has a CLI equivalent.

What we learned

A model handed a whole draft and asked what changed will summarize and invent. A model shown one scene's diff will not. Almost every quality problem we hit traced back to giving the model too much context and too little evidence.

Where you put a guard matters more than whether you have one. The cost ceiling, max_tokens, and the fallback chain all had to sit at layers where no code path could route around them. A guard that is merely present, but skippable, is not a guard.

Domain convention is load-bearing. The single most important technical fact in this project is that screenplay scene numbers stay stable across drafts because cut scenes are marked OMITTED rather than deleted. Everything aligns on that. We did not invent it, we read about how productions actually work, and the whole architecture rests on it.

Fallback is for resilience, not economy. Our chain triggers on throttling, timeouts and 5xx only, never on schema or prompt errors, which fail identically on every provider and would otherwise burn three quotas on one bug.

What's next for Bluepages

Get it into a real production office. Everything so far is validated against revision pairs we authored. The next real test is a production's own drafts, with their own conventions and their own mess.

Scheduling that solves rather than surfaces. Today the agent reports what moved on the board: day-to-night flips, omitted and inserted scenes, what each one touches. It deliberately stops short of being a constraint solver. That is the natural next layer, and it is a project in itself.

Learn each production's routing. Every production divides responsibility slightly differently. The department map is currently ours; it should become theirs, corrected by what the AD reroutes and remembered across drafts.

Close the purchasing loop, properly. With a real vendor integration and an approval that maps to an actual purchase order, the proposals become orders. The gate stays exactly where it is.

Video dailies. The same diff logic applied to what was actually shot versus what the script said, which is a different and harder alignment problem.

Built With

Share this project:

Updates

Submission history