Inspiration

A three-minute teaser is a stepping stone for film and television works, and also a touchstone for newcomers to this industry. It may actually be composed of mostly unnamed cutaways: ocean, steam, hangar air, a slow nebula. Those shots do not need to win a beauty contest against Veo. They need to exist in volume, look like contemporary digital capture, and land in the timeline without burning the month’s Flow Quality credits.

Google AI Pro / Flow is about $20/month. That budget is priced for a handful of Quality hero shots, on the order of ten, not for the ~36 five-second pads you throw away while cutting three minutes. People upgrade to Ultra (~$100/month) because they run out of count, not because every locked-off pad must look like a hero.

For building a reliable b-roll generation workflow benefiting who have limited budgets, we already had two other pieces of the stack: ProDocuX (Word as evidence, SHA-256, no LLM in the parse) and pdx-artifact-engine (typed tool request / result contracts). The missing product was the factory: take a prompt sheet, spend Kaggle T4×2 (~30 GPU hours/week), and come back with a checksummed B-roll library so the $20 stays on the shots that actually need Flow.

Site: prodocux.com · Code: github.com/prodocux/b-roll-library-generator

What it does

B-roll Library Generator turns a Word prompt sheet into a private Kaggle T4 job, then fetches 5-second, 1024×576 locked-off T2V pads with SHA-256 manifests.

How you use it:

  1. Write prompts in Word, one paragraph per shot, ----- between shots, a trailing 5 seconds stored as duration metadata.
  2. check verifies Python, PyPI pins, and that your Kaggle token works. It prints the username and folder, never the key.
  3. run (or parse → build-notebook → push) submits a private kernel (T4, 12-hour timeout). Weights download on Kaggle, not on your laptop.
  4. status / fetch bring back MP4s plus a pdx_tool_result_v1 record.

What it is not: a Veo replacement, a public live kernel, or image-to-video at 1024×576 on T4. Embedded stills in the Word sheet are recorded, then skipped - that path OOM’d on the card we actually measured.

The repo ships ten real pads in examples/broll-10 (ocean through factory steam, nebula, distant galaxy) so judges can watch the output and verify checksums after clone. The public notebook is demo_three_prompts.ipynb. Generated job.ipynb files from real sheets are gitignored so customer prompts never land in git.

How we built it

The pipeline is boring on purpose:

Word .docx
  → ProDocuX DOCX profile (evidence, SHA-256, no LLM)
  → job_manifest.json + pdx_tool_request_v1
  → Kaggle notebook (LTX-2.3 Q3 T2V 512×288 + hires ×2)
  → kaggle kernels push (always private, T4)
  → fetch MP4 + pdx_tool_result_v1

The video generation model is LTX-Video-2.3 Q3 via airesearch-official/free-aistudio. The recipe we kept is the one that finished batches on a Kaggle T4: 512×288, 121 frames, 24 fps, 8 distilled steps, then ltx-2.3-spatial-upscaler-x2-1.1 (scale 2, 4 steps, denoise 0.35) → 1024×576. About 7.5 minutes per clip. txt_cfg / distilled guidance stay at 1.0 for this path.

Local install is a venv plus PyPI pins (prodocux==0.3.0rc4, pdx-artifact-engine==0.3.0a4). The CLI is python -m broll_library.cli. Kaggle is the GPU; this computer only holds credentials and the Word sheet.

Hugging Face ZeroGPU is optional polish (a few stills, a voice pass, a short music loop). It is not a second video farm: free ZeroGPU is minutes per day, and stills cannot feed T4 I2V 1024.

Challenges we ran into

VRAM vs wishful resolution. Native 1024 T2V and I2V 1024 on this T4 Q3 path do not fit. We had to measure a 512 + hires recipe instead of promising a number from a model card.

Chasing a newer checkpoint. Newer LTX builds looked like the “hackathon upgrade.” On this hardware they were slower and did not finish a full Word batch. The product is a library you can fetch, not a blog post about a weight.

Kaggle’s free GPU is T4×2 - two 16 GB devices, not 32 GB, and the scarce resource is weekly hours. A session gets two Tesla T4s. They do not present as one 32 GB GPU, so a single generate still lives on 16 GB (that is why I2V 1024 OOM’d). A three-minute teaser is ~36 five-second pads at ~7.5 minutes each: a clean pass is ~4.5 hours plus retries, inside ~30 GPU hours/week and a 12-hour kernel.

Kaggle kernel titles. A generated notebook named job.ipynb became the kernel title job, which the API rejects (minimum length). Push now expands short stems to a legal title/slug so a private kernels push actually submits.

Accomplishments that we're proud of

  • A measured T2V path to 1024×576 on Kaggle T4×2, not a theoretical 1080p claim.
  • Ten checksummed example clips in the public repo, including nebula and galaxy, so the library is visible without a live kernel.
  • Word → evidence profile → typed PDX request/result → private push → fetch, with check that never prints the API key.
  • Generated job notebooks gitignored so real prompt sheets cannot be committed by accident.
  • A story we can defend in front of Flow: keep $20 for heroes; generate the unnamed pads here. 36 clips at ~7.5 min each is about 4.5 hours, inside a week of Kaggle GPU quota, with a lot of room to retry.

What we learned

Free GPU time is a capacity constraint, not a quality slider. The useful split is: Flow (or any paid hero model) for the 5-10 shots that must be that landing / that face; this factory for the rest.

“Latest model” is not a product. On a T4, the checkpoint that completes the sheet beats the one that looks better in a single lucky clip and then dies.

Evidence belongs in the document layer. Checksums and private-by-default are part of the UX for anyone who would actually paste client prompts into a notebook.

Optional clouds have to earn their place. Hugging Face ZeroGPU is real, and still the wrong farm for 36 LTX clips.

What's next for B-roll Library Generator

v0.1 is a volume executor: Word sheet in, checksummed 5-second pads out, on a free Kaggle T4×2 session. The product is the factory, not a finished trailer.

What we want next is a libre toolchain: gratis when we can, paid only when the hardware demands it. That is reliable enough for newcomers to enter this field without renting a studio or guessing a notebook.

Still as tools, and only where the hardware actually holds the job:

  • Image-to-video / refine at 1024×576 on a 24 GB class card (rent), behind the same Word → checksum → fetch contracts.
  • 1080p as an offline remaster step, not a T4×2 native mode.
  • Interrupt / resume so a 12-hour kernel keeps finished clips and only reruns the rest.

Hugging Face ZeroGPU stays optional polish (stills, voice), not a second video farm. The factory stays a Word sheet and one Kaggle T4×2 session.

Built With

Share this project:

Updates

Submission history