Inspiration

I run several YouTube Shorts channels across very different niches — hair-transformation clips, short drama, molten-glass art and fantasy adventure. Every Short starts as a text prompt for an AI video generator, and for months I did everything by hand: watching reference videos, writing prompts, counting characters, writing titles and re-explaining my rules to the AI every single session.

The breaking point came when I lost credits on four refused generations in one day. The content was fine — the prompts were simply too long, and the generator returned the same generic "I can't generate the requested content" message it uses for policy blocks. That's when I decided my creative rules shouldn't live in my head or in chat history. They should be code that checks every prompt before I spend a single credit.

What it does

YT Automation is a local pipeline that turns a reference video into a checked, ready-to-generate Shorts prompt and title:

  1. Study – A Video Reference Lab analyses a reference video locally: FFmpeg/PyAV metadata and timestamps, OpenCV motion analysis, WhisperX word-aligned dialogue, and YAMNet sound-event detection.
  2. Prepare – Each niche has its own rulebook (CURRENT_RULES.json) and its own guard script. The guard builds a packet from exactly one selected reference, so styles from different sources never bleed into each other.
  3. Draft and review – A prompt is drafted, then reviewed against the evidence and the niche rules.
  4. Check and export – The guard runs hard validators. Only a prompt that passes can be exported, together with its matching title.
  5. Remember – Owner feedback and every failed render become new rules and new regression tests, saved to shared memory so the next session (or my WhatsApp assistant) starts with the latest rules.

How I built it

  • Python guard scripts per niche, each with prepare / review / check / export commands
  • Video Reference Lab: FFmpeg + PyAV, OpenCV, WhisperX, YAMNet — all running locally and for free, no paid APIs and no uploading of source videos
  • JSON rulebooks with versioned owner corrections and preferred/rejected examples
  • Regression tests: every past failure becomes a test case so it can't come back
  • A shared cloud memory that keeps desktop and WhatsApp sessions in sync

Every prompt must pass a set of hard constraints. For a prompt $P$ and title $T$:

$$ |P| \le 2100, \qquad |T| \le 100, \qquad 3 \le #\text{hashtags}(T) \le 4 $$

and its beat timestamps $[t_0, t_1), [t_1, t_2), \dots, [t_{n-1}, t_n]$ must cover the whole runtime without gaps:

$$ t_0 = 0, \qquad t_n = D, \qquad t_{i} < t_{i+1} \;\; \forall i $$

where $D$ is the declared duration (30 s by default, vertical 9:16).

For motion, the Lab scores each frame pair by mean absolute pixel difference,

$$ m_t = \frac{1}{WH}\sum_{x,y} \left| I_t(x,y) - I_{t-1}(x,y) \right| $$

which helps me find where the action and cuts happen in a reference and pace my own beats to match.

Challenges I faced

  • Silent failures. Over-length prompts and policy blocks return the same error. Fix: always check length first — now an automatic gate.
  • Rule drift. Rules from one niche kept leaking into another (drama camera cuts showing up in hair videos). Fix: strict category isolation — each niche has its own rules, references and guard.
  • Pipeline ≠ understanding. A successful analysis run doesn't prove an action really happened in the video. I learned to keep measurements, model predictions and human-reviewed observations separate, and to leave unclear items marked needs_review instead of pretending they were verified.
  • Dead air. Generated videos had long silent endings. Fix: mandatory timestamped beats and an active final 5 seconds with a natural payoff line.
  • Robotic dialogue. Early prompts made characters talk like tutorials ("Watch the comb!"). Now dialogue is checked for natural, non-repeated, situation-appropriate lines.

What I learned

  • Rules belong in code, not in memory. If a correction isn't a test, it will eventually be forgotten.
  • Validate before you spend. A cheap local check beats an expensive failed generation every time.
  • Be honest about what your tools know. Good automation states its blind spots clearly.
  • Free local models (WhisperX, YAMNet, OpenCV) are powerful enough to build a serious analysis pipeline without paid inference.

What's next

  • One shared onboarding flow so a new niche gets its own rules, checks and tests from day one
  • Analysis of finished renders, so generated videos are audited against the prompt automatically
  • Better voice and sound design for each niche using ElevenLabs

Built With

Share this project:

Updates

Submission history