🎬 ClipCrop — Deterministic AI Video Re-Framing Engine

100% Offline • Sovereign NLE Editorial Freedom • $0.00 Cloud Cost • BIPA/CUBI Biometric Privacy Safe-Harbor > Built for the AI Content Engine Challenge 2026 — Production Category

ClipCrop Hero Banner


💡 Inspiration

Repurposing long-form podcasts, webinars, and keynote footage into short-form vertical assets (TikTok, YouTube Shorts, Reels) has become an essential distribution requirement. However, contemporary AI video repurposers trap creators in the Cloud Repurposing Trap:

  • Recurring API Tax: $30 to $99/month subscriptions with strict queue throttling.
  • Severe Privacy Liabilities: Uploading proprietary, unreleased video to remote servers risks NDA breaches and statutory exposure under biometric laws like Illinois BIPA and Texas CUBI.
  • Irreversible Flat Exports: Cloud platforms spit out flat, burned-in MP4s. If an automated crop clips the speaker's forehead or misses an action, the asset is ruined with zero recourse for timeline recovery.

We built ClipCrop to ask: What if a creator could execute an autonomous, mathematically deterministic virtual cameraman entirely on consumer CPUs for $0.00, delivering both ready-to-publish social cuts and fully editable timeline tracks for DaVinci Resolve and Premiere Pro?


🔍 What It Does

ClipCrop ingests 16:9 widescreen footage and executes a forward-only, 8-stage perception pipeline:

  1. Ingest & Stream Validation: Analyzes container tracks and stream health via sandboxed ffprobe subprocesses.
  2. Audio Transcription & VAD: Extracts native word timestamps via in-process Faster-Whisper INT8 and maps acoustic boundaries with Silero VAD TorchScript.
  3. Deterministic Candidate Selection: Evaluates pause patterns, RMS audio energy peaks, speaking rate variance, and question markers to rank top-K clip candidates (hard-capped at 10).
  4. Multi-Core Spatial Tracking: Executes MediaPipe BlazeFace face detection across candidate frames in parallel via ProcessPoolExecutor.
  5. Binary Confidence Gating: Evaluates tracking certainty against a 0.65 cutoff. High-confidence segments proceed to render; uncertain segments are logged with an auditable skip reason.
  6. Kinetic Camera Smoothing: Applies an Exponential Moving Average filter ($\alpha = 0.15$) to 2D bounding boxes, generating jitter-free 9:16 pan/zoom trajectories.
  7. FFmpeg Render & Timeline Serialization: Encodes vertical 1080x1920 MP4s with burned-in yellow kinetic captions while simultaneously serializing standard CMX 3600 EDL (HH:MM:SS:FF) and Apple FCPXML timelines.
  8. 1-Click Master Creator Pack: Packages .mp4, .edl, .xml, .json, .srt, .ass, .jpg peak thumbnail, and README_METADATA.txt into an all-in-one ZIP archive.

🖥️ Visual Grounding & Live Engine Showcase

1. Autonomous 8-Stage Forward-Only Perception Pipeline

ClipCrop Dashboard Live Execution

Real-time execution across Faster-Whisper, Silero VAD, and MediaPipe BlazeFace with 0 KB network egress in under 45 seconds.


2. Hormozi-Style Kinetic Captions & Active Word Highlighting

ClipCrop Vertical Video Output with Kinetic Subtitles

Sub-millisecond speech-onset alignment clamping with vibrant yellow active word focus and local Qwen 2.5 viral hook synthesis.


3. 1-Click Complete Creator Bundle & Non-Destructive NLE Assets

ClipCrop Complete Creator Pack Folder

All-in-one ZIP bundle containing 1080x1920 MP4, CMX 3600 EDL, Apple XML, SRT captions, cover thumbnail, and metadata.


4. Deterministic Reliability: 95 / 95 Unit & Integration Tests Passing

ClipCrop Pytest 95 Passed Suite

100% test coverage enforcing air-gap socket blockers, source immutability, and 10-candidate loop circuit breakers.


🌟 The 4 Core Architectural Pillars

1. 🛡️ Air-Gapped Privacy & Biometric Safe-Harbor

ClipCrop operates with 0 KB network egress. MediaPipe BlazeFace extracts only spatial 2D coordinates (origin_x, origin_y, width, height). Facial recognition, 3D facial meshes, voiceprints, and identity profiling are structurally excluded, satisfying safe-harbor compliance under Illinois BIPA and Texas CUBI. Outbound sockets are blocked at the runtime level.

2. ⚡ Deterministic 8-Stage Forward-Only Perception Pipeline

Rather than relying on non-deterministic LLM loops that hallucinate timestamps, ClipCrop executes a forward-only signal graph:

  • Faster-Whisper (base.en INT8): Native word-level acoustic alignment.
  • Silero VAD (TorchScript JIT): Sub-millisecond speech onset and pause boundary clamping.
  • 4-Signal Heuristic Scoring: Combines pause patterns, RMS energy peaks, speaking rate variance, and question density.
  • Multi-Core Tracking & EMA Smoothing: ProcessPoolExecutor fan-out with Exponential Moving Average trajectory smoothing ($\alpha = 0.15$) to eliminate camera pan jitter.

3. 🎬 Editorial Sovereignty & Non-Destructive NLE Timeline Export

Instead of flat, baked-in video deliverables, ClipCrop produces paired deliverables:

  1. Production-ready 1080x1920 9:16 vertical MP4 video with burned-in kinetic captions.
  2. Industry-standard CMX 3600 EDL (.edl) and Apple FCPXML (.xml) timeline files.

Editors can import these files directly into DaVinci Resolve or Adobe Premiere Pro, retaining complete keyframe control to adjust pan velocity, crop framing, or color grading.

4. 🔥 Hormozi-Style Kinetic Captions & Local SLM Intelligence

Subtitles are shifted relative to the candidate clip origin and clamped against acoustic vocalization onset, eliminating blank-screen text pops. SubStation Alpha (.ass) rendering generates Hormozi-style kinetic captions with yellow active-word highlights ({\c&H0000FFFF&}).

Viral metadata (opening hooks, 3 high-CTR titles, 5 tags) is synthesized locally via Ollama Qwen 2.5 (3B/7B) in under 2 seconds, backed by an instant deterministic keyword fallback.


🛠️ How We Built It

  • Runtime & Backend: Python 3.11 LTS (>=3.11,<3.12) with PyTorch CPU wheel mapping. FastAPI async backend with Pydantic V2 (strict=True, frozen=True) schemas.
  • Perception Stack: Standalone Faster-Whisper INT8 inference via CTranslate2, Silero VAD loaded locally via torch.jit.load(map_location="cpu"), and MediaPipe BlazeFace short-range tasks.
  • Multi-Core Worker Architecture: Python ProcessPoolExecutor isolates CPU-heavy MediaPipe frame analysis across child worker processes.
  • Frontend UI: React 19, Vite 6, and Tailwind CSS 4 with a task-first, glassmorphic dark interface.
  • Streaming Protocol: Vercel AI SDK v6 Data Stream Protocol (data: <json>\n\n) delivering Server-Sent Events (SSE) for real-time stage transitions, heartbeats, and state commits.
  • In-Process State Machine: Managed by 4 functional reducers (immutable_after_init, append_only, merge_by_key, last_write_wins). Direct state attribute mutation is blocked at runtime.
  • Zero-Egress Observability: Local OpenTelemetry SDK generating 3-tier hierarchical trace spans saved as newline-delimited JSON (outputs/traces/{session_id}_trace.jsonl).

🚧 Challenges We Ran Into

  1. Windows Multiprocessing with MediaPipe: MediaPipe’s C++ SWIG wrappers cannot be pickled across process boundaries on Windows. We resolved this by passing primitive serializable arguments (str, list[int], dict) into worker functions, instantiating the detector inside each child process.
  2. Audio-Subtitle Timing Drift & Intro Silence: When clips started with pauses or intro music, naive subtitle division caused captions to display prematurely across blank screens. We resolved this by extracting native Whisper word-level timestamps, shifting them relative to segment_start_ms, and clamping them against Silero VAD vocal onset boundaries.
  3. Windows FFmpeg Subtitle Path Escaping: The FFmpeg subtitles='...' video filter treats drive colons (e.g., C:\) as filter separators, causing immediate crashes on Windows. We implemented automatic path sanitization, escaping drive colons as C\:/ and standardizing forward slashes.
  4. Protobuf Conflicts with OpenTelemetry: Standard OTLP exporters require protobuf>=5, which conflicts with MediaPipe’s protobuf<5 constraint. Because our system is strictly air-gapped, we resolved this by eliminating network exporters entirely and writing a custom local JsonFileSpanExporter using core opentelemetry-sdk.
  5. Ollama Timeout Resilience: Cloud models can stall or fail. We wrapped local Ollama calls in a 4-second hard timeout with keep_alive: -1, backed by a silent fallback to deterministic keyword heuristics to prevent pipeline interruptions.

🏆 Accomplishments That We're Proud Of

  • 95 / 95 Passing Tests: 10 automated test suites verifying model loading, reducers, tools, controller sequencing, end-to-end integration, and guardrails.
  • $0.00 Running Cost: Complete end-to-end execution without cloud API tokens, external subscriptions, or GPU servers.
  • Sub-35s CPU Render: A full 60-second widescreen clip is transcribed, tracked, smoothed, and rendered to 9:16 vertical video in under 35 seconds on consumer CPUs.
  • Zero-Dependency CMX 3600 EDL Serialization: Generates SMPTE non-drop frame timecodes (HH:MM:SS:FF) using pure Python standard library formatting.
  • Complete Air-Gap Verification: Mechanical proof via socket monkeypatching that 0 KB of data leaves the machine during processing.

📚 What We Learned

  1. Deterministic Pipelines Outperform LLM Agents for Media: Video reframing requires pixel-level geometric accuracy and millisecond timecode synchronization. Using rule-based heuristics and signal processing eliminates hallucinations and produces repeatable, professional deliverables.
  2. Editorial Control is the True Differentiator: Creators reject AI tools that produce flat, uneditable outputs. Providing CMX 3600 EDL and Apple XML timeline files bridges the gap between autonomous AI speed and professional NLE craft.
  3. Local AI on CPUs is Production-Ready: With INT8 quantization (CTranslate2) and TorchScript JIT runtimes, modern consumer CPUs can execute multi-modal AI pipelines locally without dedicated GPUs.

🚀 What's Next for ClipCrop

  • Multi-Speaker Conversational Layouts: Implementing dynamic split-screen (stacked 9:16) rendering when two speakers converse simultaneously.
  • Dynamic Active-Speaker Switching: Adding predictive gaze tracking to cut smoothly between active speakers during debate-style podcast formats.
  • Custom Kinetic Caption Presets: Exposing customizable ASS typography themes (font styles, colors, text bounce animations) via the React dashboard.
  • Direct DaVinci Resolve Studio API Integration: Building a local Python script bridging ClipCrop outputs directly into active DaVinci Resolve timelines via its native scripting API.

🧪 Quickstart & Instructions for Judges

Judges can clone the repository, run the complete 95-pass test suite, and process a test fixture locally in under 3 minutes:

# 1. Clone & enter repository
git clone [https://github.com/piyushxlabs/clipcrop.git](https://github.com/piyushxlabs/clipcrop.git)
cd clipcrop

# 2. Sync Python virtual environment with CPU PyTorch
uv sync --extra dev --python 3.11

# 3. Pre-cache offline AI perception models into models/
uv run python scripts/download_models.py

# 4. Run the full 95-pass automated test suite
uv run pytest tests/ -v

# 5. Execute live 8-stage pipeline verification on sample media
uv run python scripts/test_live_upload.py tests/fixtures/simple_case.mp4

Built With

  • computer-vision
  • ctranslate2
  • edl
  • fastapi
  • faster-whisper
  • fcpxml
  • ffmpeg
  • local-ai
  • mediapipe
  • offline-ai
  • ollama
  • opentelemetry
  • pydantic
  • pytest
  • python
  • qwen-2.5
  • react
  • silero-vad
  • srt
  • tailwindcss
  • video-processing
  • vite
Share this project:

Updates

Submission history