Inspiration

Every creator sitting on a backlog of long-form content knows the drill: a 20-minute video has 3-4 clip-worthy moments buried inside it, but manually finding them, reframing for vertical, adding captions, and writing titles eats an hour per clip. That bottleneck is why most creators never repurpose their best content. I wanted a tool that collapses that hour into a single command.

What it does

Shorts Miner takes a YouTube URL and outputs 3 ready-to-upload Shorts:

  • Downloads the source video (yt-dlp)
  • Detects faces/subjects per frame and dynamically crops to 9:16 so the subject stays centered, not just a dumb center-crop
  • Burns in captions directly onto the video
  • Generates titles, descriptions, and hashtags per clip using the Gemini API
  • Runs through a Streamlit UI, so no CLI knowledge needed - paste a link, get 3 files back

How I built it

  • Streamlit for the interface — fast to iterate on during the hackathon window
  • yt-dlp for reliable video acquisition
  • OpenCV for face detection, driving a subject-aware crop window instead of static center-crop
  • ffmpeg-python for the actual crop/encode/caption-burn pipeline
  • Gemini API for metadata generation (titles/descriptions/tags) conditioned on transcript context per clip

Challenges I ran into

  • Face detection is noisy frame-to-frame — a naive per-frame crop caused visible jitter. Had to smooth the crop window across frames so the reframe feels intentional, not shaky.
  • Balancing caption burn-in timing against actual speech timing without a full ASR pipeline slowdown.
  • Keeping the whole pipeline (download → detect → crop → caption → metadata) fast enough to feel usable inside a Streamlit session rather than a background job.

What I learned

Subject-aware cropping is a much harder problem than it looks — smoothing and confidence-thresholding the detection box mattered more than the detection model itself. Also spent real time on the gap between "technically works" and "actually usable," which is what pushed the Streamlit wrapper.

What's next

  • Auto-selecting the "best" 3 moments from a long video (currently user-picked timestamps) using transcript-based virality heuristics
  • Direct upload integration with YouTube's API for one-click publish
  • Support for multi-speaker framing (currently optimized for single-subject talking-head content)

Built With

  • computer-vision
  • content-creation
  • ffmpeg
  • gemini-api
  • google-ai
  • opencv
  • python
  • streamlit
  • video-processing
  • yt-dlp
Share this project:

Updates

Submission history