Inspiration

Creating short-form content from long videos is still a time-consuming process. Creators often spend hours watching footage just to find a few moments worth clipping. ClipAI was inspired by the idea of using AI to understand a video first, identify its best moments, and reduce the manual work required to create short-form content.

What it does

ClipAI analyzes long-form videos using AI and detects moments with strong short-form potential, such as funny, emotional, surprising, action-packed, educational, and reaction moments. It provides timestamps, descriptions, categories, and AI scores so users can quickly decide which moments are worth turning into clips.

Users can also analyze a public YouTube video from its link before uploading the original video, saving unnecessary upload time and bandwidth.

How we built it

ClipAI uses React and TypeScript for the frontend, Supabase for authentication, database, and private storage, and Gemini for multimodal video understanding.

Video processing is separated from AI analysis. A Dockerized FFmpeg worker deployed on Render handles media processing and clip generation, while the application is deployed on Vercel. Signed URLs and HMAC authentication are used to securely communicate between services.

Challenges we ran into

One of the biggest challenges was realizing that understanding a video and actually editing it are completely different problems.

We also encountered challenges with large video uploads, Gemini processing time, Supabase RLS policies, signed storage uploads, authentication, environment variables, FFmpeg worker deployment, and securely connecting multiple services.

Designing the system so that large files didn't need to be uploaded unnecessarily was another major challenge

Accomplishments that we're proud of

We're proud of building an end-to-end pipeline that combines multimodal AI with real video processing rather than stopping at AI-generated recommendations.

ClipAI can analyze videos, identify and rank meaningful moments, display their timestamps, and use a separate FFmpeg processing system to generate actual video clips.

We're especially proud of the analyze-first workflow, where users can analyze a YouTube video before committing to uploading a potentially multi-gigabyte source file.

What we learned

Building ClipAI taught us that creating an AI product involves much more than connecting to an AI API.

We learned about multimodal AI, prompt engineering, structured AI outputs, background jobs, FFmpeg, Docker, Supabase RLS, signed URLs, resumable uploads, HMAC authentication, deployment, error handling, and asynchronous system design.

Most importantly, we learned how different services can be combined into a reliable end-to-end AI application.

What's next for ClipAI

The goal is to evolve ClipAI from Find → Cut into:

Find → Cut → Edit → Publish

Future improvements include automatic captions, 9:16 smart reframing, dynamic zooms, sound effects, improved moment ranking, faster video analysis, more editing controls, and eventually automated publishing to short-form content platforms

Built With

Share this project:

Updates

Submission history