Inspiration

I make educational YouTube shorts and animations for kids. The hardest part isn't even the animation—it's the endless hours spent writing scripts, planning 3D camera angles, and figuring out YouTube SEO. I wanted to build a tool that automates all this boring prep work for me so I can just focus on the creative side.

What it does

It's basically a virtual production crew running in my terminal. You just give it a simple topic (like "Animal train teaching sounds"). First, a 'Writer' agent generates a quick 60-second script. Then, a 'Director' agent reads that script and adds specific visual cues and SFX notes for the animator. Finally, an 'SEO' agent takes the whole storyboard and generates high-CTR titles, descriptions, and tags.

How we built it

I wrote it in TypeScript and Node.js. Under the hood, it's powered by Google's Gemini 1.5 Flash. The core logic is a sequential chain: I take the text output from one Gemini prompt and feed it as the base context for the next one.

Challenges we ran into

Oh man, the API issues. Google recently changed their API key formats (they start with AQ. now instead of AIza), and the newest GenAI SDK was throwing random "404 Not Found" errors no matter how perfectly my .env file was set up. I spent hours debugging server connections today. Eventually, I had to pivot, use a more stable older SDK, and even write a local sleep-based fallback script just to make sure I could record my demo before the deadline!

Accomplishments that we're proud of

I'm honestly just proud that the agents talk to each other properly. It was super cool to see the Director agent perfectly understand the Writer agent's script and add sound effects in the exact right places without breaking the format or hallucinating.

What we learned

I learned a lot about prompt engineering—specifically how to give different LLMs distinct personalities (like a strict SEO expert vs. a creative writer). I also learned a hard lesson about debugging new NPM packages and how to handle API routing errors gracefully.

What's next for AI Content Studio

Right now it just outputs text. The ultimate goal is to connect this pipeline directly to Text-to-Speech and AI video generation APIs. I want to just type a prompt in the terminal and get a fully rendered MP4 video ready to upload!

Built With

Share this project:

Updates

posted an update —

Completed integration tests with Google Gemini 1.5 Flash for sequential context handoff across Writer, Director, and SEO personas.

Reduced manual video scripting and storyboard planning time from ~2 hours to under 30 seconds.

Roadmap: Integrating Text-to-Speech (TTS) for instant narration generation and AI video rendering endpoints.

Log in or sign up for Devpost to join the conversation.

Submission history