Inspiration

In the modern media landscape, breaking news in technology, artificial intelligence, cryptocurrency and pretty much anything happening in the world moves at lightspeed. Yet, the process of producing high-retention, broadcast-quality short-form video remains painfully manual.

I've always been obsessed with tech content on social media: What new AI model did Google just drop? What powerhouse GPU is Jensen Huang unveiling next? I love watching these stories break in my feed, especially when they are explained through clean, high-retention motion graphics.

Naturally, I wanted to start my own tech news channel. But the moment I tried, the operational reality hit me like a brick wall:

  • I had to turn on every notification, constantly doomscrolling and monitoring RSS feeds to catch breaking stories before they got old.
  • I had to spend hours researching, fact-checking claims, and sorting through clickbait.
  • I had to write scripts, storyboard scenes, manually keyframe motion graphics, synthesize voiceovers, line up audio waveforms, render video files, and post consistently day in and day out.

Every time I tried, life got in the way. It was exhausting, updates slipped past me, and maintaining that level of daily consistency felt impossible as a solo creator.

When I saw this hackathon, something clicked: Why not build a system that solves this exact problem for me—and everyone else—completely autonomously?

That was the birth of COLONY-V. Instead of another shallow chatbot or generic text-generator, I set out to build a true autonomous AI newsroom: an agentic network that monitors what's happening in the niches I care about, digs deep into claims with a self-healing web scraper, writes punchy scripts, designs and renders stunning motion graphics using Remotion and Google Gemini, and publishes them on an automated schedule without me having to lift a finger.


What it does

COLONY-V is an always-on, multi-agent editorial and media production network powered by Google Gemini 3.5 on Vertex AI. It autonomously executes a continuous 7-stage production lifecycle:

  1. Discovery Agent (Monitor): Scans live topic feeds and Google News search streams. It filters duplicate content using cryptographic hashing ($\text{SHA-256}(\text{title} \parallel \text{url})$), deduplicates historical stories, and scores each candidate on three factors — relevance, novelty, and urgency — combining them into a single ranking score: $$\text{Score}_{\text{story}} = \text{Relevance} + \text{Novelty} + \text{Urgency}$$
  2. Research Agent: Scrapes deep article content using a self-healing extractor. It isolates verifiable facts, flags claims with source citations, identifies conflicting narratives across outlets, and packages verified editorial assets.
  3. Research Integrity Check: Every claim the Research Agent extracts is only marked verified if it's confirmed by two or more independent sources — unverified claims travel downstream flagged as such rather than silently treated as fact. If research generation fails outright, the pipeline skips the story rather than risk building a script on an unconfirmed foundation.
  4. Scriptwriter Agent: Transforms structured facts into high-retention video scripts featuring an explosive hook, 3–6 distinct visual beats, CTA, target timestamp pacing, and tailored color branding.
  5. Art Director Agent: Reads the research report and script to choose a semantic accent color and map every beat to a motion graphic template — Stat Reveals, Mechanism Flow Diagrams, Timelines/Roadmaps, Ken Burns Editorial Cards, or Kinetic Typography — before handing the finished visual plan to the Producer.
  6. Visual Producer (Remotion + Edge-TTS): Synthesizes narration per scene, measures audio waveforms via ffprobe, clamps frame allocations ($F_i = \max(1, \text{round}(T_{\text{narration}} \times 30\text{ FPS}))$), dynamically synthesizes procedural sound effects (whooshes, pops, dings), and renders crisp portrait 1080x1920 MP4s via headless Chromium.
  7. Publisher Agent & Analyst Loop: Automatically uploads the finished video to YouTube Shorts with optimized SEO titles, tags, and descriptions. The Analyst Agent reviews performance and feeds learned strategic adjustments back into future Discovery prompts.

All of this is managed through an interactive Glassmorphism Mission Control Dashboard featuring real-time WebSocket log streaming, live DAG agent execution visualizers, background scheduling, and cooperative operator controls (Pause via SIGSTOP, Resume via SIGCONT, and Stop).


How we built it

  • Agent Intelligence & Orchestration: Powered by Google Gemini 3.5 Flash on Google Cloud Vertex AI (using Application Default Credentials for enterprise-grade keyless authentication). Orchestrated through custom Google ADK (Agent Development Kit) tool patterns with strict jsonschema verification contracts between every stage.
  • Motion Graphics Engine: Built on Remotion 4 (React 18, TypeScript, Tailwind CSS), utilizing a custom Dark Tech design system, Lucide icons, and GPU-accelerated Angle GL rendering.
  • Audio & Sound Design: Edge-TTS neural speech synthesis combined with mathematical pure-Python PCM wave synthesis for procedural SFX.
  • Backend & Event Architecture: FastAPI with asynchronous Python 3.11 runtimes, non-blocking executor threadpools, atomic filesystem writes, and Google Cloud Pub/Sub push architecture for distributed stage execution.
  • Persistence & Cloud Infrastructure: Deployed natively on Google Cloud Run with Google Cloud Firestore for state persistence and log buffering.

Challenges we ran into

  1. Making Remotion + Gemini Actually Look Good, Not Just Render: This is the challenge that ate the most of my time by far. It's one thing to get an AI agent to output a JSON visual plan and pipe it into Remotion — it's another thing entirely to get the result to look like a professionally designed motion graphic instead of a generic template. Early attempts produced layouts that were technically correct but visually flat or inconsistent from shot to shot. We solved this by refusing to let the Art Director generate freeform layouts — instead we built a constrained layer library (Paper, Plate, Copy, Data, Callout, Brand) behind a strict VisualPlan JSON schema, so the agent chooses which composition to build from a fixed, pre-designed set rather than inventing one from scratch. That trade of creative freedom for consistency is what finally made the output look intentional.
  2. Sub-Millisecond Audio-to-Visual Synchrony: Early prototypes suffered from visual drift because LLM duration estimates did not match real speech length. We eliminated static scene timing entirely: the engine generates dedicated narration audio per beat, measures true waveform durations using ffprobe, and passes dynamic frame maps directly into Remotion via runtime props.
  3. Headless Chromium Rasterization in Containers: Rendering rich React motion graphics inside resource-constrained Docker containers initially triggered rasterizer crashes due to /dev/shm shared memory limits and unconstrained spring velocities. We re-engineered the rendering pipeline with Angle GL software rendering, clamped spring physics, and stripped heavy SVG filters for rock-solid stability.
  4. Cooperative Control over Native Subprocesses: Pausing or stopping an autonomous AI pipeline while heavy subprocesses (scrapers, audio generation, video renderers) are executing proved difficult. We implemented an idempotent process registry tracking all active PIDs, paired with a monotonic pause clock in pipeline_runtime.py that discounts paused time from execution timeouts.
  5. Factuality & Metric Extraction: LLMs often hallucinate or misformat statistics in rapid narrations. We implemented strict regex tokenizers that differentiate real numerical metrics from calendar years and bind them to animated spring counters.

Accomplishments that we're proud of

  • True Autonomy: Built a system that goes from zero to a published, high-quality 1080x1920 video with synchronized voiceover, animated charts, and verified sources in under 3 minutes without human intervention.
  • Closed-Loop Intelligence: Developed a functioning feedback loop where the Analyst Agent's post-run insights directly influence story discovery parameters on subsequent runs.
  • Enterprise Security: 100% keyless authentication using Google Cloud Vertex AI ADC, removing exposed API keys from the entire stack.
  • Resilient Architecture: Containerized production deployment running on Google Cloud Run with comprehensive error boundaries, self-healing scrapers, and cooperative runtime control.

What we learned

  • Agent Contracts are Mandatory: LLM agents cannot communicate reliably via free-form text. Deterministic schema validation and sanitization layers between every agent handoff are essential for unattended stability.
  • Physics-Driven Visual Timing: In programmatic video synthesis, visual timing must always be downstream of audio physics—audio is the master clock.
  • Cooperative Concurrency: Asynchronous event loops require rigorous isolation from synchronous blocking calls (such as subprocesses and network requests) via threadpool executors to prevent dashboard lockups.

What's next for COLONY-V

  • Omnichannel Publishing: Expanding direct automated delivery beyond YouTube Shorts to TikTok, Instagram Reels, and X/Twitter.
  • Multi-Speaker Conversational Formats: Introducing dual-agent debate and interview video layouts with dynamic speaker transitions.
  • Real-Time Retention Fine-Tuning: Hooking YouTube Analytics retention curve APIs directly into the Scriptwriter's prompt engineering to optimize hook pacing based on second-by-second drop-off data.

Built With

Share this project:

Updates

Submission history