Inspiration

Organizing Google Developer Group (GDG) events meant drowning in unpaid operational grind: cleaning attendee CSVs without leaking PII, outpainting speaker headshots, resolving date conflicts with Polish holidays, balancing schedules, scanning multi-currency receipts (EUR/USD/PLN), and drafting endless LinkedIn posts.

I asked: What if an autonomous AI crew handled 95% of community ops? That led to Community AI Studio, built on Google Agent Development Kit (ADK) 2.0 and Vertex AI / Gemini.


What It Does

A unified workspace where specialized agents collaborate across the event lifecycle:

  • Orchestrator (root_agent): Manages shared state and delegates tasks.
  • Video Editor (video_editor): Outpaints 1:1 speaker photos to 9:16 vertical video with Gemini Image, animates via Veo 3.1 / Gemini Omni Flash, and exports 4K MP4/GIFs.
  • Receipt Scanner (receipt_scanner): Multimodal OCR reconciles EUR/USD/PLN against live NBP/Pekao FX rates and builds Google Docs expense reports.
  • Event Planner (event_planner): Checks clashes via Nager.Date API and local tech calendars.
  • Registration Manager (registration_manager): Deduplicates CSVs with strict PII guards; exports .docx rosters.
  • Agenda Formatter (agenda_generator): Auto-snaps talk slots and buffer times.
  • LinkedIn Writer (linkedin_post_generator): Drafts Engaging, Concise, and Professional posts with anti-hallucination guardrails.
  • Venue Secretary (office_secretary): Drafts security access and room booking requests.

How I Built It

  • Runtime: Python 3.12, Google ADK 2.0.
  • Model Tiering:
    • gemini-3.5-flash-lite: Low-latency orchestration, tool calling, and text formatting.
    • gemini-3.7-flash: High-precision multimodal OCR and LLM-as-a-judge evaluation.
    • gemini-3.1-flash-lite-image: 9:16 portrait outpainting.
    • veo-3.1-fast-generate-001 & gemini-omni-1.1-flash: Generative speaker animation.
  • A2A Protocol & Backend: Distributed microservices with IAM/OIDC auth for heavy video rendering and Google Workspace integrations.
  • Frontend: Svelte 5 + Vite SPA streaming agent thoughts via SSE.

Challenges & Solutions

  1. Multi-Currency Reconciliation: Receipts vary in format and currency. We combine gemini-3.7-flash structured outputs with deterministic math: $$\text{Total}{\text{PLN}} = \sum{i} \left( \text{ItemPrice}_i \times \text{FX_Rate}(\text{Currency}_i, \text{Date}) \right)$$ Any mismatch against printed totals triggers an immediate review flag.
  2. Dynamic Kinetic Typography: To prevent text clipping across variable title lengths, we scale duration dynamically before GSAP/FFmpeg rendering: $$\text{Duration} = \max\left(T_{\min}, T_{\text{base}} + \alpha \cdot \frac{\text{CharCount}}{\text{ViewportWidth}}\right)$$
  3. Zero Context Bloat: Passing large files through prompts causes latency $\mathcal{O}(\text{Context} + \text{Tokens})$. We pass file handles (session.state) by reference rather than binary payloads.

Accomplishments & Key Learnings

  • <60s Pipeline: Raw assets $\rightarrow$ 4K animated teaser & verified Google Docs expense report in under a minute.
  • Model Tiering is Critical: Routing routine tasks to gemini-3.5-flash-lite keeps latency sub-second and costs minimal, saving heavy models for vision/video.
  • Decoupled Architecture: Offloading video rendering to A2A servers keeps the core orchestrator instant and responsive.

What's Next

  • 1-Click Publishing: Direct integrations with LinkedIn, X, and Meetup APIs.
  • Live Q&A Sub-Agent: Real-time Sli.do question clustering during live talks.
  • Multi-Chapter Analytics: Unified quarterly budget tracking across chapters.

Built With

Share this project:

Updates

Submission history