Inspiration
Organizing Google Developer Group (GDG) events meant drowning in unpaid operational grind: cleaning attendee CSVs without leaking PII, outpainting speaker headshots, resolving date conflicts with Polish holidays, balancing schedules, scanning multi-currency receipts (EUR/USD/PLN), and drafting endless LinkedIn posts.
I asked: What if an autonomous AI crew handled 95% of community ops? That led to Community AI Studio, built on Google Agent Development Kit (ADK) 2.0 and Vertex AI / Gemini.
What It Does
A unified workspace where specialized agents collaborate across the event lifecycle:
- Orchestrator (
root_agent): Manages shared state and delegates tasks. - Video Editor (
video_editor): Outpaints 1:1 speaker photos to 9:16 vertical video with Gemini Image, animates via Veo 3.1 / Gemini Omni Flash, and exports 4K MP4/GIFs. - Receipt Scanner (
receipt_scanner): Multimodal OCR reconciles EUR/USD/PLN against live NBP/Pekao FX rates and builds Google Docs expense reports. - Event Planner (
event_planner): Checks clashes via Nager.Date API and local tech calendars. - Registration Manager (
registration_manager): Deduplicates CSVs with strict PII guards; exports.docxrosters. - Agenda Formatter (
agenda_generator): Auto-snaps talk slots and buffer times. - LinkedIn Writer (
linkedin_post_generator): Drafts Engaging, Concise, and Professional posts with anti-hallucination guardrails. - Venue Secretary (
office_secretary): Drafts security access and room booking requests.
How I Built It
- Runtime: Python 3.12, Google ADK 2.0.
- Model Tiering:
gemini-3.5-flash-lite: Low-latency orchestration, tool calling, and text formatting.gemini-3.7-flash: High-precision multimodal OCR and LLM-as-a-judge evaluation.gemini-3.1-flash-lite-image: 9:16 portrait outpainting.veo-3.1-fast-generate-001&gemini-omni-1.1-flash: Generative speaker animation.
- A2A Protocol & Backend: Distributed microservices with IAM/OIDC auth for heavy video rendering and Google Workspace integrations.
- Frontend: Svelte 5 + Vite SPA streaming agent thoughts via SSE.
Challenges & Solutions
- Multi-Currency Reconciliation: Receipts vary in format and currency. We combine
gemini-3.7-flashstructured outputs with deterministic math: $$\text{Total}{\text{PLN}} = \sum{i} \left( \text{ItemPrice}_i \times \text{FX_Rate}(\text{Currency}_i, \text{Date}) \right)$$ Any mismatch against printed totals triggers an immediate review flag. - Dynamic Kinetic Typography: To prevent text clipping across variable title lengths, we scale duration dynamically before GSAP/FFmpeg rendering: $$\text{Duration} = \max\left(T_{\min}, T_{\text{base}} + \alpha \cdot \frac{\text{CharCount}}{\text{ViewportWidth}}\right)$$
- Zero Context Bloat: Passing large files through prompts causes latency $\mathcal{O}(\text{Context} + \text{Tokens})$. We pass file handles (
session.state) by reference rather than binary payloads.
Accomplishments & Key Learnings
- <60s Pipeline: Raw assets $\rightarrow$ 4K animated teaser & verified Google Docs expense report in under a minute.
- Model Tiering is Critical: Routing routine tasks to
gemini-3.5-flash-litekeeps latency sub-second and costs minimal, saving heavy models for vision/video. - Decoupled Architecture: Offloading video rendering to A2A servers keeps the core orchestrator instant and responsive.
What's Next
- 1-Click Publishing: Direct integrations with LinkedIn, X, and Meetup APIs.
- Live Q&A Sub-Agent: Real-time Sli.do question clustering during live talks.
- Multi-Chapter Analytics: Unified quarterly budget tracking across chapters.
Built With
- a2a-protocol
- cloud-run
- computer-vision
- fastapi
- ffmpeg
- firebase
- gemini
- gemini-flash
- google-adk
- google-cloud
- google-docs-api
- google-drive-api
- gsap
- javascript
- multimodal
- ocr
- python
- rest-api
- server-sent-events
- svelte
- typescript
- veo
- vertex-ai
- vite
Log in or sign up for Devpost to join the conversation.