Inspiration

I am a mother who treasures the ordinary moments that make a family life. The memories are already in our phones, but a growing collection of short clips is difficult to revisit and turn into something we will actually watch. I built Family Vlog Agent so a parent can choose family videos, tap once, and receive an actual locally rendered Vlog instead of editing suggestions or a list of timestamps.

The problem

Creating a family Vlog normally means reviewing every source, finding useful moments, deciding what belongs together, arranging a story, trimming the footage, and waiting for an export. That is a messy, multi-step chore that competes with everyday family life.

Family Vlog Agent turns that friction into a complete Taskmaster workflow. After source selection, the parent taps Create Vlog once. The foreground, user-started run prepares the media, understands the footage, plans a story, executes the edit, inspects the output, and publishes the MP4 without step-by-step guidance.

What it does

  1. Android Photo Picker returns the selected source videos.
  2. Kotlin inspects each source and adaptively prepares model inputs. Compatible sources keep their original bytes when they fit; otherwise the app first tries keyframe-aligned remuxed segments and creates an H.264/AAC proxy only as a final fallback.
  3. Every prepared input receives a fresh schema-constrained video-understanding call.
  4. Kotlin validates and remaps segment-relative results into authoritative source coordinates, then merges them into video_understanding.json.
  5. One separate, fresh JSON-only story-planning call selects validated event IDs, story roles, and reasons.
  6. Kotlin restores the authoritative source IDs and time ranges into edit_plan.json.
  7. Android Media3 trims and sequences the original media, renders the composition, technically inspects the MP4, and publishes it locally through MediaStore.
  8. After the original Vlog succeeds, the user may optionally start a separate on-device Chinese-English subtitle workflow.

How I built it

The application is written in Kotlin for Android API 36 with Jetpack Compose and coroutines.

ADK Kotlin Android 0.8.0 runs two named, isolated roles: video_understanding_agent and story_planning_agent. Every invocation receives a fresh InMemoryRunner and session, with no tools, memory, sub-agent routing, or shared conversation history.

A custom ADK Model bridge uses Firebase AI Logic to reach the Agent Platform Gemini API. The production model is gemini-3.6-flash in global. Each prepared video input receives its own understanding request. After all results are merged, exactly one separate JSON-only planning request selects event IDs.

The model never receives file-operation authority. Kotlin owns validation and authoritative source/time restoration. Android Media3 owns trimming, sequencing, rendering, output inspection, storage, and final export. The repository contains no project-owned cloud media store or remote renderer.

Public Cloud Run evidence

Cloud Run provides a public, read-only verification surface for the project. The service exposes a frozen aggregate from one real-device run—including key metrics, timing limitations, the immutable evidence commit, and SHA-256 provenance—while accepting no media or application payloads. It is intentionally isolated from the Android Vlog runtime, ingests no live telemetry, exposes no mutating routes, and does not participate in Vlog creation, editing, or rendering.

Public evidence endpoint: https://family-vlog-evidence-513339907677.us-central1.run.app/v1/evidence/latest

Architecture and autonomy boundary

Gemini performs semantic video understanding and structured story planning. ADK provides the isolated agent roles and one-shot orchestration. Kotlin converts model decisions into an executable, validated plan, while Android retains deterministic media authority.

This separation is deliberate: the model decides what the footage contains and which validated events form the story, but it cannot invent paths, open files, or write media. The app resolves every selected event against validated understanding data before Media3 acts.

Real-device result

A post-migration real-device run processed:

  • 4 original clips / 46.8 seconds of source footage
  • 6 uploaded analysis windows, including overlap
  • 11 events identified
  • 5 related events selected by the story planner
  • 20.9-second final local edit
  • 3 min 40.3 s from the likely Create Vlog touch to the original output file mtime
  • 75.5 s total across seven observed Cloud request-to-choice spans
  • 6.2 s from the final plan choice to the original output file mtime

The timing figures use different endpoints and are not additive.

The matching Google Cloud activity group contains seven consecutive gemini-3.6-flash traces: six video-understanding calls followed by one JSON-only story-planning call. The public evidence snapshot includes checksummed Cloud activity records, device JSON artifacts, event-selection comparisons, Android logs, MediaStore metadata, timing derivations, and integrity hashes.

Challenges I ran into

The first challenge was fitting real video into an inline request budget without uploading everything to a project-owned media store or compressing every source by default. The adaptive preparation path preserves original bytes whenever possible, prefers verified remuxing when segmentation is enough, and transcodes only as a final compatibility or size fallback.

The second challenge was making model-generated event times executable across segmented and overlapping windows. Kotlin maps segment-relative coordinates back to authoritative source coordinates and restores the final edit plan deterministically.

The third challenge was separating semantic autonomy from media authority. Gemini makes content decisions, while the Android application owns files, cancellation, rendering, storage, and publication.

What I learned

  • Do not compress video by default; preserve original bytes when they fit.
  • Separate perception from editorial judgment so the planning request remains JSON-only and auditable.
  • Make model output executable through validated references instead of model-invented paths or timestamps.
  • Keep semantic autonomy and deterministic media authority separate.
  • Local rendering is not the same as fully offline processing.

Data and privacy boundary

Video-understanding requests send the selected video's audio-visual bytes, prepared as necessary, plus readable metadata and the understanding task. Story planning receives only the merged understanding JSON and editing brief.

Sources, temporary preparation files, final JSON artifacts, render intermediates, and exported MP4 files remain on the phone's product pipeline. Firebase AI Logic monitoring can retain sampled request and response activity in Cloud Logging; the verified retention at evidence capture was 30 days. This project does not claim zero cloud retention or fully offline processing.

Repository and evidence

Source code and reproduction guide: https://github.com/qiuqiuaiweb3/family-vlog-agent-public

Sanitized real-device Android and Google Cloud evidence: https://github.com/qiuqiuaiweb3/family-vlog-agent-public/tree/submission-v1/execution-evidence

The submitted build is a personal-use Android 16 sideload for one owner-controlled phone, not a public consumer release. Model understanding and edit quality depend on the selected media and still require human review.

Built With

Share this project:

Updates

Submission history