Artifact is a production tool with an AI agent system at its center. Think of it as the art department's fabrication partner — one that reads a scene the way a prop master would, proposes several buildable directions for the object the camera will study, and — once a human approves a direction — compiles the turnaround sheets and build specifications a workshop actually needs.

You give Artifact a screenplay moment (paste it, upload it, or fill in a short brief). It distills a structured hero-prop brief — what the object is, the world and era it belongs to, what it does on screen, and the physical and stunt constraints — then generates divergent concept designs with photoreal images. You review the options in a museum-style "vitrine" view, choose one with a recorded rationale, and Artifact finishes the job: seed-locked multi-angle turnarounds, detail callouts, and a materials/finish/build spec ready for the floor.


The System — A Web App and an Agent System, Built as One

Artifact is two things designed together: a web application and an AI agent system. Neither is useful without the other.

The web app is what the crew sees — a login-gated studio dashboard for onboarding a brief, reviewing concept plates, walking human-in-the-loop review gates, and exporting build dossiers. It is clean, responsive, and browser-based, and it persists the team's production profile and prop slate so work survives between sessions.

The agent system is the intelligence. It extracts the brief, generates concept options, records the human selection, and produces the final asset package — calling Gemini models for reasoning and image generation along the way. The two halves talk over a small internal REST contract: the app sends the brief and requests; the agent returns options, images, and specs, which the app displays and stores.

The Agent System — One Orchestrator, Three Specialists

The agent core is a multi-agent architecture built on the Google Agent Development Kit (ADK). A root agent coordinates three specialists, and a strict state machine enforces the order of work so nothing runs (or spends money) out of turn.

  • OptionsGenerator — turns the brief into several distinct silhouette/design directions and renders draft concept images with the fast Nano Banana 2 image model.
  • SelectionRecorder — captures the human's chosen option, who chose it, and why (the review gate), making the decision auditable for franchise continuity.
  • AssetFinisher — for the selected option only, generates high-fidelity, seed-locked hero imagery with Nano Banana Pro, plus turnaround angles and a structured build/CMF spec.

The same ADK agents run in-process locally (for tests and offline development) or on Vertex AI Agent Engine in the cloud, chosen automatically — the code does not change between the two.


The Problem It Solves

Turning a director's note into something a workshop can build is slow and bespoke. A prop master reads the scene, sketches directions, sources references, waits on concept art, reconciles it against franchise canon, and only then writes a build spec. Each hero prop can absorb days of back-and-forth before fabrication even starts — and hero props are the ones the camera lingers on, so they cannot be rushed or made generic.

Artifact compresses that front half. It reads the scene, proposes buildable directions with real images in minutes, keeps a human firmly in the decision seat, and emits a spec the floor can use — while guarding continuity so sequels stay faithful to what an earlier film established.


How It Works

1. You give it a scene or a brief. Paste screenplay text, upload a file, or fill the form. The agent extracts a structured brief (prop, world, era, on-screen function, constraints).

2. It generates options. The OptionsGenerator produces several divergent concepts with draft images. You see them as vitrine plates with rationales.

3. You approve a direction (the gate). You pick one and record why. Nothing expensive happens before this — hero-grade imagery is only generated after a human commits.

4. It finishes the asset. The AssetFinisher renders seed-locked multi-angle turnarounds and detail callouts and compiles the build/CMF spec, which you can review and export.


What Artifact Does — The Core Capabilities

Brief extraction from a scene. It reads screenplay text and produces a structured, editable brief rather than making you fill in every field from scratch.

Divergent concept generation. It proposes several genuinely different directions — not variations of one — so the team has real choices, each with a photoreal draft image and a stated rationale.

Human-in-the-loop review gates. Approvals are explicit and recorded. The system will not proceed to expensive finishing until a person has chosen and justified a direction.

Build-ready finishing. For the approved option it generates seed-locked turnaround sheets, six-part detail callouts, and a materials/finish/mechanism spec — the artifacts a fabricator actually needs.

Continuity guardrails. Selections and specs are captured so franchise sequels can stay faithful to established canon (silhouette, symbols, dimensions, finish).


How the Technology Fits Together

A request enters the frontend (an nginx-served React SPA on Cloud Run). It calls the backend (Express on Cloud Run), which verifies a short-lived JWT and proxies to the agent (FastAPI + ADK on Cloud Run). The agent runs its subagents on Vertex AI Agent Engine, calls Gemini through the Google Gen AI SDK for reasoning and for Nano Banana image generation, writes images to Cloud Storage (served as IAM-signed URLs), and the backend mirrors state into Firestore, streaming live updates to the browser. Everything scales to zero at idle; credentials live in Secret Manager; containers build via Cloud Build → Artifact Registry. An optional Confluent Cloud event backbone (or a lightweight HTTP webhook) carries prop-lifecycle events for durability and real-time fan-out.


Findings & Learnings

This project reached its final shape by reversing several early assumptions once reality disagreed. The interesting parts are the why and the payoff.

Why Google ADK over a hand-rolled agent loop. A custom orchestration would have meant building our own runners, session handling, and — crucially — our own path to production. Google ADK gave us a standard LlmAgent + subagent model, an in-process InMemoryRunner for offline tests, and a one-line path to deploy the same agents to Vertex AI Agent Engine. The benefit is dev/prod parity: the exact agents we test locally are the ones that run managed in the cloud, selected by a single environment variable, with no rewrite. That is what the 3-tier platform adapter (Stub / LocalADK / AgentEngine) delivers — fast offline tests and managed cloud scaling from one codebase.

Why the pipeline is staged the way it is. Hero imagery is expensive, so the state machine deliberately splits the work: economic draft options first (Nano Banana 2), a human approval gate, then high-fidelity hero finishing (Nano Banana Pro) only for the chosen option. The API also separates "start" from "run": creating a prop returns immediately with a job id, while generation runs in the background and streams progress. The payoff is cost discipline and responsiveness — no money is spent on hero renders before a human commits, and the UI never blocks.

The image API correction. Nano Banana models are Gemini image models, so we call generate_content(response_modalities=["IMAGE"]), not generate_images (which is for Imagen). Getting this right is what makes images render at all.

Model/region reality. gemini-2.0-flash 404'd in our region; we standardized on gemini-2.5-flash and generate images on the global location.

Serverless everywhere for near-zero idle cost. Cloud SQL cannot scale to zero, so it was a monthly floor; we migrated the backend to Firestore behind a thin compatibility shim that left the route handlers unchanged. Combined with min-instances 0 on all three Cloud Run services (and --cpu-boost so the agent's heavy ML imports survive cold start), the whole stack idles at ~$0

Robustness fixes we learned by testing live. Enabling Firestore ignoreUndefinedProperties and merge-writes fixed both an undefined-field crash and a race where two writers clobbered each other's fields.


Who It Is Built For

  • Production Designers shaping the visual language of a film, TV series, or game cinematic.
  • Prop Masters who must turn a script note into buildable hero-prop directions quickly.
  • Fabricators who need a real materials/finish/mechanism spec and turnaround sheets, not mood boards.
  • Franchise/continuity teams who must keep a hero object faithful across sequels.

Technologies used

Frontend: React 19, Vite, TypeScript, TailwindCSS (served by nginx on Cloud Run). Backend: Node.js + Express + TypeScript (auth, proxy, persistence, SSE). Agent core: Python 3.12 + FastAPI + Google ADK, deployed to Vertex AI Agent Engine. Models: Gemini 2.5 Flash (reasoning) and Nano Banana image models (gemini-3.1-flash-image, gemini-3-pro-image) via Vertex AI. Data plane: Firestore (Native, serverless), Google Cloud Storage, Secret Manager, Cloud Build, Artifact Registry; optional Confluent Cloud (Kafka) event backbone.

Data sources

User-supplied screenplay text and uploaded briefs; a bundled set of sample production scenes; generated images in GCS (delivered as IAM V4 signed URLs); prop/profile state in Firestore.

Built With

Share this project:

Updates

Submission history