We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Film and TV productions constantly need fictional interactive screens for individual scenes — biometric scanners, medical monitors, spaceship consoles, investigation terminals — that exist for exactly one scene and then never again. Building them normally means pulling in a designer and a developer for something that's often on screen for thirty seconds. We wanted to see if a director could just describe the prop instead, and get something real back — not a mockup, an actual working interactive component.

The harder question was what "agentic" should actually mean here. It would have been easy to build a thin wrapper around a single LLM call that outputs some JSX and call it a day. We wanted the agent to be responsible for the outcome — to check its own work against what was actually asked for, and fix it when it's wrong, without a human in the loop.

What it does

A director describes a prop in plain English — what it looks like, how it behaves, what happens when an actor interacts with it. Propire's agent plans concrete requirements from that description (states, interactions, pass/fail conditions), generates a working React component, renders it live in a sandboxed preview, and then deterministically verifies the rendered result against every stated requirement by actually clicking and inspecting the DOM — not by asking an LLM to judge a screenshot. If a requirement fails, the agent patches its own code and re-verifies automatically, up to two attempts, before reporting back. Directors can then iterate conversationally — "make it feel more 1990s," "add a lockdown state" — and the agent revises the existing prop rather than starting over, with full version history and rollback.

How we built it

Next.js/React/TypeScript on the frontend, with every generated prop rendered as a self-contained component inside a fully sandboxed (srcDoc, no allow-same-origin) iframe. The agent itself is built natively on Google Cloud's Agent Development Kit (@google/adk), running on Gemini, rather than a raw API wrapper — planning, generation, and self-healing are each a distinct LlmAgent call, with the planning step using structured (Zod-validated) output instead of hand-parsed JSON. When a director's request implies real-world grounding — "an authentic 1980s military terminal," "realistic ICU vital ranges" — the agent calls Parallel's Search API at runtime to pull reference material before generating, rather than that integration being a bolted-on afterthought. Verification is deliberately not another LLM call: generated props follow a strict contract (state reflected via a data-prop-state attribute, interactions exposed via data-action attributes), so the app can click and inspect the actual rendered DOM and get a deterministic pass/fail.

Challenges we ran into

  • The Replit track wasn't what we assumed. The original plan was for the agent to deploy each generated prop to a fresh Replit project at runtime. There's no public API for that — Replit Agent is a first-party IDE product, not something a third-party orchestrator can call to spin up hosted apps on demand. We caught this before building around it, and pivoted the whole architecture: props render as sandboxed components inside one app instead of separately deployed micro-sites, and we moved to the Parallel track, whose real-world search integration turned out to genuinely strengthen the product rather than just satisfy a rule.
  • Our own code broke the agent framework. Partway through, prop generation started throwing Context variable not found: 'state'. Root cause: ADK's instruction field runs a Jinja-style {variable} template pass when given as a plain string, and our system prompt's worked JSX example contained literal {state}, {count}, {statusText} expressions — which ADK read as required session-state variables that didn't exist. Fixed by passing the instruction as a function instead of a string, which skips that pass entirely, without touching a word of the actual prompt.
  • Serverless statelessness broke self-healing intermittently. Our project store lived in a plain in-memory map. That's invisible in local dev (one process), but on Vercel, the request that creates a project and the later request that patches it during self-healing can land on different function instances that don't share memory — causing sporadic "unknown project" errors specifically on the self-heal path. Fixed by having the client carry its own plan/code state directly into the patch request instead of depending on server-side memory for correctness.
  • Free-tier quota on the Gemini Developer API is easy to exhaust during active development — added key rotation across multiple keys and standardized on gemini-flash-lite-latest, a self-updating alias, rather than pinning a model version that ages out.

Accomplishments that we're proud of

Getting the verify → self-heal loop to actually work live, not staged — the agent catching a real failed requirement, patching its own generated code, and re-verifying, in front of a camera, with no human intervention. Building on Google's Agent Development Kit natively rather than a wrapper. Making the Parallel integration something the product actually needs on certain requests, not a checkbox. And shipping working props across genuinely different domains — a sci-fi biometric scanner, a medical monitor, a spaceship navigation console, and a full investigation/forensics system — from the same unmodified agent architecture.

What we learned

That verifying agent output deterministically (real DOM checks) beats asking a second LLM call to "look at" the result — it's faster, cheaper, and doesn't have its own failure mode. That it's worth deeply validating a platform's actual capabilities before architecting around an assumption about them. And a very concrete lesson about agent frameworks specifically: instruction text isn't just a prompt, it can be part of a templating system with its own parsing rules — worth reading the source, not just the docs, when something breaks in a way that doesn't match your mental model.

What's next for Propire

More curated prop categories, persistent storage so version history survives across serverless instances rather than just within a session, export options so a generated prop can be dropped directly into a real on-set playback rig, and multi-user sessions so a production designer and a director can iterate on the same prop together in real time.

Built With

Share this project:

Updates

Submission history