Inspiration

Creative agents should direct animation through stable, semantic controls instead of guessing where to click. Stagehand began as a rigging experiment, but rigid puppets produced fragile seams and inconsistent results under a deadline. I pivoted to a simpler idea that matches traditional limited animation: distinct drawings held on exact frames.

What it does

Stagehand is a browser-based 2D frame-by-frame animation studio with 27 native WebMCP tools. A person or agent can inspect a project, write a multi-scene storyboard, manage assets, author and duplicate held cels, set an exact playhead, regenerate estimated lip-sync mouth cues, design procedural sound effects, generate voice through a restricted local bridge, validate the project, inspect any rendered frame, export PNG, and render an audio-bearing WebM.

The release includes three original multi-scene demos: Ship the Brick, One More Deploy, and Land Money. Each has character dialogue generated locally with distinct voice designs plus deterministic procedural sound effects. The submission narration uses a consented local OmniVoice clone of my own voice; the reusable reference and voice prompt never enter the public repository.

How I built it

The app uses Next.js, React, TypeScript, Canvas 2D, Web Audio, and MediaRecorder. A shared held-cel evaluator selects the latest authored drawing at or before an integer frame with no interpolation. Preview, storyboard thumbnails, frame inspection, PNG export, and WebM rendering all use that same evaluator.

The WebMCP surface is registered with document.modelContext.registerTool(...). Every mutation accepts an optional expected revision and idempotency key. Invalid targets return structured failures without incrementing state. Read tools omit binary asset payloads, and export tools return bounded metadata.

The local voice bridge binds to loopback, restricts browser origins, caps input size, detects an OmniVoice-compatible endpoint, and preserves an offline preview fallback. WebM export mixes decoded dialogue and optional music with deterministic procedural SFX.

Challenges

The biggest challenge was recognizing that the technically ambitious rigging path was the wrong product constraint. Switching to held cels improved visual stability and made scene direction, lip sync, and rendering much easier to reason about. Other difficult parts were keeping UI and agent mutations on one revisioned state model, mixing browser audio into WebM, timing local speech against scenes, and testing native WebMCP registration rather than only testing JavaScript in isolation.

Accomplishments

  • Exactly 27 public WebMCP tools with concurrency and idempotency guards
  • Three multi-scene animations with dialogue, lip sync, and procedural sound
  • One deterministic evaluator shared across preview, inspection, PNG, and WebM
  • Responsive desktop and 390 px layouts
  • Public MIT-licensed repository with no credentials or private voice material
  • Hosted browser and native WebMCP smoke tests against the production URL

What I learned

The right agent interface is less about exposing every low-level capability and more about choosing stable creative primitives. Held drawings, scenes, captions, cues, and exact frames give both humans and agents a vocabulary they can verify. I also learned to treat generated media as provenance-bearing project data and to keep personal voice material local by design.

What's next

I want to add denser cel packs, more backgrounds and prop variants, automatic phoneme alignment, editable exposure sheets, reusable style bibles, and richer nonlinear scene assembly while preserving the same deterministic tool contract.

Built With

  • canvas-api
  • codex
  • mediarecorder
  • next.js
  • omnivoice
  • react
  • typescript
  • web-audio-api
  • webmcp
Share this project:

Updates

Submission history