Inspiration
Hyphened came out of a much larger attempt to build a living open-world simulation. The path looked roughly like this:
At first, I had a game concept and a lot of questions about simulated worlds. I wanted worlds with history, physical space, autonomous actors, and events that continue without the player watching.
Then I spent a long time trying to define the simulation framework itself. I explored journals, world laws, causal graphs, actor perception, action possibilities, branching, replay, and ways for people or models to participate without directly changing world state. Some of this produced useful ideas. Some of it produced an absurd amount of prose.
The crowd simulation became the first serious reality check. Large populations, navigation, changing regions, and long-running activity forced the framework ideas to deal with actual runtime problems.
Core Time grew out of the time-related part of that work. It became responsible for logical clocks, playback, history, seeking, branching, and replay. This deserved its own framework because the same time problems existed outside open-world simulation.
WebGPU Engine grew from a different set of pressures. The browser needed a real way to run persistent GPU systems. Physics, rendering, crowds, and machine learning all needed to share resources and execute in a known order.
The physics work made those requirements much stricter. A large GPU simulation cannot depend on JavaScript rebuilding84 reby rebuilding its state every frame. That pushed the engine toward persistent GPU state, fixed capacities, runtime commands, and clearer ownership.
Around the same period, I started experimenting with machine learning on the GPU. That included tensor execution, policy learning, and an AI4Anim motion-matching demo. A lot of that code later disappeared, but it proved that learned behavior could run inside the same engine as simulation and rendering.
Then NVIDIA released ARDY. It was an autoregressive diffusion model for human motion, driven by language and spatial constraints. It happened to need almost every system I had already been building.
ARDY gave the work a concrete body. Time became a motion timeline. GPU state became a moving character. Spatial constraints became waypoints. History became the previous motion needed to generate the next motion. Rendering and physics could finally operate on the same actor.
Hyphened came from turning that technical work into something I could actually use. It is a focused motion-production environment inside the larger open-world effort.
WebMCP changed the project again. An agent could finally enter the same scene I was watching, use its real authoring operations, inspect, I could inspect the timeline, change something, play the result, and look at what it made.
That last step is probably the clearest expression of what I was reaching for in the early simulation documents. The original idea involved model-driven actors with bounded access to a living world. Hyphened makes a small but real version of that possible.
What it does
Hyphened is an experimental 3D neural motion studio that runs in the browser.
It uses NVIDIA's ARDY model to generate human motion through autoregressive diffusion. I can arrange prompt-conditioned motion on a timeline, place waypoints, direct multiple actors, cut between cameras, add physical objects, play the scene, and capture the result.
The model is about 380 MB. It downloads into the browser and runs locally through WebGPU. The GPU handles model execution, motion reconstruction, character skinning, physics, and rendering.
Hyphened also exposes its authoring operations through WebMCP. An agent can inspect the current scene, read the timeline, add actors, change their motion, edit waypoints and cameras, control playback, and capture visual evidence.
The person and the agent work on the same scene. That is the part I care about most.
How I built it
The first browser implementation focused on matching the released ARDY model. I worked through checkpoint loading, transformer execution, diffusion sampling, tokenizer decoding, motion reconstruction, skeleton transforms, and skinning.
The early version loaded the real model before it could produce useful live motion. Click-to-move was the first meaningful interaction. A click on the floor became a root-position constraint at a future frame, and ARDY generated the movement needed to reach it.
That exposed the hardest part of the model. ARDY generates motion in short windows. Each new window depends on previous motion, prompt conditioning, random state, spatial constraints, and coordinate recentering. The new result then has to join the existing motion at exactly the right frame.
I rebuilt the runtime several times while trying to make that work continuously.
One version used a large application scheduler. Another used workers and service layers. One used a state machine with more than a thousand lines of control logic. Later versions tried shared generation queues for multiple actors.
Most of that machinery is gone.
Each failed design exposed a real rule. Slow generation could not control playback speed. Each actor needed independent history. Generated motion needed to remain one continuous sequence. JavaScript could issue commands, but it could not remain the owner of the generated result.
The current design gives Core Time responsibility for authored time, playback, seeking, and history. WebGPU Engine owns GPU execution, resources, physics, and rendering. The learned-motion code generates retained motion on the GPU. Hyphened composes the scene around those systems.
I directed the project while several generations of Codex agents wrote most of the code. That allowed the project to move very quickly. It also meant that entire architectures could appear before anyone had properly reviewed them. Deleting bad machinery became as important as building new capabilities.
Challenges I ran into
Motion failures are extremely visible. A small error in history or timing can make an actor jump, freeze, drift, disappear, or twist into a broken pose.
Performance has been difficult for the same reason. Model inference and rendering share one GPU. Earlier runtimes performed too much model work near the display loop, so a slow diffusion step could become a visible pause or frame jump.
The browser adds more limits. The model takes time to download and load onto the GPU. WebGPU support still varies. The original ARDY text encoder is much larger than the motion model, so Hyphened currently relies on a library of prepared prompt embeddings and a separate encoder for new prompts.
Agent authoring introduced another set of problems. A tool call can succeed while the visible result is wrong. The agent needs temporal readouts, useful errors, stable scene identity, and images that let it judge its own work.
The project history is also messy. There are hundreds of ARDY-related commits, several large rewrites, and some changes made under severe time pressure. I am still reviewing and simplifying that work.
Accomplishments that I'm proud of
The real ARDY checkpoint runs in the browser through WebGPU. It does not depend on a hosted motion-generation API.
The project grew from one actor performing one prepared prompt into a timeline-based environment with multiple actors, motion prompts, waypoints, cameras, physics, persistent scenes, and agent authoring.
I am proud that Core Time, WebGPU Engine, physics, and learned motion now participate in one visible product. These systems came from different parts of the open-world project, and Hyphened gives them a practical reason to work together.
I am also proud of the WebMCP loop. An agent can work with the actual scene, move through its timeline, make changes, and capture what happened. That was an idea in the earliest simulation-framework documents. It now exists in a form I can use.
What I learned
Architecture shows up on screen.
Bad time ownership becomes a frame jump. Bad history ownership breaks a motion transition. Bad GPU ownership stalls the interface. Bad authoring boundaries let an agent report success while the scene is visibly wrong.
I also learned that the browser can support a serious combination of machine learning, simulation, physics, and rendering. Coordinating them is much harder than getting each one to run alone.
The early simulation work makes more sense to me now. An agent needs access to a world's real objects, time, available actions, and observations. Giving it a chat box outside the world does not create that relationship.
What's next for Hyphened
Hyphened still needs faster scene loading, a clearer workspace, better characters, more motion, and stronger visual review.
The learned-motion code also needs to finish becoming a reusable WebGPU Engine capability. Hyphened is where I am proving and refining it.
The broader project remains open-world simulation. I want authored scenes to happen inside worlds that continue running around them. Physics, background actors, and unrelated events should remain active while a directed sequence plays.
Hyphened is one focused project within that larger effort. It gives me a practical place to develop learned actors and human-agent authoring, but it is far from the last thing I want to build.
Built With
- arktype
- effectts
- react
- typegpu
- typescript
- webgpu
Log in or sign up for Devpost to join the conversation.