Inspiration

Everyone has had the good version of learning exactly once. You were standing somewhere with a person who knew the place, asking dumb questions, getting real answers. It is the fastest learning there is and almost nobody gets access to it. The rest of us get a textbook chapter, a Wikipedia rabbit hole at 1am, or a museum placard written for the average visitor, who does not exist.

Two things were missing. You cannot be in the place, and you cannot ask. So the material is fixed in advance, and the questions you actually have are never the ones the chapter answers.

Both of those constraints just broke. Single-image 3D reconstruction can build a place you can walk through. Realtime voice models can talk to you about what you are looking at while you look at it. So we put a historian inside a photograph.

What it does

Walk the Past turns any historical source into a world you can walk through with an expert you can talk to. No syllabus, no tour route. You go where you want and ask whatever you want.

Bring anything. A photograph, a video, a written memory, or one sentence describing a place. Five minutes later it is a navigable 3D world. WASD, mouselook, gravity, collision. You walk it, you do not orbit it.

Hold space and ask. The voice historian answers out loud in a sentence or two, like someone walking beside you. It knows which world you are in, what built it, and what is under your crosshair right now. Who built this? What was this street used for? Why is the stone that colour? What happened here a century later? There is no list of approved topics. Ask anything, then follow the answer wherever it goes.

Click anything it says. As the historian speaks, the captions light up every person, place, site and period it names. Click one and the world freezes mid-stride. A card slides in with a summary and the reference article, and for anywhere real, a globe spins up and flies you down to it. Hit back and you are standing exactly where you left off. Tangents cost one click. That is the whole point, because curiosity that is expensive to follow does not get followed.

It admits what it made up. A reconstruction has to invent things. A photograph only shows one side of a street. So we classify every Gaussian in the world against the source image as source-visible, occluded-inferred, or unsupported. Press E and the world recolours by how much evidence is actually behind it. The historian will say it out loud too: that facade is in your photograph, the courtyard behind it is a guess. A tutor that never says "I'm not sure" is not a tutor.

How we built it

  • Worlds. World Labs Marble turns a source image into a Gaussian splat (.spz). Text-only prompts get a written world guide first, then a painted photograph from it, then Marble.
  • Renderer. Three.js and Spark, no react-three-fiber, so React never touches the per-frame path. Custom dyno shaders classify provenance on the GPU. A GLB collider gives walls, gravity, stairs and slopes.
  • Voice. OpenAI Realtime (gpt-realtime) over WebRTC. Ephemeral client secrets are minted server-side, so the key never reaches the browser. The historian has typed tools: getCurrentWorld, getCurrentEvidenceState, getNearbyPOI, getHistoricalContext, and linkHistoricalEntity, which is what turns spoken names into clickable cards and maps.
  • Narration. gpt-4o-mini-tts for the voice, whisper-1 word timestamps against that same audio, so captions track the real waveform instead of a words-per-minute guess.
  • Maps. MapLibre GL with a globe projection for the knowledge portal.
  • Backend. A Cloudflare Worker serves every /api route, R2 holds world assets, and generation runs as a Workflow. A five-minute Marble run with polling outlives any request, and a failed stage should retry from that stage, not from a paid re-generation. Clerk gates everything that costs money.

Challenges we ran into

  • Getting the model to behave like a guide, not a disclaimer machine. Early versions opened every answer by apologising for the source image. Nobody learns from that.
  • Stopping it from inventing. Tools return explicit not_available, so it says "that isn't mapped yet" instead of conjuring a monument.
  • Captions drifting. Timing them to estimated speech rate looked fine for one sentence and was wrong by the fourth. Now we transcribe our own generated audio for real word timestamps.
  • Provenance versus level-of-detail. Classification indexes splats by file order and Spark's LoD tree reorders them, so we ship two meshes and pay in GPU memory instead of framerate.
  • Five-minute jobs on serverless, with a 1 MiB cap on step returns. Splats stream straight into R2 inside their step and photographs pass by key.

Accomplishments that we're proud of

  • A real, open-ended conversation inside a place that did not exist an hour ago.
  • It tells you which parts of the view it made up.
  • Every name it speaks is one click from an article and a map, without leaving the world.
  • Photograph to walkable world in five minutes, in a browser.

What we learned

  • The bottleneck in a learning tool is not how much the model knows. It is how cheap it is to follow a tangent.
  • Voice kills the typing. Walking kills the search query. Clickable captions kill the alt-tab. That is the actual product.
  • Honesty is a feature. People trust a guide more once it has told them what it is unsure about.

What's next for Walk the Past

  • Spatially mapped points of interest, so "what's that building" resolves to a named monument with claim-level citations.
  • Multiple source photographs per world, so evidence coverage grows as you feed it more.
  • Light structure over the open session. Thread tracking, recall questions, a record of what you covered, without putting any of it on rails.

Built With

Share this project:

Updates

Submission history