Paste the following into the project story field. Adjust the first-person wording if needed to match your team.

Inspiration A photograph captures a view, but revisiting a place means remembering more than what fit inside the frame. There’s the path you walked, the details you noticed, and the little things that never made it into a caption.

Memory explores a simple question: what if your photo album gave you a way to step back inside a moment?

We wanted to connect world models to something personal: the places people already care about.

What it does Memory turns five to ten photographs into a personal photo journey. You can arrange the stops, add recollections, and return to your favorite moments.

Three optional AI features let you explore further:

Spatial photos: Modal runs Depth Anything V2 to estimate relative depth, which Three.js uses to give a photograph a subtle sense of perspective as you move. Landscape scenes: Runware generates a wider scene using a source photograph and up to two additional references. Generated images stay separate from the saved source photos. Live walks: Reactor’s LingBot World 2 uses one selected photograph to start an AI-generated world. You can move with WASD, look around with arrow keys, pause, and save a generated frame. The journal is local-first. Saved photo copies and notes live in your browser, with export and import for portability. Cloud actions require an access code and explicit consent to send photos to the selected provider. ow we built it We built the application with Next.js, React, and TypeScript, using IndexedDB for local persistence.

The browser prepares resized photo copies and removes embedded metadata. Server-side API routes validate requests with Zod, normalize provider-bound images with Sharp, and keep provider credentials out of the client.

A Python worker on Modal runs Depth Anything V2 on an L4 GPU. The resulting depth maps drive the Three.js spatial viewer.

Runware’s FLUX.2 Pro integration creates optional landscape scenes. Reactor’s LingBot World 2 SDK handles interactive world generation and streams the result into the browser.

We used Vitest and Playwright to test data contracts, local journal workflows, consent boundaries, and provider failure handling.

Challenges we faced One important challenge was distinguishing a photograph from an AI interpretation. Estimated depth is not a measured 3D reconstruction, and a generated walk cannot guarantee what actually exists beyond the frame. We kept those distinctions visible and preserved source photos separately from generated scenes.

Real-time controls also needed careful handling. Movement commands persist until stopped, so releasing a key, switching tabs, pausing, or ending a session must clear the input.

Cloud processing introduced another challenge: a request can fail after other work has already succeeded. Spatial preparation saves completed photos as it goes, allowing the user to retry unfinished stops without losing progress.

Accomplishments that we’re proud of We connected a personal photo journal to three distinct AI workflows while keeping the journal useful without cloud processing.

We also made preservation part of the experience: source photos remain separate, generated content is labeled, and memories can be exported instead of being tied to one browser forever.

What we learned World models need clear boundaries around what they generate. A convincing scene can feel like a memory even when its details are invented.

We also learned that session management, consent, and recovery deserve as much attention as the generation itself. They determine whether someone can comfortably use the feature with photos that matter to them.

What’s next for Memory We want to improve scene continuity, make longer photo journeys easier to organize, and offer a clearer comparison between the source photograph and its generated interpretation.

The goal is to make revisiting a memory more expressive while keeping the original record intact.

Built With

Share this project:

Updates

Submission history