Inspiration

Bridging the gap between 2D concepts and immersive 360° reality.

Almost nobody can read a floor plan. Architects can. Contractors can. But the people who will actually use a building, the students in the school, the families in the clinic waiting room, the members of a community centre, look at a blueprint and see lines on paper.

That matters, because those are exactly the people asked to approve things. A school shows parents a renovation plan. A neighbourhood association votes on a redesigned pavilion. A small business owner signs off on a fit-out they can't picture. Everyone nods, and then the space gets built, and only then does anyone realise the corridor is too tight, the entrance has a step, the room feels nothing like it looked on paper.

The same gap exists at neighbourhood scale. When a city publishes a ten-year development plan, residents get charts and zoning maps. Almost nobody can convert that into "what will my street actually look like."

The technology to close this gap exists. Research on immersive technologies in construction shows that letting people experience a project's scale, lighting and spatial "feel" before construction significantly improves approval rates and reduces design misunderstandings. But producing an immersive walkthrough traditionally means a specialist, 3D modelling software and a multi-week rendering cycle, which prices out every school, nonprofit and small business that needs it most.

We wanted to compress that cycle from weeks to minutes, and drop the skill requirement from "trained 3D artist" to "can upload a photo."

What it does

XRPlot is an AI-powered spatial computing platform for creating immersive 360° XR worlds out of interconnected nodes. Each node is a distinct environment: a room, hall, street, office or any spatial zone. Connecting them builds a complex multi-zone virtual world. Create, connect, generate, swap, edit and build. The graph is the floor plan.

Node-Based World Orchestrator. A visual, n8n-style canvas where you lay out environmental zones and wire them together. Drawing an edge from Hall to Room means you can walk from the hall into the room.

Spatial 360° Synthesis. Upload images, sketches, plans or reference photos to any node, including ordinary phone photos of a real room shot up, down, left, right and centre. A vision model first checks your coverage and names the direction you're missing, then a cascade of image models renders the inputs into a single high-fidelity 4096×2048 equirectangular 360° environment.

Generative Node Refinement. Modify any zone with natural language or a reference image. Turn a "Sketch" node into a "Modern Minimalist" interior by prompting it. The AI updates that zone while maintaining spatial consistency, and the original is preserved so you can always revert.

Context-Aware World Assembly. The system understands transitions between zones. Connecting a "Hall" node to a "Room" node generates the transition surface between them, so spatial logic, lighting continuity and boundaries blend rather than cut.

Rapid Prototyping Workflows. Zones support one-to-one, one-to-many and many-to-many relationships with bidirectional connections, plus direct, adaptive and recursive swaps. Swap a "Waiting Room" for a "Reception" and the surrounding connections re-map around it, so you can iterate on the arrangement of a building without rebuilding its contents. Every swap runs through a transactional manager with snapshot-and-rollback and a validator that rejects dangling edges and duplicate IDs.

Predictive Spatial Futures. Drop a pin anywhere on a map and XRPlot projects the next decade of development within a 1 km radius, then lets you walk through it. A simulation model produces a ten-year built-up-index (NDBI) trend and urban density estimate for the site. Gemini synthesizes a comparative past-decade-versus-next-decade report with projected counts for hospitals, schools, shopping centres, parks, road kilometres and population. Those become categorised zones (Main Commercial Street, Residential District, Healthcare Zone, Education Campus, Green Space, Tech Hub), and each one is rendered as a navigable 360° panorama and wired into a world. The output is not a chart about your neighbourhood in ten years. It is your neighbourhood in ten years, and you can walk down the street.

Integrated World Previewer. A unified rendering pipeline aggregates every node and connection into one explorable walkthrough. You drop into the first zone in first person, drag to look, scroll to zoom. When a zone opens onto several others you get a choice of exits with thumbnails, plus a minimap showing where you are in the graph.

AI Chat and Voice Interfaces. Complete platform control through natural language. Create zones, connect them, rename them, restructure a world, all by typing or speaking. A user who can't comfortably use a mouse can still build. A stakeholder reviewing a design can just ask where the accessible entrance is.

Beyond the original scope we added zones that are generated video, for showing a proposal that doesn't exist yet; zones that are genuinely playable 3D scenes with collision and physics; and a 15-language interface covering English, Hindi, Tamil, Telugu, Bengali, Marathi, Spanish, French, German, Portuguese, Arabic, Chinese, Japanese, Korean and Russian, because a consultation that only happens in English isn't a consultation.

Key use cases

Architecture and interior design · real estate and property visualization · construction and urban planning · smart-city development · XR and metaverse experiences · game and virtual-world development · industrial simulation · education and training · film and media production · virtual tourism and immersive experiences.

How we built it

Next.js 15 App Router with React 19, MongoDB via Mongoose, Clerk for secure multi-tenant auth, Cloudinary for media, deployed on Vercel. Roughly 16,000 lines of JavaScript across 82 source files.

The canvas is React Flow with one custom node type and one custom edge type. Child nodes communicate upward by dispatching a custom window event rather than threading callbacks through props, and that decision is the only reason a 1,900-line canvas component stayed readable. Saves are debounced rather than reactive. An early version re-saved on every render and hammered the database.

The previewer is hand-written three.js, not a library. A SphereGeometry scaled by -1 on the X axis so we render the inside of the sphere, the panorama applied as an equirectangular texture, camera at the origin, and hand-rolled drag-to-look with latitude clamped to ±85° so the camera can't roll over the poles. Navigation hotspots are projected from 3D world space into screen coordinates every frame and written imperatively to DOM elements, because routing that through React state at 60fps was visibly janky.

Every generative path is a cascade, never a single call. 360° synthesis tries three image models through OpenRouter, then two more directly through Google, and if every model fails it falls back to a deterministic sharp composite. Playable-3D generation tries a reasoning model on NVIDIA's endpoint, then a list of Qwen models, then Gemini, then a built-in scene. This was our most important architectural decision: a tool that degrades is worth more than a tool that errors.

The prediction pipeline chains four stages that each have to survive the previous one being imperfect. Simulate the NDBI trend and density for the coordinate. Have Gemini turn that into a structured comparative report. Normalise the returned hotspots so the six required zone types always exist even if the model omits one. Then synthesize a 360° panorama per zone and assemble them into a connected world. The normalisation step exists because LLM JSON is never reliably complete, and a missing "Healthcare Zone" would otherwise produce a world with a hole in it.

Spatial Q&A is graph-first. Every zone gets a short AI-written description stored on the node. When you ask a question we rank zones by keyword overlap and pull the top matches plus their directly connected neighbours, using the world's own edges as the retrieval expansion. That gave us spatially-aware answers without a vector database.

The multilingual layer walks the live DOM with a TreeWalker, batches visible text 12 at a time with 4 batches in flight, and writes translations back in place. Three WeakMaps track each node's original English, the last value we wrote, and its current language, which is what stops the layer from re-translating its own output. A debounced MutationObserver catches new content and disconnects itself while we write. If the translation service fails it returns the original English with a degraded flag rather than blanking the interface.

Challenges we ran into

Equirectangular projection is unforgiving. Naively compositing five photos side by side looks fine flat and tears catastrophically the moment it's mapped to a sphere. Seams split, poles smear. We wrote a nine-rule prompt describing exactly what the projection requires, and we validate every returned image with sharp to reject blank or single-colour results, because failed generations came back as a convincing flat grey rectangle rather than an error.

Recursive swaps were the hardest logic in the project. Re-mapping edges when a zone is replaced is easy in one direction and subtle in both. Our first version reassigned A to B and then immediately B to A, quietly undoing itself. Snapshot-and-rollback plus a graph validator is what made swaps safe to ship.

Getting images back out of chat APIs. Image models return pictures inside text responses in at least three different shapes: a markdown image tag, a bare data URI, or unlabelled base64. Our extractor has to try all three before it can even tell whether a generation succeeded.

The translation layer fought back. Version one translated its own output on the next pass, so text drifted further from the original with every DOM change. Version two hit an infinite loop because the MutationObserver fired on our own writes. The WeakMap of last-written values fixed the first, and disconnecting the observer during writes fixed the second.

Real-time voice has no comfortable abstraction. We tried a standalone WebSocket proxy, a Next.js route handler, and a direct browser connection. The route handler was a dead end, because the App Router can't perform WebSocket upgrades. We ended up with raw AudioContext at 16kHz, manual PCM capture and a playback queue, which is far more low-level code than we expected to write.

Accomplishments that we're proud of

  • The 360° previewer is real three.js written from scratch, hotspot projection maths included. We understand every line of it.
  • Nothing in the pipeline hard-fails. Every AI path has a fallback and the final fallback is always deterministic.
  • The prediction pipeline turns a map pin into a walkable world through four chained AI stages and still produces a complete, connected result every time.
  • Direct, adaptive and recursive swaps work transactionally, with rollback. This was the feature we were least sure we could finish.
  • Three completely different interfaces (canvas, chat and live voice) drive one shared tool schema, so a capability added once works in all three.
  • It runs. Photograph a room, and walk through it a few minutes later.

What we learned

  • Graceful degradation is a feature, not error handling. Committing to a fallback for every AI call early on shaped the entire architecture, and it is the reason the project is demonstrable at all.
  • Never trust the shape of a model's output. Every stage that consumes LLM output needed a normaliser, because "mostly correct JSON" is the default and a missing field silently breaks the stage after it.
  • Retrieval doesn't require embeddings. The graph already encoded which zones relate to each other, and walking its edges beat generic vector similarity at a fraction of the complexity.
  • React's rendering model isn't built for 60fps. Knowing when to step outside it and write to the DOM directly was our biggest performance lesson.
  • Transactions belong anywhere state can be edited from two places at once. Once the AI could restructure a world while a user was dragging nodes, "the current state" stopped being obvious.

Impact and benefits

Benefits

  • Cost-effective visualization. High-fidelity, real-time site previews at a fraction of the cost of traditional architectural rendering services.
  • Drastically reduced design timelines. Compresses the multi-week 360° rendering cycle into minutes of AI-driven spatial generation.
  • Democratizes 360° creation. Empowers non-technical users to build immersive, professional-grade XR worlds without complex modelling software.

Impacts

  • Accelerates project buy-in. Replaces static 2D blueprints with fully navigable walkthroughs, significantly increasing approval rates.
  • Eliminates communication gaps. Bridges the mental disconnect between a 2D sketch and the finished build, so stakeholders genuinely understand a project's "vibe" before money is spent.
  • Rapid iterative feedback. Instant design changes and style updates to specific zones via a simple text prompt, so feedback from a community meeting can be reflected before the meeting ends.

Who this reaches locally: a school putting a proposed library layout in front of parents; a community centre showing members a renovation before the vote; a small business planning a fit-out it can't afford to get wrong; a nonprofit showing donors the space their money will build; residents at a planning consultation who can finally see what a decade of development does to their own street. All of them make spatial decisions today using drawings most of the people affected cannot read.

Supporting facts

  • Generative AI accelerates architectural iteration. Integrating generative AI into architectural workflows allows rapid exploration of design variations, reducing schematic visualization from weeks to hours and letting designers focus on creative decision-making rather than manual rendering. Source: Adobe, Generative AI in Architecture
  • Immersive visualization improves stakeholder buy-in. VR and immersive 3D walkthroughs significantly improve project approval rates in construction and real estate by letting clients experience scale, lighting and spatial feel before construction begins, minimizing design misunderstandings. Source: Ubicuity, The Impact of Immersive Technologies in Construction

How we used AI

AI is both the product and part of how we built it. Disclosed in full.

As product features: Google Gemini (image generation and editing for 360° synthesis, vision for photo-direction classification, urban report synthesis, and the Live API for real-time voice), StepFun and Qwen via NVIDIA NIM and OpenRouter (reasoning, 3D scene code generation, spatial Q&A, translation), Seedance and Google Veo (video zones), Anam AI (spoken video guide), and Firecrawl (web-grounded research).

On the prediction feature specifically: the ten-year built-up-index trend and density figures come from our own simulation model, not from live satellite imagery. The projections are a modelled scenario for visualization and discussion, not a forecast, and the interface presents them that way.

As a development tool: we used AI coding assistants while building, for scaffolding, debugging and refactoring. Every architectural decision described above (the node-graph model, the cascade-with-fallback pattern, the four-stage prediction chain with hotspot normalisation, the transactional swap manager, the WeakMap approach to translation state, and imperative hotspot rendering) was ours, and we can explain and defend any part of this codebase.

Credits

Next.js, React, three.js, React Flow (@xyflow/react), Mongoose, sharp, Clerk, Cloudinary and Google Maps Platform, plus the AI services listed above. Equirectangular projection follows the standard approach documented in the three.js examples. NDBI (Normalized Difference Built-up Index) is a standard remote-sensing measure of built-up land. Supporting research cited above.

What's next for XRPlot

  • Replace the simulated NDBI trend with real Google Earth Engine satellite time-series, so predictions are grounded in observed change rather than a model.
  • Right-to-left layout so Arabic is properly usable, not merely translated.
  • Translating placeholder, title and aria-label attributes so screen-reader users get a translated interface too.
  • Guided mobile capture, with on-screen prompts walking a first-timer through photographing a room correctly, since coverage is where most worlds fail.
  • Public share links, so a world can be sent to a community without requiring an account.
  • Building a world with one real local organisation, to find out what actually breaks in use.

Built With

Share this project:

Updates

Submission history