Inspiration
Almost every building has a 2D floor plan. Almost none have an accurate 3D model. Making one today means LiDAR scanners, photogrammetry rigs, or hours in a 3D modelling tool, and the model is out of date the moment someone moves a door or repaints a wall. Facility managers, safety inspectors and first responders end up relying on flat drawings that no longer match reality.
We asked: what if the phone in your pocket and the floor plan you already have were enough? Walk the building once, take a few photos, and get a 3D model that stays in sync with the real space. That is Spatial Sync.
What it does
1. Floor plan → 3D walls, automatically. Upload a floor plan image. Spatial Sync detects every wall with computer vision, keeps the doorways open, and even works out the scale (pixels per metre) from standard door widths. The result is an instant 3D shell of the building.
2. Live indoor tracking on the plan. Stand at the entrance and tap Calibrate. From then on, the phone's AR tracking follows you through the building and draws your path live on the blueprint. No GPS, no beacons, no extra hardware.
3. Photos pinned in space. Tap Capture and a pin drops on the map exactly where you stood, with a cone showing which way the camera faced. Tap any pin to see its photo.
4. Each photo updates only the walls it can see. This is the core idea. For every photo, Spatial Sync works out exactly which walls, and which part of each wall, were in the camera's view, while ignoring walls hidden behind others. Open-source vision models then read that part of the photo: is the wall really there? What colour is it? Where are the doors, windows, exit signs and fire extinguishers? Only that stretch of those walls is updated. Everything else stays untouched.
5. A live 3D digital twin. On the laptop, a 3D viewer rebuilds changed walls within seconds. They flash so you can see what each photo changed. Doors and windows become real openings, objects are marked, and each photo is projected onto the wall it shows. You can orbit the model or walk through it in first person.
How we built it
Spatial Sync has three parts that talk over HTTP on the same network.
📱 Phone app: Unity 6, C#, AR Foundation + ARCore
- ARCore's visual-inertial SLAM (camera features fused with the motion sensors) gives the phone's position and orientation.
- Calibration stores the pose at the entrance as an origin, without resetting the AR session. Each frame, the movement is rotated into that frame (right/forward of the start) and projected onto the plan:
plan = entrance + scale × rotate(heading, (right, forward)). - Heading comes from the camera's forward direction flattened onto the floor, with a fallback for when the phone points down. This avoids gimbal-lock flips.
- The map, the live path (a custom mesh with miter joins), the pins and the view cones are all drawn in Unity's UI system. Each capture records the plan position, heading, field of view (from the camera's projection matrix), pitch and camera height, and uploads them with the photo. Uploads are queued and retried.
💻 Server: Python, FastAPI, OpenCV, PyTorch
- Wall extraction: thresholding, then morphological filtering to erase text and door arcs, then connected components for straight walls and a Hough transform for diagonal ones. Corners are closed by ray-snapping wall ends. Scale is estimated from the most common doorway gap.
- Visibility: one ray per photo column, spaced like a real pinhole camera:
angle(u) = heading + atan((2u − 1) · tan(FOV/2)). Each ray is cast against every wall segment and the nearest hit wins. This maps photo columns to exact positions along each visible wall. - Vision, fully offline: SegFormer semantic segmentation labels every pixel (wall, door, window, furniture…) and gives the wall colour. YOLOE open-vocabulary detection finds objects named in plain text, with no training. An optional hybrid mode sends only low-confidence walls to Claude's vision API and gets structured JSON back.
- Geometry: photo rows become real heights using camera pitch, vertical FOV and distance. The models decide what is there; geometry decides where.
- State: each update is merged only into the visible stretch of each wall. A worse view can't overwrite a better one, and each stretch keeps its best photo as texture. Walls carry revision numbers, and a job queue with a worker thread keeps the API responsive.
🖥️ 3D viewer: Unity 6, C#
- Polls the server's versioned state and rebuilds only walls whose revision changed.
- Generates wall meshes with real door and window openings.
- Projects photos onto walls by mapping each vertex back into the photo (projective texturing).
We also built a one-click Unity setup wizard that creates and wires both scenes, a keyboard simulator to test without a phone, and a synthetic floor plan plus a ray-casting "fake camera" to verify all the geometry end to end.
Challenges we ran into
- Coordinate systems everywhere. Image pixels (y down), Unity's 3D world, the AR frame, the calibrated frame and photo coordinates all differ. One sign error moves a wall across the building. We put every conversion into small, pure functions and checked them against a synthetic renderer.
- Knowing which walls a photo really sees. A simple angle check kept updating walls hidden behind other walls. Casting one ray per photo column with nearest-hit occlusion solved it.
- Vision models vs. reality. Glass doors were labelled as both doors and windows, and a window reflected in a glass table appeared at floor level. We added rules to merge and reject these cases.
- Real-world setup. Windows install issues, a missing dependency for the detector, a black camera feed until the AR background renderer was enabled, and campus Wi-Fi blocking phone-to-laptop traffic.
Accomplishments that we're proud of
- The full loop works on a real phone: walk, capture, and watch exactly the walls in view update in 3D seconds later.
- No training data and no 3D modelling. A floor plan image and phone photos are the entire input.
- Measured on synthetic data: every wall and doorway recovered, auto-scale within 3%, photo-to-wall positions within about 1%, and 3D projection matching the photo within 1 pixel.
- Runs fully offline with open-source models; the cloud AI is only an optional fallback.
What we learned
- SLAM tracks relative motion very well; the biggest error source is the human, through the heading at calibration.
- Let AI answer "what", let geometry answer "where." Vision models are poor at judging distance, but precise camera geometry turns their 2D answers into accurate 3D positions.
- Clean interfaces paid off. Swapping the AR tracker for a keyboard simulator, or the cloud model for local ones, needed no changes anywhere else.
What's next for Spatial Sync
- Accessibility mode for blind and low-vision users. Combine the semantic 3D map with live on-device segmentation and depth to give spoken, spatial guidance: "door, 2 metres, slightly left", "obstacle ahead", and turn-by-turn directions between rooms.
- Use phone depth sensing (ARCore Depth / LiDAR) to correct wall positions, not only their appearance.
- Tap-to-set start points, multi-floor buildings, and several phones mapping together.
- Export to standard formats (glTF, IFC) for architects and facility-management software.

Log in or sign up for Devpost to join the conversation.