Scale
Inspiration
We built Scale because we wanted to put any real object into any real room at its true size. When we looked at what already ships, we kept running into the same wall. IKEA Kreativ already does LiDAR room scan, accurate dimensions, and place-at-scale, and it does them well. Houzz and Roomvo do the same thing over their own partners. Measurement accuracy is not really an open problem anymore.
What none of those tools can do is handle an object that is not in their catalog. They are retailer tools, and they can get all the data they need to make it possible because the catalog is the product. They cannot place the desk from your parents' basement or a couch you found on Marketplace, on the floor in your new apartment.
So we set out to build the opposite: a system where the catalog is an input rather than the product. That gave us two goals. First, ingest an object that is for sale nowhere and place it at its measured size. Second, make your own possessions searchable, so that a request like "find something that fits the 80 cm gap beside my desk and matches its wood tone" is a query the system can actually answer. We also wanted to show that a Cloudflare edge architecture could carry the whole pipeline, from capture through generation to retrieval, without a traditional server anywhere in the middle.
What it does
Scale scans your room and any physical object, then composes them together at true measured scale. You can view the result on the phone in AR, or stand inside it in a Meta Quest at 1:1.
It does five things:
- Measures a room and an object. An Expo app with custom Swift modules wraps Apple's RoomPlan for the room, and uses a LiDAR raycast-and-grow pass for objects. One tap returns real width, depth, height, and rotation in under a second, with no machine learning involved. A second path runs Apple Object Capture for a full textured mesh when you want one.
- Finds real items/furniture on the open web. We pull live listings from verified Shopify storefronts, extract dimensions out of messy per-merchant data and validate results.
- Generates 3D models from a single product photo from the Shopify storefronts. A product image goes to Stable Fast 3D on Baseten, and the generated mesh is then bound to the measured dimensions. This is the core technical idea: the generative model supplies shape, the depth sensor supplies size, and we bind them exactly once.
- Searches by fit and by style at the same time. Fit is an integer-millimetre range filter, so an object that does not fit is never returned. Style ranks only what survives that filter, using image embeddings and perceptual colour distance.
- Arranges the room on request. An agent turns a spoken request into an objective and a set of constraints, and an OR-Tools solver decides where things actually go. The language model never emits a coordinate.
Two things make it feel different to use. The first is that the blob of an object appears in the room within about a second, while its mesh is still generating, so you are never staring at a spinner. The second is that you can talk to it. On the phone you can point at an object and ask whether it fits beside your desk. In the headset you can say "find me a lamp under 1.5 m" and watch storefronts get searched in parallel, then pick a result and see it drop into the room at its listed size.
Every number the assistant speaks is checked against a tool result before it is said out loud.
How we built it
We built Scale as four parallel workstreams against a single shared contracts file that defined every schema, route, and storage key.
- Cloudflare is the backbone. Workers are the only front door, with D1 for rooms and objects, R2 for meshes and frames, Vectorize for 768-dimension SigLIP2 embeddings, Queues and Workflows for generation jobs, and a Durable Object per room for live updates over SSE. A second Worker on the Agents SDK runs the designer agent, one Durable Object per room, holding the decision log and version history and calling OR-Tools for placement. Before it plans, it runs a deterministic cleaning pass over the room state for duplicate placements, objects with no usable size, and door arcs pointing outside the room, and writes each decision into a log you can read in the headset. Tunnels connect the Python services back to the edge.
- Baseten runs Stable Fast 3D for image-to-3D generation, plus a binding library that rescales the generated mesh to the measured box, moves the origin to bottom-centre, and then re-parses its own exported bytes to verify every axis before anything downstream sees the file.
- Browserbase recovers the dimensions that product APIs do not expose. Shopify's
/products.jsondoes not serve metafields, which is where most merchants keep their dimensions, so we render the actual product page and read schema.org JSON-LD, then a labelled spec block, then page text. It pulled real dimensions for 262 products the API had nothing for. - Shopify storefronts are the live catalog, and getting trustworthy numbers out of them is the hardest part of the system. Dimensions arrive as prose in
body_html, as variant titles like60" x 30", in metafields the API does not serve, inside a spec diagram image, or nowhere at all, in inches or centimetres or mixed units. We escalate through four sources in cost order: a regex pass over text, an LLM pass over the remaining prose under a strict schema, a rendered page pass through Browserbase, and a VLM pass on spec images. Every source, model or regex, hands its answer to the same validator. It checks unit sanity, category priors judged against every plausible reading of what the product is (an L-shaped sectional is 2.97 m deep and would fail a sofa prior), and whether width and depth were swapped. What survives carries a confidence score that starts from which source produced it, and below a threshold the UI says "unverified fit" instead of showing a number as fact. We verified 28 candidate merchants this way and built a 100-product catalog across seating, surfaces, storage, lighting, and sleeping. - ElevenLabs handles voice in the headset. The Quest browser has getUserMedia but no Web Speech API, so we record push-to-talk audio and proxy STT and TTS through a Worker that holds the key.
- The rest: Expo and Expo Router with four custom Swift native modules on iOS, three.js and WebXR with physics for the Quest, and FastAPI services in Docker for catalog ingest, search ranking, and the fit solver.
One architectural decision did most of the work. The scale binding happens exactly once, in one component. If two parts of the system both rescale a mesh, the error compounds and nobody notices until it's placed. Writing that rule down before writing code meant every other component could treat a mesh as already correct.
Challenges we ran into
The hardest problems were about trusting the numbers, not about feature count.
- Merchant data is much worse than it looks. A store can serve a perfect product API and still carry no dimensions in it, because the numbers live in metafields that the endpoint does not return. Four of the stores we verified were green on every reachability check and worth nothing. We added a rendered-page pass through Browserbase to recover them, which pulled dimensions for 262 products that the API had nothing for.
- LLM output needed strict guardrails. A language model reading dimensions off a page is a less trustworthy source than a regular expression. We constrained every model call to a strict schema, refused answers with no stated unit, and sent everything through the same validation layer regardless of where it came from. That layer checks unit sanity, category priors, and whether the axes were swapped.
- Cold starts nearly cost us the demo. The first generation request against a fresh Baseten deployment took around five minutes to activate and another sixty seconds to answer. Once a replica is warm, the same request takes 1.19 seconds.
- Cloudflare Workers on the same account cannot call each other over their public URLs. The platform answers with its own 404 page, which looks exactly like a routing bug and is not one. We moved to service bindings.
- Coordinate conventions caused the subtlest bugs. Converting a LiDAR depth pixel into a world point requires a sign flip between two different camera conventions, and a wrong rotation sign does not show up as a wrong size, it shows up as a correctly-sized object facing the wrong way.
Accomplishments that we're proud of
We got a generated mesh to match a measured object to within a millimetre, verified by re-parsing the exported file rather than by trusting the export step. The measured error on our real test artifact was effectively zero.
We built a search where a beautiful match that does not fit is never returned. Fit is a hard filter and style only ranks what passes it. This is the specific failure the product exists to prevent, and it is enforced by a test rather than by convention.
We built a voice assistant that structurally cannot invent a measurement. The first model pass has no tool results in its context, so it has nothing to state a number from, and a second check verifies every number spoken against the tool results for that turn.
We assembled a 100-product catalog with real validated dimensions across five furniture categories, with more than half recovered from sources the product APIs did not expose.
Nothing in the pipeline guesses. A dimension with no stated unit is refused, a range is refused, a number that fails validation outright is rejected, and one that merely disagrees with its category prior is kept with its confidence lowered and the disagreement recorded against it. Being able to say how sure it is, and why, is the difference between an agent and a scraper.
What we learned
We learned that the useful question is not what an incumbent does badly, but what its business model prevents it from building at all. Reframing around that took an hour and set the direction for everything afterwards.
We learned that a language model doing geometry does not fail loudly. It fails confidently and plausibly, which is much worse in a product whose entire claim is dimensional honesty. The fix was structural rather than prompt engineering: the model turns intent into an objective and constraints, and a solver places things.
We learned to measure before optimising, and also before deciding not to. We tested a lower texture resolution expecting a speedup, found it saved almost no time and visibly degraded the result, and kept the higher setting. That was a few minutes of instrumentation that stopped us from making the output worse for nothing.
Finally, we learned to decide collisions in advance. We knew exactly which file two people would fight over, and we resolved it in the first two hours by having one person own the native side and expose a plain interface for the other to build against. That one decision saved us an entire category of merge problems later.
What's next for Scale
Moving day is the obvious next product. "Will my furniture fit in the new place" is a recurring, genuinely painful question that nothing serves well, and every piece we need already exists: your possessions are measured, the new room is scannable, and the solver already explains why an arrangement is impossible instead of just failing.
Beyond that, we want to index catalog items when they are written rather than when their mesh is generated, so the vector search path is live from the first query. We want to close a gap in lighting data, where two of our merchants publish height but not depth. We want to move the assistant's API key off the device and behind a proxy route. And we want to support more than one room per user, which is a schema change we have already scoped and deliberately deferred.
Built With
- baseten
- browserbase
- cloudflare
- cloudflare-agents-sdk
- cloudflare-d1
- cloudflare-r2
- docker
- durable-objects
- fastapi
- object-capture
- or-tools
- python
- react-native
- roomplan
- shopify
- swift
- three.js
- typescript
- vectorize
- webxr






Log in or sign up for Devpost to join the conversation.