About the project

Inspiration

A language model can describe any structure you want. It has no idea what the machine in front of it can physically do. Ask one for a design and you'll get blocks floating in mid-air, blocks the gripper can't fit between, and blocks sitting inside the robot's own chassis.

Blocks turned out to be the perfect testbed for closing that gap. Putting a block on a block is the same problem as laying a course of masonry or stacking mixed cases on a pallet: a lattice, support constraints, a placement order, a reach envelope, and a grasp that has to hold. It also fails visibly from across the room, which is what you want at 3am.

What it does

Give it a sentence or a photo. Either one becomes a real build.

1. Design. Prompts go through Claude with structured output. Photos get downsampled onto the build grid and reduced to buildable line art, either with a vision model or with a deterministic tracer that needs no model access at all. Cats, dogs, rockets and houses have all gone through both paths.

2. Check. Every design hits a deterministic validator before anything moves: schema, support, connectivity, finger clearance, colour inventory, reach. When a design fails, the system repairs it, re-checks, and reports every block it dropped and why. A negative control lives in the repo so we can show the checker failing on purpose.

3. Plan. The structure becomes staged positions, a placement order, a tool yaw per block, and a joint trajectory from joint-limited IK, gated on a reach map measured from the actual arm and filtered against the robot's own chassis.

4. Build. A learned pickup policy and contact-aware placement run the plan cube by cube in contact simulation, with free-body cubes and nothing welding them to the gripper. The same trajectory streams to the real BracketBot arm over bbos.

build_from_description.py "a small house"
image_to_structure.py photo.png        # same pipeline, from a picture

The real use case

Every time an arm gets deployed into a new cell, an engineer does the same week of work by hand: survey the workspace, confirm what's reachable, tune the grasp, sequence the placement order, then discover halfway through that the bench is too high. Change the part size or move the table and you do it again. That cost is why arms live in fixed cages doing one task forever, and it's the real blocker for robots on job sites and in small warehouses, not the manipulation.

MasonBot automates that week. It probes the arm's reach itself, excludes its own chassis, validates a design against those measurements, and refuses work it can't do with the reason attached. Our own run is the proof: re-probing found our table at 0.78 m had 146 reachable cells against 542 at 0.5 m, because the arm runs out of vertical travel and stretches flat. That's a workspace verdict in minutes rather than a wasted day.

The same loop drives robotic block laying and dry-stack walling on site, mixed-SKU palletizing where stack stability and order matter, and modular fixturing. All of them are lattice assembly under support, order and reach constraints. And because the verification comes from measurements rather than assumed precision, it works on a cheap arm, which is what makes it deployable outside a cage.

The part we're proudest of: how we evaluated pickup

Pickup is learned, not scripted. A two-layer neural policy outputs bounded corrections to tool XYZ, yaw and jaw on top of a deterministic phase controller, trained by behaviour cloning on 128 episodes (18,724 state-action pairs).

The evaluation is the part worth stealing. We'd already looked at our original test seeds while developing an earlier checkpoint, which quietly poisons them. So we froze the model first, then reserved a fresh 32-episode test set, binding the checkpoint and physics hashes so it can't be overwritten.

On that frozen set the policy went 32 for 32, and the scripted expert went 32 for 32 on identical scenarios, so the policy matches its teacher rather than quietly underperforming it. Success isn't "it looked fine": each episode needs a 40 mm lift, a one second hold, 0.5 N at each finger throughout, under 2 mm slip, and no solver warnings. Open-finger and no-motion controls fail, as they should. Scenarios shift the cube up to 30 mm and 0.15 rad and vary friction and mass, so the policy is recovering from misalignment rather than replaying one trajectory.

Placement closes the loop: the runner catches insertion collisions against cubes already down, corrects pose on the way in, and detects a wedged cube before release, inside 3 mm and 5 degrees. A two-layer fixture lands at 0.405 mm worst case.

(Fill in your whole-build number before submitting, and see the notes at the bottom.)

How we built it

The bug we're most glad we caught. IK converges perfectly happily at points inside the robot's own body, and nothing warns you. 147 of 542 mapped cells sat inside the chassis, and since the planner built "as close to the robot as possible," most of one design was going inside the machine. We now filter against a base footprint measured off the URDF meshes, kept separate from the reach map because one is a fact about the arm and the other about the chassis.

The workspace is a ring around the base, not a patch to one side, so "smallest y" was never "nearest the robot." Origins now rank by radial distance, then margin.

Grasp geometry. Clipping the real gripper mesh triangles showed we were gripping near the block edge, because the nominal tool origin sat over the block centre. The grasp point now carries an 18 mm offset onto the face where the foam actually makes contact.

Contact physics. MuJoCo with CoACD-decomposed geometry, free cubes under gravity, no cheat constraint. Nominal cases place within 0.75 mm against a 3 mm gate, and the controls fail correctly: open fingers fail the lift, low friction misses at 3.46 mm, zero friction gets flagged numerically invalid instead of counted as a pass.

59 sim tests, 52 compiler tests and 10 learning regression tests pass.

Challenges we ran into

  • Silent IK. Solvers won't tell you the pose is inside the robot. It looked fine in Rerun because the chassis mesh isn't a collider.
  • Probe windows lie. Our first reach map didn't measure the arm, it measured the box we probed in.
  • Contaminated test seeds. Throwing out seeds we'd already inspected cost a full evaluation cycle, and it's the one decision that makes 32 for 32 mean anything.
  • Placing beside what you already placed. Pickup in open space is the easy half. Builds break descending into a gap next to existing cubes, and that took collision detection during descent, not a smoother trajectory.
  • A shared robot. The machine's app directory is shared with other teams and isn't version controlled, so we wrote push and pull scripts to restore our state after someone else has been on it. ### What we learned

Most of the work in getting a language model to drive a robot has nothing to do with motion planning. It's building the thing that tells the model no, and those checks are what transfers to a different arm, workspace or part. We also learned to be ruthless about evaluation hygiene. A number from a seed you've already stared at isn't a result.

What's next for MasonBot

  • Held-block detection with no new hardware. Gripper current already shows up in the arm state, and a close that hits its commanded angle with no current rise means an empty jaw. Cheapest closed loop available, and it turns a simulated build into a verified physical one.
  • Supervised hardware trials. 18 of 20 clean pick-and-place cycles before repeated multi-cube builds.
  • Perception, so the robot can work from blocks scattered and unsorted rather than pre-staged.
  • Full-arm dynamics in the loop, including arm, table and chassis collisions.

Built With

  • bracketbot
  • ik
  • il
  • llm
  • mujoco
  • python
  • rerun
  • rl
  • voxels
Share this project:

Updates

Submission history