Inspiration
Every robot simulator lets you command a robot. None of them let you change one.
So the moment the arm can't reach the top shelf, you leave the tool. You open the URDF in a text editor, edit the XML by hand, recompute inertias, push all the child joints out, and hope the origins still line up. Reload, find out you got the rpy wrong, do it again.
Redesign is the loop. Reach, fail, make the arm longer, try again. We wanted that to be a sentence instead of an afternoon.
The setup cost around it doesn't help either. Doing one thing in simulation usually means installing ROS, fighting Gazebo, writing a URDF by hand, and scripting a controller before anything moves at all. It's tooling built for people who already know robotics, which is a problem if you're trying to learn it.
What it does
Drop in a URDF and it renders as a live 3D model with the full joint tree, limits, and an end effector you can drag around with IK. Upload your own or pick one from the library.
Then you talk to it. There are two kinds of things you can say.
The first kind changes what the robot does. "Pick up the green cube and stack it on the blue block." "Drive forward for 3 seconds, turn 90 degrees left, repeat 4 times." "Move forward at 0.5 m/s until the front sensor sees something within 30 cm, then avoid it." The planner gets a live JSON view of the scene with joint angles, object poses, and distances, and sends back a validated action graph instead of prose. The sim runs it.
The second kind changes what the robot is. "Add a parallel gripper to the wrist." "Extend the forearm by 20 cm." "Let the elbow rotate further." The model gets rebuilt and hot reloads in place, keeping whatever joint state it was already in. Then you run the command that just failed and watch it work.
Scenes work the same way. Describe one in a sentence, like a 3x3 warehouse grid with four red hazard cubes and a blue loading dock in the far corner, or place things by hand with drag and drop.
How we built it
The frontend is Next.js 15 and React 19. urdf-loader parses the uploaded robot into a link and joint tree and React Three Fiber renders it, so what you see is a live kinematic model with real joint limits, not a video feed. Zustand holds the scene graph, sim state, and chat log. A FastAPI backend runs alongside it.
The planner is Claude, called through the Vercel AI SDK with structured output. Nothing reaches the simulation until it clears a Zod schema. It returns typed actions like MOVE_VELOCITY, ROTATE_JOINT, and AWAIT_CONDITION, never free text, and the prompt carries a compact view of the scene: joint angles, object poses, distances, and the robot's current chain.
The whole thing rests on one decision. The model never touches URDF XML. It emits an operation instead, something like {"op": "extend_link", "target": "forearm", "axis": "z", "delta": 0.20, "anchor": "base"}, and plain code does the edit against the parsed tree. The uploaded file is never modified. Every change is a patch replayed on top of it, which is where undo, diffs, and reset come from for free.
We tried it the other way first. Ask a model to rewrite 400 lines of XML and it will drop a closing tag, or rename a link and leave the three joints referencing it pointing at nothing. Ask it to pick an operation and fill in two numbers and it's reliable. That line is why the rebuild loop works at all.
Edits get validated before hot reload: one root, no cycles, positive masses, no self collision in the default pose. Failures go back to the planner to retry.
Challenges we ran into
Getting a language model to drive a simulator was the hard part, and it's the part we're still working on. The model has to know enough about the scene to plan against it, but a full state dump is too much context and a summary that's too thin produces plans referencing objects that aren't there or joints the robot doesn't have. Most of our time went into that seam: what goes into the prompt, how tightly the schema constrains what comes back, and what to do when a plan validates but still doesn't make physical sense.
Giving the planner enough ROS grounding was a separate fight. A language model knows what a revolute joint is in the abstract, but it doesn't know that your URDF measures in meters and radians, that joint names in a real robot description follow conventions it can't guess, or which of the actions in our schema apply to a wheeled base versus an arm. We kept widening what the prompt carries. The spec constraints that actually matter, worked examples of valid action graphs, the robot's real chain rather than a description of it. Every addition fixed one class of bad plan and cost us context budget, and finding the line between too little grounding and too much was most of the tuning.
Coordinate frames ate the rest. Mounting a gripper means lining its approach axis up with the parent flange's frame, and if the rpy is off you get a gripper bolted on sideways. It looks like a physics bug. It isn't.
Extending a link turned out to be four things, not one. URDF geometry is centered on its own origin, so growing a box means resizing it, shifting the visual and collision origins by half the delta to keep the base anchored, pushing every child joint out by the full delta, and recomputing inertia from the new dimensions. Skip any one and the arm either floats off its parent or sinks into it.
The tool center point caught us late. IK has to aim at the gap between the fingertips, not the wrist flange, or every grasp misses by exactly the length of the gripper.
Real URDFs are a mess. Unresolved package:// paths, missing inertials, zero mass links, joints with no limits. We wrote a repair pass that fixes the mechanical stuff automatically and only asks about genuinely ambiguous cases.
Accomplishments that we're proud of
We pivoted mid-hackathon. The plan was hardware, and when that stopped being viable we dropped it and rebuilt as pure software with most of the clock already gone.
What we're proud of is that the pivot made the project better rather than smaller. Cutting hardware is what let us chase the robot-modification loop, which turned out to be the interesting idea and the one nobody else is demoing.
Beyond that, the rebuild pipeline works end to end. A sentence becomes a structured edit operation, code performs the surgery on the parsed URDF, the result gets validated, and the robot reloads in place without losing its joint state. That's the part we expected to be flaky and it isn't.
What we learned
Models are good at deciding what should change and bad at carrying out the change. That sounds obvious written down, and it cost us a few hours to actually believe it.
Everything in this project that works sits on the right side of that split. The model picks the operation and fills in the numbers. Code does the surgery, checks the result, then reloads. Every time we handed the model something structural and asked it to produce the artifact directly, we got output that looked correct and quietly wasn't, which is the worst thing to debug because nothing catches it.
The other thing we learned is that URDF is stricter than it looks. A robot is a tree with real geometric constraints, and most of our bugs weren't model bugs at all. They were us getting an origin offset wrong by half a delta.
What's next for NL-Robot
Planner reliability is the thing we're still on. Right now we improve it by hand, tightening the scene summary and the action schema against cases we've watched fail. The better version is a task suite we can run in batch so a prompt change gets measured instead of eyeballed.
Closing the loop is next after that. Plans currently run open loop, so if a block slips out of the gripper nothing notices. We want to check sim state after each step, ask whether the block is actually held, and replan when it isn't.
Then real IK and collision aware motion planning instead of interpolating joints, and a ROS 2 bridge. URDF is already what we take as input, so the same plans should drive a real robot without a translation layer in between.
Built With
- anthropic
- claude
- drei
- fastapi
- framer-motion
- next.js
- python
- radix-ui
- react
- react-three-fiber
- robotics
- tailwindcss
- three.js
- typescript
- urdf
- urdf-loader
- uvicorn
- vercel-ai-sdk
- webgl
- zod
- zustand
Log in or sign up for Devpost to join the conversation.