Hydrogen bomb (Astra, Fable) vs. coughing baby: who understands object permanence?
Inspiration
Frontier models like Fable seem more intelligent in almost every way. They write software, explain advanced mathematics, and work through hard problems. But on one fundamental concept, a roughly one-year-old baby has something to teach them: object permanence.
Object permanence is understanding that things keep existing when we can't see them. It lets us remember what's behind an obstacle, follow where it went, and look for it again. Babies show it long before they can explain it.
This isn't a claim that babies beat AI on a benchmark. It's the gap that interested us. A model can explain object permanence, but a robot still needs a way to keep it in a changing physical world. A detector losing sight of something shouldn't mean the robot forgets it exists.
We built Permanence to turn "I detected it" into "I still know where it is."
What it does
Permanence gives a robot a persistent 3D memory of objects. The rover scans the table by itself, identifies the objects, and keeps track of where they are as they get hidden or leave the camera's view.
Our clearest demonstration is a shell game. Show the robot colored balls, cover them with cups, shuffle the cups, then reveal. Permanence tracks which cup holds each hidden ball and checks its answer against the real reveal. It called the right cup in 7 of 7 recorded games and 3 of 3 live.
You can also:
- Talk to it. "Hey rover, the cup next to the green ball": a 3D vision-language model picks the object from the scan and shows how confident it is.
- See its memory. A live 3D dashboard and an x-ray display on XREAL glasses show hidden objects through the cups covering them.
- See what's seen versus remembered, including how long each object has been out of sight.
- Send it to a remembered target. "Hey rover, drive": the rover plans around the other cups and goes to its target, even when nothing on camera shows where the ball is.
The cups are a demonstration, not the whole idea. The goal is a robot that acts on what it knows, not just on what's in frame.
How we built it
Two brains, one Jetson. Everything runs on an NVIDIA Jetson Orin Nano (8 GB). An iPhone on the rover is only the sensor: it streams color, LiDAR depth and its own position over USB through Record3D.
- The reasoning: Qwen-3D, a 3-billion-parameter 3D vision-language model. It reads 18 views of the scan and turns language like "the cup next to the box" into a specific object, in about 4.5 s. Three identical cups have no names; only a model that understands language and space can tell them apart.
- The fast reflexes: YOLOE. It's prompted with words (no training on our props) and compiled to a 22 MB TensorRT FP16 engine that keeps up with the camera's 15 fps.
- The memory between them: a 3D tracker we wrote. It matches detections across frames, follows each object through occlusion, and knows that a ball under a cup moves with the cup.
Fitting it in 8 GB. CPU and GPU share the Jetson's 8 GB, and Qwen-3D alone was 7.1 GB. We wrote our own 8-bit quantization in plain PyTorch, loading the model one tensor at a time so the full-size version never exists in memory: 4.5 GB. The pipeline allows one memory peak at a time. Qwen-3D picks the target, hands it to the tracker, and its process exits to free its memory before tracking starts. A full run peaks at about 6.1 of 7.4 GB.
Interfaces. A Three.js dashboard shows the robot's calibrated view, a top-down memory map, the camera feed, the confidence behind each pick, and live Jetson memory. A web HUD on XREAL glasses brings the x-ray view and key state to a wearable display. Voice commands use Chrome's speech recognition.
Movement. The Jetson drives an Arduino Uno and L298N motor driver and steers from the phone's position. The rover scans by stop-and-shoot: turn, stop, take a sharp view, roll, repeat. It learns how much power its cheap motors need to start moving, and stops on its own if commands stop arriving. We also built a separate head-steering interface with a Pico W and GY-521 motion sensor.
Challenges we ran into
- 8 GB is small for two models. Loading both crashed the board, and an earlier run left behind can eat the room. We added the handoff, our own quantization, and cleanup of leftover processes.
- A wobbly camera. With the phone mounted high on the rover, scanning while moving gave blurred frames. Stopping for each shot fixed it.
- Cheap motors. They stall when started too gently and hum instead of turning, so the rover measures its own speed and adjusts power. When its left and right turned out to be swapped, it detects that and corrects it in software.
- Disappearing isn't the same as hidden. A missed detection doesn't mean the ball went into a cup. We needed geometry, motion and containment checks, and matching in 3D so identical cups don't swap identities when they cross.
- Freshness over completeness. An overlay that shows where the world was seconds ago is useless, so every stage takes only the newest frame.
Accomplishments that we're proud of
- The full loop on a palm-sized edge board: perception, memory, language, visualization and physical action, with a 3B vision-language model and a real-time tracker together in 8 GB.
- The reveal. The robot commits to where the hidden ball is, and we lift the cup to check. Nobody has to take our word for it.
What we learned
- Recognizing an object is only the beginning. Useful intelligence needs continuity: what was there, what changed, and what should still be there when you can't see it.
- Memory is a design constraint, not an afterthought. Deciding what runs when mattered as much as which models we picked.
- The strongest demo is a claim you can check: show it, hide it, move it, reveal it.
What's next for Permanence
- Prediction: a learned world model that guesses what happened out of view, instead of only remembering what it saw.
- Bigger hardware: Jetson Thor, so the language model can stay loaded and answer at any time.
- Better hardware and integration: wheel encoders, and ROS 2 / Isaac ROS.
- Beyond the tabletop, with clearer uncertainty when memory gets old.
- Longer term: assistive robots and glasses that help people find misplaced things.
The coughing baby had this figured out in its first year. Now our robot does too.
Out of sight should not mean out of memory.
Built with: NVIDIA Jetson Orin Nano, TensorRT, CUDA, PyTorch, Qwen-3D, YOLOE, Record3D, iPhone LiDAR, Python, Three.js, JavaScript, Web Speech API, XREAL, Arduino, L298N, Raspberry Pi Pico W.





Log in or sign up for Devpost to join the conversation.