Inspiration
No AI today can improve itself. But a robot can catch itself being wrong, if it has a third-person view of itself.
Every robot we've built or used runs on a fixed model of its own hardware. Swap a motor, reverse a wire, and it keeps driving confidently on a model that is no longer true, until a human notices and recalibrates it. Machines can't tell that their own body has changed.
We wanted the opposite: a robot that is never told how it works, learns its body from scratch, and works out for itself when that knowledge has gone stale. The trick is the camera. It gives Darwin an outside view of its own motion, so the gap between what it predicted its wheels would do and what the camera saw them do becomes a training signal it can act on.
What it does
Darwin is a two-wheel robot driven by two abstract motor commands and watched by one overhead camera. It has no wiring diagram and no calibration file.
- Learns its body. It runs 24 probe pulses, watches itself move, and fits a model of what each command does. The camera supplies the labels. It then scores that model on 12 held-out trials it never trained on.
- Drives. Click anywhere in the arena and it plans with its own learned model.
- Gets sabotaged. Mid-drive we reverse both wheels, swap them, or weaken one. Darwin is never told.
- Notices. Its predictions stop matching the camera. After three failed predictions in a row it stops itself. One noisy pulse can't trigger it.
- Experiments. It freezes the old model and designs its own experiments, choosing the commands it knows least about.
- Relearns and resumes. It refits, validates on fresh held-out trials, and finishes the drive it started.
Break it again with a different sabotage and it runs the whole loop again. We verified two different mutations in a single run: three model generations, target reached.
Darwin also thinks out loud. A live inner-monologue panel narrates what it is doing in one tiny first-person line, for example "Who scrambled my controls?". OpenAI phrases each thought from measured runtime facts and ElevenLabs speaks it aloud. This layer is narration only. It never touches the motors, and any rewrite containing a number the runtime facts don't support is thrown away.
Results (45-case simulation benchmark, 180 navigation runs):
- Stale model: reached the goal 5 / 180 times.
- Relearned model: reached the goal 180 / 180 times.
- Prediction-error reduction: 89.95% minimum, 97.73% mean.
How we built it
A complete machine learning loop with no human in it. Darwin collects its own training data, labels it with the camera, validates on held-out trials, deploys the model, monitors it, detects when it has gone stale, designs new experiments, and retrains.
- Model. A small affine ridge regression maps the two motor commands to forward speed and yaw rate,
[uL, uR, 1] → [v, ω]. Six coefficients, closed-form, fit on CPU between motor pulses: $$\beta = (X^\top X + \lambda D)^{-1} X^\top y, \quad D = \mathrm{diag}(1,1,0)$$ It is deliberately small: interpretable, identifiable from 24 samples, and every fit is gated by held-out validation before it is allowed to drive. A model that can't beat the stationary baseline is rejected. - Concept-drift detection. A sequential residual detector compares each prediction with the camera. It needs a normalized error above 0.25 for 3 consecutive pulses and a minimum amount of evidence, so noise doesn't cause false recoveries.
- Active learning. Recovery experiments are chosen by leverage over the action space, using the information matrix (XᵀX + λI)⁻¹, so the robot spends its few safe pulses where it is least certain.
- Continual learning. Every model generation is versioned, checkpointed, and compared to the frozen pre-change model on the same held-out set. That comparison is the ablation.
- A blind learner, enforced in code. The learner never receives the simulator's hidden map, the mutation type, or any privileged inverse. It sees only the commands it sent and the camera-measured result. The UI separates what the operator did from what Darwin inferred, and never shows "change detected" just because a button was pressed.
- Vision. An overhead OAK camera tracks an ArUco marker mounted on the robot. We calibrate the marker plane with a homography and recover pose in metres. On synthetic frames the tracker measured 0.946 mm position error over 150 frames.
- Hardware. An Arduino Uno R3 and a TB6612FNG motor driver, with firmware written for fail-closed behavior and reviewed line by line. It never auto-flashes.
- Safety. Boots disarmed. STOP, stale frames, serial loss, operator-lease loss, boundary limits, and the firmware watchdog all cut motor output.
- Same code, sim and hardware. Only the camera and actuator adapters differ. The runtime, learner, controller, safety state machine, recorder, replay, and dashboard are identical.
- Stack. Python, NumPy, OpenCV (ArUco), FastAPI, WebSockets, and a dependency-free HTML/CSS/JavaScript dashboard. Everything except the optional voice runs locally.
Challenges we ran into
- Keeping the learner honest. It is easy to leak the answer. Every visualization and every line of narration had to be built from public learned state only, with the mutation map fenced off from the learner.
- Telling a real change from noise. Our first detectors either missed changes or fired on jitter. Requiring consecutive normalized exceedances plus minimum evidence fixed both.
- Camera-as-sensor is unforgiving. The marker sits above the robot, so the marker plane and the floor plane differ. Getting metric pose right needed a proper plane-aware calibration, and the calibration is invalid if the camera or arena moves.
- Real camera failures. The OAK stream threw an
X_LINK_ERRORand a native crash mid-session. Losing the camera has to stop the motors, so we made stale frames a hard stop rather than a warning. - Making an LLM safe to narrate. Language models invent numbers. We discard any rewrite whose figures aren't backed by the runtime facts, and keep the narration off the control path entirely.
- Not fooling ourselves. We kept failures in the denominator, and we separate what we measured on hardware from what we measured in simulation.
Accomplishments that we're proud of
- A robot that diagnoses itself from outside evidence and recovers, with the learner blind to what changed.
- 5/180 to 180/180 navigation successes after adaptation across a 45-case benchmark (3 seeds x 3 plant variants x 5 mappings), with failures kept in the denominator.
- Two different sabotages in one run, three model generations, target reached.
- Real hardware evidence: hundreds of ACKed serial commands at roughly 3 to 7 ms latency, a recorded run of 177 actions and 175 camera-measured transitions, and a physical learned-navigation run that ended 5.7 cm from the target after 10 actions.
- Held-out non-overlap is enforced: the code raises if train and test pulses overlap.
- 240+ automated tests, deterministic replay of every run, and a read-only replay mode that can never be mistaken for live hardware.
- A voice for the robot that never has a path to the motors.
What we learned
- The sensor you add is the model's real teacher. Giving the robot an outside view of itself turned an unsolvable introspection problem into ordinary supervised learning.
- A small, honest model beats a big vague one. With 24 samples, ridge regression is identifiable, checkable, and fast enough to refit between motor pulses, and we could prove each fit was valid before letting it drive.
- Detection matters as much as learning. Recovering was the easy half. Knowing when to recover, without false alarms, took most of the work.
- Separate what the operator knows from what the machine has inferred. Doing that visibly made the demo more honest and more convincing.
What's next for Darwin
- Real uncertainty. Replace our RMSE-based confidence with a proper posterior, σ²(XᵀX + λI)⁻¹, so the active-experiment choice is provably information-maximizing.
- Richer bodies. Add quadratic terms and choose between model classes by held-out error, so Darwin can adapt to nonlinear damage such as friction, dead zones, and slip.
- Camera-loss hardening. Reproduce and fix the OAK disconnect path and prove that losing the camera always stops the motors.
- More physical trials. Run supervised live mutations on hardware across all mappings and publish the physical results next to the simulation benchmark.
- A memory that persists. Match new bodies to previously learned ones across sessions instead of relearning from scratch.
- A 3D sensorimotor landscape of the learned command-to-motion surface, so anyone can watch Darwin's model bend when its body changes.
Built With
- active-learning
- arduino
- arduino-uno
- aruco
- c++
- canvas
- computer-vision
- css3
- depthai
- elevenlabs
- fastapi
- html5
- javascript
- luxonis-oak
- machine-learning
- numpy
- openai
- opencv
- pyserial
- python
- ridge-regression
- robotics
- tb6612fng
- uvicorn
- websockets
Log in or sign up for Devpost to join the conversation.