Inspiration
After a building collapse, the first hours decide who survives, and the most dangerous job is going in to look. We wanted a rover that could do the looking: drive into a disaster zone, map it as it goes, notice a person (a shout, a figure in the rubble), go to them, and talk to them until human rescuers arrive. Not a remote-controlled camera, an agent that decides.
What it does
C.H.U.D. is an autonomous rescue rover with an LLM brain. It:
- Explores and maps unknown terrain with frontier-based exploration over an occupancy grid built from LiDAR, camera depth and ultrasonic readings.
- Detects survivors by camera (field of view + line of sight) or by microphone (a call for help gives a noisy bearing and distance).
- Decides to go there. The detection wakes the brain, which calls a bounded
approach(bearing, distance)verb that plans an A* path around rubble and stops a metre short of the person. - Narrates every decision on a live dashboard: 3D driving view, 3D occupancy map, sensor gauges, the brain's tool-call feed and the two-way transcript.
Everything the brain does is one of 17 tools. It cannot invent a motor command.
How we built it
Two loops, two speeds. Cloud calls take 0.5 to 3 s, far too slow for control, so the robot runs a fast reflex loop (10 to 30 Hz, classical, no network: obstacle avoidance, frontier exploration, path following, watchdog) and a slow deliberative loop (the LLM brain, event-driven). The reflex loop always wins at the arbiter. If Wi-Fi dies, the rover still stops at walls.
The brain is Backboard. One persistent thread with tool calling, RAG over real FEMA/USAR
rescue protocols, structured memory of past encounters, and Gemini as the planner model
(BYOK). Verbs are bounded and self-completing: forward(0.5) returns completed or
stopped_by_obstacle, and that result is fed straight back to the model (Inner Monologue)
so it can react within the same episode. A separate survivor-persona agent on Backboard
plays the person being rescued, with a hidden condition, so the conversation is genuinely
two agents talking.
Hardware pipeline. A DJI RoboMaster EP Core captures webcam video and mic audio with FFmpeg and streams both over TCP to a laptop, which runs monocular depth (Depth-Anything-V2-Small) and local transcription (faster-whisper) and sends detections back. Sensor fusion lands in a log-odds occupancy grid streamed to a React + three.js dashboard over WebSocket.
Virtual demo. The same brain, tools, RAG, memory, TTS and MongoDB tracking drive a simulated rover in a 3D disaster scene in the browser. Survivors spawn at random free locations, sensors are modelled with range, FOV and noise, and the loop of explore, detect, approach, converse, and log runs end to end.
Training loop. A scenario generator produces messy disaster episodes, a trace The collector records the brain's decisions, an evaluator scores protocol adherence, and the Good traces become SFT data for Baseten fine-tuning. Essentially, our encounters and incidents get sent to MongoDB, which our agent takes, extracts, and starts training on. This is a sense of mini reinforcement learning to further our model's capabilities and quality in being able to detect when someone is unwell or if there are any obstacles in the way of reaching that person, as we've improved our model routes over the course of the past 12 hours.
Challenges we ran into
- Streaming video fast enough to steer on. Getting frames off the Pi, through depth inference on the laptop, and back as detections with low enough latency to matter.
- Keeping the LLM out of the control loop. Every time we let the brain touch velocity directly it either stalled the robot or drove it into things. The fix was architectural: intent-only verbs, an arbiter, a watchdog.
- Making the brain act reliably. Tool outputs that took long to submit came back as duplicated tool calls; the rover once announced "Mission initialized" five times. We now dedupe by tool-call id and run tool execution off the event loop.
- Navigation that actually gets there. Reactive steer-away logic oscillated along rubble edges forever. Planning A* on the observed grid and re-planning every second as the map fills in fixed it.
- Hardware that fought us. Ultrasonic sensors were unusable from chassis vibration, so camera depth became the only obstacle sensor.
Accomplishments that we're proud of
- A rover whose every action is an auditable LLM tool call, sitting on top of a safety layer that never needs the cloud.
- Two AI agents holding a real rescue conversation in two voices, driven by sensor events.
- Frontier exploration + A* on a map the robot built itself.
- Pi-to-laptop video, audio and depth pipeline over plain TCP.
What we learned
- Fast video streaming over TCP and where the latency actually hides.
- Cloud inference routing: which calls belong in the slow loop, and how to make the robot degrade gracefully when they fail.
- Hierarchical control is the whole game for LLM robotics: SayCan-style gating, Code-as-Policies verb sets, and Inner-Monologue feedback.
- Fine-tuning a small command parser on Baseten to skip the big model for simple commands.
What's next for C.H.U.D.
- Close the loop on the physical RoboMaster with the same verbs the simulator uses.
- Real thermal and IMU sensors, and microphone direction-of-arrival for audio bearing.
- Multi-rover coordination through shared Backboard memory.
- Human ratings of the rover's speech fed back into the fine-tuning loop.
Built With
- backboard
- baseten
- elevenlabs
- gemini
- lora
- mongodb
- python
- qwen
- raspberry-pi
- react
- robomaster
- three.js
- typescript
- websockets
- whisper

Log in or sign up for Devpost to join the conversation.