-
-
The real physical MVP we built — ESP32, L298N driver, 3D-printed capture mechanism — closing in on a test roach.
-
A closer look at our real hardware: ESP32 + L298N wiring, battery pack, and the 3D-printed capture box we built and tested.
-
YOLO dockroach detection
-
(Simulation) Roach Hunter in standby — a clean, ready-to-deploy interface before the patrol begins.
-
(Simulation) Patrol active: the robot autonomously searches the room in SEARCH mode.
-
(Simulation) Real-time YOLO-style vision lock — the onboard camera detects and tracks a cockroach as the robot moves to intercept.
-
(Simulation) Report a sighting in plain language — our Fetch.ai agent understands the location and dispatches the robot to investigate.
-
(Simulation) When a target is lost, the robot captures the moment and the AI agent narrates what happened with a photo and anexplanation.
-
(Simulation) The infestation heatmap reveals where roaches tend to escape from (red) and where they get caught (green).
Roach Hunter
Inspiration — The Problem
When a cockroach appears in your home, there are usually two problems. First, a lot of people are simply too scared or unwilling to kill it themselves. Second, even when someone tries, cockroaches are fast — one missed hit can scare it away, and suddenly the bigger problem is not killing the cockroach, it's figuring out where it went.
That's what inspired Roach Hunter: an autonomous robot that detects, tracks, chases, and captures a cockroach before it disappears.
Cockroach spotted → immediate response → continuous tracking → capture
Instead of forcing someone to find a brave roommate, grab a slipper, or keep watching the cockroach until they can deal with it, Roach Hunter responds immediately and keeps tracking the target. If it escapes — especially while nobody's home — the system should remember where it was last seen instead of starting from zero.
We built two things to prove this out:
- A physical robot MVP — ESP32, motors, and a 3D-printed capture mechanism, running a real YOLO vision pipeline, proving the detect → chase → capture loop works in hardware.
- A full-vision simulation and AI agent — a 3D demo of the complete intended experience, including a real conversational agent (built on Fetch.ai's uAgents, registered on Agentverse, and discoverable through ASI:One) that a user can talk to in plain language to report a sighting.
How We Built It
Hardware
The robot runs on an ESP32 connected to an L298N motor driver, DC motors, a camera, and a servo-powered capture mechanism. The ESP32 is the real-time control layer: it receives commands and turns them into motor actions (forward, turn left/right, stop, trigger capture), with the servo providing the final grab once the target is close enough.
Python Vision & Autonomous Chase
The intelligence runs through our Python vision program. We use YOLO11n with OpenCV to process the camera feed, extracting each detection's bounding box, center position, and relative size. If the target is left of center, turn left; if right, turn right; once centered, move forward and keep reevaluating. A growing bounding box means the robot is getting closer — once it crosses our capture threshold, it stops chasing and triggers the capture mechanism.
Camera → YOLO detection → target position → chase decision → ESP32 → motors → new camera frame
ESP32 Control System
A separate Arduino/C++ program on the ESP32 exposes movement and capture commands while directly driving the motor driver and servo. Splitting the physical controller from the Python vision system let us iterate on detection and pursuit logic without rewriting low-level motor code — Python decides what the robot should do, the ESP32 makes the hardware do it.
User Interface, Simulation & AI Agent
We built a working simulation and a real AI agent to demonstrate the product experience beyond the physical MVP. In the simulation, you can type a message like "I just saw a roach near the sofa!" — not a mockup. That message goes to a real agent built on Fetch.ai's uAgents, registered on Agentverse, and discoverable/chattable through ASI:One (our entry for the ASI:One Agent Challenge). It uses ASI:One's own LLM (asi1-mini) to understand the location from free-form language, with a local keyword-matching fallback if the LLM is unavailable, and dispatches the simulated robot to stake out that exact spot.
We also built the "what if it escapes" half of the vision. If a tracked roach disappears from view, the simulation captures a still of the onboard camera at that moment and asks the agent's LLM to narrate what happened, grounded in real facts (location, whether it was fleeing or foraging, where it fled toward). That photo and narration appear right in the chat. We also built a toggleable infestation heatmap over the top-down map — red marks where roaches tend to vanish from, green marks where they get caught, accumulated over a session.
These run on the same detect → track → chase → capture state machine the physical robot uses, even though the physical MVP stayed focused on making that core loop work in hardware first.
Specifically: How We Built It
1. A Robot That Could See and Move
We started with the physical foundation: ESP32, L298N motor driver, DC motors, a chassis, and a webcam. Before thinking about cockroaches, we proved the basic perception-to-action loop worked — connecting the ESP32 to the motors, wiring Python to the ESP32, and streaming webcam frames into our vision pipeline. We tested on easier objects first: detect it, figure out left/center/right, translate into motor commands.
Webcam → object detection → steering decision → ESP32 → L298N → motors
2. Designing and 3D-Printing the Capture Mechanism
Detection was only half the problem — we needed a way to physically catch the target. We designed the capture mechanism in Autodesk Inventor and manufactured the parts on a Bambu Lab A1 mini. A servo, controlled by the ESP32, triggers the grab once the vision system confirms the target is close enough. This was one of the most experimental parts — software decisions had to translate into precise mechanical timing.
3. Teaching Roach Hunter to Recognize Cockroaches
Once the robot could see and move, we replaced test targets with the real thing: a YOLO model fine-tuned on cockroach images. Bounding box position drives steering; bounding box size is our proxy for distance. That turned an object-following prototype into an actual cockroach hunter.
4. Integrating Everything and Tuning the Capture
The hardest part was making it all work together — vision pipeline, chase logic, ESP32 controller, motors, servo, and capture mechanism, cycling through SEARCH → ALIGN → CHASE → CAPTURE. We spent a large amount of time tuning steering thresholds, motor speeds, chase timing, confidence/size thresholds, stopping distance, and servo timing. A parameter that looked right on a laptop could behave completely differently once the robot had real momentum — the final few inches were much harder than the first few feet.
5. The Full User Journey — Simulation and a Real AI Agent
Finally, we built a working simulation and agent to demonstrate the complete product beyond the hackathon MVP. A message like "I saw a cockroach near the kitchen counter" is understood by our Fetch.ai uAgent (via ASI:One's LLM), which sends the simulated robot there, searches, and autonomously tracks and captures once a target is detected — on the same state machine the physical robot runs. If the target escapes, the system preserves the last known location and a real captured frame with an LLM narration, so the search doesn't start from zero. The infestation heatmap accumulates these encounters, starting to reveal real hotspots.
This lets us demonstrate the full human → AI agent → physical robot loop end-to-end today, while the hardware stayed focused on proving the hardest part first: actually seeing, chasing, and capturing a target in the physical world.
Challenges We Faced
We came in wanting to build with the components and time we actually had, not design around ideal hardware — which created real constraints. Our motors, servo, motor driver, and ESP32 were enough for an MVP, but not always enough for the precise mechanical behavior we imagined.
Getting the capture mechanism to behave consistently was one of our biggest challenges: ESP32 power-delivery issues, motor-control problems, and conflicts between analogWrite() motor control and the PWM resources used by ESP32Servo. Even after each piece worked individually, making everything work together was its own problem — a robot that can detect a cockroach isn't necessarily one that can catch it. Small changes in camera position, motor speed, detection thresholds, timing, and servo behavior could completely change whether the robot approached its target or drove past it.
What We Learned
Building a physical AI system is fundamentally different from building software alone. A vision model can output the right bounding box while the robot still fails because a motor turns slightly too fast. A control algorithm can be logically correct while an unstable power supply makes the hardware behave unpredictably. A servo can work perfectly alone and fail once it shares resources with the rest of the system.
We learned to think of the project as one complete feedback system, not separate AI, software, and hardware pieces. Most of all, we learned how much engineering hides inside the word "autonomous" — it's not just detecting an object, it's repeatedly perceiving, deciding, acting, observing the result, and correcting in an imperfect physical environment, whether that's a motor controller or an agent deciding where to send the robot next.
What's Next
We deliberately prioritized a hardware MVP: find the target, chase it, demonstrate physical capture. Our simulation already shows the fuller product vision; the next steps are mostly about bringing that vision onto the physical robot.
Bringing human + machine collaboration to the physical robot. Our simulation already lets a user report a sighting and watch the robot navigate, search, track, and capture, saving the last-known location and a frame if the target escapes. The next step is wiring the same agent into the physical robot's ESP32 loop, so a real report sent through ASI:One dispatches the real hardware.
Smarter navigation with LiDAR. Right now the robot mainly reacts to what it sees. LiDAR and proximity sensors could let it map its surroundings, navigate around furniture and walls, avoid collisions, recognize hazards like stairs, and — combined with the last-known-location system and heatmap we already built — search a room systematically instead of just chasing whatever's in front of the camera.
Infrared and low-light detection. Cockroaches are often active at night, exactly when an RGB camera is weakest. We'd like to experiment with infrared or other low-light sensing, designed with strict safety constraints around people, pets, and furniture.
Beyond cockroaches. At its core, Roach Hunter is a general loop: detect → identify → track → navigate → intercept → act. Changing the perception model, hardware, and capture mechanism could target other pests, like mice, or inspire larger purpose-built systems for other animals — each with its own safety requirements, not just a scaled-up version of this robot.
Ultimately, we imagine Roach Hunter evolving from a robot that catches a cockroach into an autonomous system that understands its environment, remembers where pests were seen, safely navigates a home, and acts when humans aren't around.
Log in or sign up for Devpost to join the conversation.