PIP

First-Responder Reconnaissance Rover

Scout the room. Map what is visible. Give responders a clearer way in.

The problem

In a fire or other dangerous indoor emergency, responders may not know the layout beyond the entrance, whether anyone is inside, or which route through the room is passable. Entering without that information can cost time and expose people to additional risk.

PIP explores how a small rover carrying an iPhone’s LiDAR and camera can gather indoor spatial information before responders commit to entering. It turns the rover’s observations into a shared 3D map, marks candidate person detections, and shows routes from the mission entry point toward detected people. The goal is to give a response team useful orientation and a starting point for planning a search.

PIP supports the responder’s decisions; it does not replace them. The map only describes what the sensors have observed. Unseen areas remain unknown, and a proposed route is not proof that a real emergency scene is safe.

What PIP does

A rover carries an iPhone 17 Pro. The phone streams camera, depth, and pose data to a backend while the rover is driven or explored. The dashboard builds a persistent, colored 3D reconstruction from the received sensor data and shows the rover’s location, its recent path, and the space it has observed.

When the vision system proposes that it sees a person, the dashboard can mark that location and draw a route from the mission’s starting point toward each detected person. Showing separate routes matters: a search team should be able to see multiple reported locations rather than having the interface silently choose one.

The map distinguishes observed space from unknown space. A route can only be reasoned about using the available map and rover-clearance rules; the dashboard should not make an unobserved corridor look verified or safe. Responders remain responsible for assessing conditions and choosing whether to follow a suggested route.

How it works

1. The phone captures spatial observations

The iPhone app uses SwiftUI and ARKit to capture camera images, LiDAR depth, and device pose. Image, depth, and pose must refer to the same ARKit frame so measurements are placed consistently as the phone moves. Frame and map-session identifiers help prevent delayed observations from being applied to a different mapping session.

2. The backend places measurements in the room

The Python and FastAPI backend validates incoming observations and uses camera intrinsics, depth, and pose to place measured points in a shared world coordinate system. Measurements without usable depth can still support a 2D image detection, but should not be presented as a measured 3D location.

3. The dashboard builds the scene

The React, TypeScript, Three.js, and React Three Fiber dashboard renders the accumulated colored reconstruction, rover position, trajectory, detections, and navigation information. The map grows as new sensor data arrives; it is not a complete model of surfaces the phone has not seen.

4. Navigation uses the mapped space

The navigation software represents space as free, occupied, or unknown, accounts for the rover’s footprint, and can plan toward reachable goals. During Explore, it evaluates where to move and scan next as observations change. Obstacle avoidance and scan pacing are active engineering work: the intended behavior is to make useful progress, detour around passable obstacles, hold safely when blocked, and continue gathering map data when conditions permit.

A plan on the dashboard is a recommendation from the current map. Physical movement depends on the rover, phone relay, wireless connection, and embedded controllers all working together.

Why this is useful

A video gives a responder a sequence of views. A spatial map can preserve how those views fit together: where the entrance is, what has been observed, where the rover traveled, and where a person was reported. This can help a team coordinate around a common picture of the space instead of relying only on descriptions from someone looking at a live camera feed.

The central idea is modest and practical: use a small remotely operated or autonomously exploring rover to reduce uncertainty before a person enters an unfamiliar indoor area. The system should show its evidence and its gaps so responders can judge what to do next.

What we have built

PIP connects an iPhone sensing app, a backend for depth processing and navigation, a browser-based 3D dashboard, and a rover control chain. The dashboard supports inspection of the spatial reconstruction and rover state. The project also includes planning and autonomous Explore software, alongside manual control and safety checks in the command path.

Person detection and autonomous driving are still being refined. In particular, detections can be incorrect, and the rover’s ability to make repeatable progress around obstacles has not yet been established through repeatable physical runs. We treat those as open engineering challenges, not proven rescue capabilities.

The project is a prototype developed and tested in ordinary indoor environments. The rover and phone have not been validated for operation in an active fire, smoke, extreme heat, or other hazardous conditions. PIP is not fire-rated equipment, and it must not be represented as ready to enter a burning building. The fire-rescue scenario communicates the problem the project is designed to help with; demonstrating the current prototype safely is a separate matter.

What we learned

A useful map must carry the difference between evidence and assumption. A missing surface is not necessarily open space. A candidate person detection is not a confirmed person location. A route through mapped cells is not a guarantee of safety. A rover that has started moving is not necessarily making useful progress.

Those distinctions shape both the software and the interface. PIP should show where observations came from, preserve the map as the camera turns away, update it when new evidence arrives, and make blocked or uncertain navigation clear to the operator.

Our next milestone is repeatable physical Explore behavior in controlled indoor tests: visible forward progress, appropriate obstacle detours or a clear blocked hold, continued mapping, and reliable operator stop. Better person-detection validation and measured rover and phone mounting geometry are also needed before the map or routes can support stronger claims.

PIP aims to give first responders a clearer picture of an unfamiliar indoor space before they decide how to search it.

Built With

Share this project:

Updates

Submission history