Inspiration
Search and rescue has not technologically evolved as it should. Sure drones are great! Until you're in a impoverished disaster zone and have to wait days for drones to ship in, or you've been located by a drone but can't receive immediate care. Human responders are always necessary but they bear a huge responsibility. We wanted to catch those critical moments that a human might miss, but a computer wouldn't and in fact, coordinate rescue efforts so seamlessly that we could do it today on stage :)
What it does
Responders scan a QR code and join with Beacon. Each phone becomes a tracked camera, streaming its position, orientation, and frames to a central hub and receiving instructions back.
The operator console shows every live feed, a minimap with every responder, and a heatmap of where a missing person still likely is. At the same time, keyframes reconstruct the landscape in 3D in real-time so within seconds, the console is showing a building nobody had a map of.
Finding a specific person starts with a reference photo, then every frame from every phone is matched against it, again in real-time. How do the responders coordinate? Through a commander model we post-trained ourselves: it analyses all the signals, including what each responder has covered, significant visual cues determined by a fine-tuned ViT, and current positions, to provide instructions. We trained it in a custom simulator, which rehearses an entire search at a specific location, including obstacles like fire, smoke and blocked doors. This also doubles as the RL environment for the commander!
How we built it
We built Beacon's mobile interface through Swift and Expo, using ARKit to process phone positions and orientations in real-time. Each phone streams pose and frames over WebSocket to a FastAPI hub, which serves the operator console and fans work out to GPUs on Baseten.
- Mobile: Swift, React Native, Expo, ARKit
- Hub: Python, FastAPI, WebSocket, Three.js
- Person/object detection: YOLOE + OSNet on a Baseten L4
- 3D reconstruction: Optimized VGGT-Ω running on Baseten H100
- Commander: Qwen3-4B + rank-16 LoRA, RL-trained on Baseten
- Visual memory: fine-tuned CLIP-B/16, trained on Baseten
Why these models (FOR BASETEN)
Qwen3-4B + rank-16 LoRA - rescue commander (trained on Baseten)
This is the model we actually trained, and we want to start off with why a language model at all:
The commander's input includes subjective signals. "I called into E7, no answer" and "I searched E7" are the same place but completely different evidence, and a Bayesian coverage planner cannot tell them apart because the difference is semantic. That's the decision we wanted learned.
Next, why RL and not SFT on examples? We have no expert transcripts to imitate due to the unique and urgent nature of this field, and imitating our own planner would just overfit. So the reward is the outcome (successful recovery - time over the deadline), with reinforce over groups of four rollouts of the same incident, each scored against the mean of the other three and NO imitation labels.
Why it's constrained: the model emits a single masked token selecting 1 of 10 actions. It cannot invent coordinates, assign a busy responder, or declare completion before a physical arrival, enforced as invariants, not behaviours we hope training produces.
The trained model rescues 100% of the people in 404 of 1024 incidents to greedy's 182 with the same average latency. Overall, it completes 47.6% more rescues than greedy across 1024 unseen incidents, winning 473, losing 32. We trained on Baseten Training Jobs with 120 updates, 1,687 runner seconds and served on an L4.
CLIP-B/16 - visual memory (trained on Baseten)
We started with a Qwen3-VL-Reranker-2B cross-encoder plus LoRA and ran it three times. All three failed our first validation gate (76.25% → 76.25%, 77.5% → 80%, 77.5% → 82.5%). The failure was structural since a cross-encoder has to score every candidate against every query, so it can't index a collection that grows every second of a live search.
But a bi-encoder can, where CLIP-B/16 embeds each observation once into a shared 512-dim space with the query text, so retrieval is a vector search over everything every phone has ever seen. Fine-tuning it on 3,072 sources from a mix of our own videos and online datasets increased retrieval accuracy from 86.1% to 92.4%.
VGGT-Ω - live 3D reconstruction
Frames from a moving crowd have no guaranteed overlap and no baseline, but VGGT-Ω predicts cameras and depth jointly in a single feed-forward pass, so we stream preprocessed keyframes to continuously generate the 3D model. These updates take 0.9/1.8/3.6/5.5/8.0 s at 12/24/44/64/88 views on Baseten's H100, with 85.6% of predicted pixels passing consistency tests. Caching depth between updates and checking cross-view agreement at half resolution cut the reconstruction's postprocessing from 14.1 s to 4.4 s.
YOLOE + OSNet - object detection/recognition
We experimented with fine-tuning object detection/image recognition models but it didn't provide an advantage over using industry standard CV models. Both models share one Baseten L4 so a frame doesn't take a second network hop between detection and matching: 98–147 ms warm round trip from a laptop. We cache the active class set between requests to avoid re-initializing the YOLOE prompt and predictor every call.
Challenges we ran into
- Getting and processing accurate position and orientation data from multiple phone connections in real-time
- Optimizing the inference of the 3D reconstruction model (VGGT-Ω) to generate 3D spaces in real-time
- Designing the proper RL environment to train the commander model (how to model a rescue mission)
Accomplishments that we're proud of
- A terrain that maps itself in 3D while you search it in mere seconds (no floor plan, no LiDAR, no setup)
- A trained commander that cut average mission time 14.9% over 200 simulated incidents
- Real-time latency, where each phone's frames go to a cloud GPU and comes back matched in about 130 ms
- It holds up with a crowd with > 30 phones streaming at once (perfect for a Finalist's demo!)
What we learned
- A tuned Bayesian planner is a strong baseline. The learning with RL belongs where the planner is structurally blind like ambiguous reports, not route optimization which it already solves
- Native software has the most support for advanced features like accurate position and orientation tracking
What's next for Swarm Sight
- Building this into an AR interface like smart glasses or a helmet
- Incorporating audio as a signal for decision making
- Gathering better data from historical events, rescue experts, and more realistic environments and simulations to train the models
- Supporting the full platform through edge inference and compute (allowing for no network deployments which are common in disaster relief areas)
- Setting up continuous learning by tracking points of success and failure in specific environments
Built With
- baseten
- expo.io
- fastapi
- python
- pytorch
- qwen
- react-native
- rl
- swift
- typescript
- vggt
Log in or sign up for Devpost to join the conversation.