We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

We wanted to make a system that would allow people to explore places that they wouldn't want to physically access themselves. For example, news reporters may not want to go into dangerous places. We were also exploring the limits of transmitting a video feed into a backend that could use it for different purposes. In our case we stream the video from the GoPro, and at the same time sample JPEG frames to analyze them with Gemini. It then explains through ElevenLabs what it inferred.

How we built it

FFmpeg pulls the GoPro's UDP stream and copies it, no re-encode, into MediaMTX running in Docker for retransmission to multiple parties: the browser takes WebRTC off it for live video. Separately, a separate sampler pulls JPEG frames off the same server and sends them to Gemini, whose text goes to the browser over Server-Sent Events and to ElevenLabs for audio. Backend is Python with no web framework, just the standard library HTTP server and asyncio. Frontend is React and Vite.

Challenges we ran into

Being able to stream video from our GoPro was the first challenge. From then on, we found hard to give meaning to our project. We acknowledge the potential behind this prototype, but didn't have a very clear goal in mind. Eventually, we came up with a good focus, which is mentioned in the Youtube Video Demo.

Accomplishments that we're proud.

The live video and the analysis pipeline are fully independent, so a slow or failed AI provider cannot delay or interrupt the picture.

Our eleven selectable personalities can change tone only; the response format and safety rules are applied separately and cannot be overridden, including by a personality the user writes themselves.

What we learned

A camera pointed at the world is a prompt-injection surface. Anything someone holds up in front of the lens lands in the model's context, so we had to explicitly tell it that text in images is scene content, not instructions.

What's next for RC-AI Surveyor

We intend to make the analysis rate adaptive. The model already reports when it has nothing new to describe, and using that signal the system could sample less frequently in static scenes and more frequently when conditions change. At present a stationary vehicle incurs the same cost as an active one.

We would like to add persistent site memory, so that the system can report differences between visits rather than only describing the current view.

Built With

Share this project:

Updates

Submission history