Inspiration
We get frustrated losing things in our world, and that "where was this last seen" is not something our phones excel at
AI loves chatting and searching the web, but is comparatively bad at remembering our physical world. Smart glasses and phones know what we see, and little else gets that stream into a queriable, accessible format for later questioning
which is why we built
SIGHTLINE - an AI memory layer for the physical world
See through the camera, remember what was important. Ask about a specific item, or event, and when it happened. And when things go wrong, see this info surfaced in a calm, accessible manner to aid in judgement.
We currently have SIGHTLINE set up to take input from either your iOS device or your META Ray-Ban glasses!
We also wanted to let people at the hackathon know what SIGHTLINE is about quickly, so we built a fun 45 second personality quiz called SIGHTLINE Sidekick which uses HokieAI.
What is it
SIGHTLINE consists of three modes:
Live
Use your phone's camera over Sightline Link to Mission Control, where we do lightweight detection, connection state, and memory building
Recall
The conversational mode. Ask "Where is my black laptop?" or "What story did I talk to David about?". Rewind through the visual memory to the relevant moment, highlight it, and answer based on the location and time. Conversation questions are summarized with Gemini across the relevant parts of the conversation transcript, rather than verbatim quotes
Guardian
The physical world safety observer, not a judgement system but an observation layer for humans:
fall detection (motion freefall → impact)
vehicle crash detection (speed >25mph → rapid deceleration)
possible distress observation (motion + audio triggers)
calm incident cards with confidence value, time, and location
all DEMO incident cards do not make real-world calls
Guardian Mode is a key product focus, as the same physical memory stack can observe both the story you want to tell, and the situation needing a helping hand nearby
The other modes:
Voice: speech in, spoken confirmation or alerts with ElevenLabs, not readback
Sidekick: ~45 second personality quiz → shareable SIGHTLINE Sidekick profile
Try it:
website: https://sightline.surf · sidekick: https://sidekick.sightline.surf
How we built it
| Layer | Stack |
|---|---|
| SIGHTLINE Link | Swift / iOS — camera, CoreMotion, GPS, WebSocket relay |
| Mission Control | TypeScript, Express, React — Live / Recall / Guardian |
| Vision & memory | COCO-SSD, LocateAnything, Gemini, MongoDB Atlas |
| Voice | On-device / Scribe STT · ElevenLabs TTS |
| Guardian | Fall + crash detectors (speed → Δv / motion → impact) · observation cards |
| Sidekick | Next.js + HokieAI (Gemini / local fallback) on Vercel |
We reworked Mission Control away from a debugging console - the focus is on having a compact status view, having Recall as the wow factor, and putting the browser-based Demo simulations under Demo. Guardian is its own safety product surface.
Challenges we faced
Venue: getting the phone ↔ laptop connection to work reliably over the venue's secured Wi-Fi. We used personal hotspots, Bonjour, QR deeplinks, and a public facing landing page (which opened the local session when it was up)
Black memory thumbnails - the iPhone JPEG EXIF rotation was interfering with the image display until we fixed it based on image dimensions after rotation
Smart conversation recall, e.g. asking about a "story told to David" rather than a weak keyword. We rank names and topics detected before using Gemini summaries
Browser audio: getting TTS to work reliably. It worked well server side, but in the browser, we had to do HTTP mp3 streaming + unlock AudioContext instead of big base64 WebSockets
Two public facing websites: static sightline landing on GitHub Pages (sightline.surf) vs sidekick's Next.js on Vercel (sidekick.sightline.surf)
Guardian: Not being scary, but being an observation layer for human judgement. Having the DEMO badge, and reserving red for real/simulated danger
What we learned
Physical-world computer vision is fiddly with regards to orientation, autoplay, CGNAT, and model IDs, and doesn't play nice with CORS-to-localhost. Building this has taught us that the wow is not in another dashboard, but in a single unified chain of seeing → remembering → asking → rewinding back to the moment.
As well as a similar chain for sensing → observing → alerting → letting humans decide what to do next
That's Live, Recall, and Guardian - the memory and judgement layers for the physical world, and the world we move through every day.
Built at VTHacks 14 · Egoists


Log in or sign up for Devpost to join the conversation.