Inspiration

We get frustrated losing things in our world, and that "where was this last seen" is not something our phones excel at

AI loves chatting and searching the web, but is comparatively bad at remembering our physical world. Smart glasses and phones know what we see, and little else gets that stream into a queriable, accessible format for later questioning

which is why we built

SIGHTLINE - an AI memory layer for the physical world

See through the camera, remember what was important. Ask about a specific item, or event, and when it happened. And when things go wrong, see this info surfaced in a calm, accessible manner to aid in judgement.

We currently have SIGHTLINE set up to take input from either your iOS device or your META Ray-Ban glasses!

We also wanted to let people at the hackathon know what SIGHTLINE is about quickly, so we built a fun 45 second personality quiz called SIGHTLINE Sidekick which uses HokieAI.

What is it

SIGHTLINE consists of three modes:

Live

Use your phone's camera over Sightline Link to Mission Control, where we do lightweight detection, connection state, and memory building

Recall

The conversational mode. Ask "Where is my black laptop?" or "What story did I talk to David about?". Rewind through the visual memory to the relevant moment, highlight it, and answer based on the location and time. Conversation questions are summarized with Gemini across the relevant parts of the conversation transcript, rather than verbatim quotes

Guardian

The physical world safety observer, not a judgement system but an observation layer for humans:

  • fall detection (motion freefall → impact)

  • vehicle crash detection (speed >25mph → rapid deceleration)

  • possible distress observation (motion + audio triggers)

  • calm incident cards with confidence value, time, and location

  • all DEMO incident cards do not make real-world calls

Guardian Mode is a key product focus, as the same physical memory stack can observe both the story you want to tell, and the situation needing a helping hand nearby

The other modes:

  • Voice: speech in, spoken confirmation or alerts with ElevenLabs, not readback

  • Sidekick: ~45 second personality quiz → shareable SIGHTLINE Sidekick profile

Try it:

website: https://sightline.surf · sidekick: https://sidekick.sightline.surf

How we built it

Layer Stack
SIGHTLINE Link Swift / iOS — camera, CoreMotion, GPS, WebSocket relay
Mission Control TypeScript, Express, React — Live / Recall / Guardian
Vision & memory COCO-SSD, LocateAnything, Gemini, MongoDB Atlas
Voice On-device / Scribe STT · ElevenLabs TTS
Guardian Fall + crash detectors (speed → Δv / motion → impact) · observation cards
Sidekick Next.js + HokieAI (Gemini / local fallback) on Vercel

We reworked Mission Control away from a debugging console - the focus is on having a compact status view, having Recall as the wow factor, and putting the browser-based Demo simulations under Demo. Guardian is its own safety product surface.

Challenges we faced

  1. Venue: getting the phone ↔ laptop connection to work reliably over the venue's secured Wi-Fi. We used personal hotspots, Bonjour, QR deeplinks, and a public facing landing page (which opened the local session when it was up)

  2. Black memory thumbnails - the iPhone JPEG EXIF rotation was interfering with the image display until we fixed it based on image dimensions after rotation

  3. Smart conversation recall, e.g. asking about a "story told to David" rather than a weak keyword. We rank names and topics detected before using Gemini summaries

  4. Browser audio: getting TTS to work reliably. It worked well server side, but in the browser, we had to do HTTP mp3 streaming + unlock AudioContext instead of big base64 WebSockets

  5. Two public facing websites: static sightline landing on GitHub Pages (sightline.surf) vs sidekick's Next.js on Vercel (sidekick.sightline.surf)

  6. Guardian: Not being scary, but being an observation layer for human judgement. Having the DEMO badge, and reserving red for real/simulated danger

What we learned

Physical-world computer vision is fiddly with regards to orientation, autoplay, CGNAT, and model IDs, and doesn't play nice with CORS-to-localhost. Building this has taught us that the wow is not in another dashboard, but in a single unified chain of seeing → remembering → asking → rewinding back to the moment.

As well as a similar chain for sensing → observing → alerting → letting humans decide what to do next

That's Live, Recall, and Guardian - the memory and judgement layers for the physical world, and the world we move through every day.


Built at VTHacks 14 · Egoists

Built With

Share this project:

Updates

Submission history