Inspiration
The original idea for this project stems from the backend application I built for my capstone project involving Meta Aria Gen 2 glasses. The streaming pipeline made for the glasses with fixed dimension images was adapted to a batch process for a generalized web application capturing images from any camera. The idea was to help develop an application that helps the users track their belongings based on periodic snapshots of their surroundings, especially useful for Alzheimer's patients. The original project can be found at https://github.com/ericbeier253/VT-Capstone-Summer-2026.
What it does
There are 3 parts: 1) Streaming Server: this is setup on a website to allow any device with a browser to take pictures and store the images on cloud. 2) Batch Enrichment: It processes all new images on the cloud to a local DB after running object detection, crop objects using bounding boxes and store objects as image and text embeddings (vector indices) alongside environment information. 3) Chatbot viewer: A chatbot application that takes in user query and searches the DB for the object's last seen information alongside their images.
How I built it
The streaming server was built using Django framework. The enrichment pipeline uses OpenAI GPT-6 Luna as the Generative model for object detection and metadata generation, DINOv2 for image embedding using GPU and OpenAI embedding model for text embedding generation. Finally, the chatbot viewer was built using Dash and Panel UI libraries for Python. All parts were created using GPT-5.6 Sol model through Codex.
Challenges I ran into
- The security issues with the streaming server were difficult to bypass as camera permission through web application are stringent.
- The frontend for the camera app was the hardest since this was completely new and developed for the first time during the hackathon.
- The bounding boxes had to be dynamic instead of static since the camera can take pictures of any dimensions based on their specs.
Accomplishments that I'm proud of
- The frontend app is quite robust and independent.
- The pipeline is working well even after adapting to new models as they were originally tested for Gemini 3.5 Flash and Gemini embedding 2.
- The Local DB Chroma is working well as the original used Firestore which was cloud based.
What we learned
- The camera permissions are quite stringent requiring SSL certificates and HTTPS end-to-end via a web server,
- All embedding models work well as long as the pipeline is well crafted.
What's next for VectorCompare
- The current project uses batch enrichment which can be upgraded to streaming using event notifications and queues. More robust system design, basically.
- Also, the objects detected can be misidentified using the current parameters. This can be researched further.
Log in or sign up for Devpost to join the conversation.