Inspiration
Every time we go grocery shopping, we look at the price tag, but we rarely look at the planet tag. Food production is responsible for over a quarter of global greenhouse gas emissions, yet as consumers, we have absolutely no idea what the carbon footprint of our weekly grocery haul actually is. We realized that we can't reduce what we can't measure. We were inspired to build a tool that brings environmental transparency directly to the checkout line, gamifying sustainability and empowering everyone to make greener choices, one grocery trip at a time.
What it does
Carbon Receipt Scanner is a web application that instantly turns your grocery receipt into an actionable environmental scorecard.
Users can either upload an image of their grocery receipt or use their live web camera to snap a picture on the go. The app instantly extracts the food items, ignores irrelevant data (like taxes and non-food items), and calculates the estimated carbon footprint in grams of CO2 equivalent (g CO2e).
We don't just give users a boring number—we translate that data into relatable metrics:
- 🚗 Miles driven equivalent
- 🌳 Trees needed to plant to offset the impact
- 💸 Micro-transaction simulator showing exactly how many pennies it would cost to neutralize the impact via a certified carbon offset fund.
- 🏆 A Gamified Grade from A to F based on the sustainability of the shopping trip.
Finally, a History Dashboard persistently tracks your scans, giving you a visual breakdown of your footprint over time.
How we built it
We built the application using a full-stack architecture optimized for speed and reliability:
- Frontend: Built with React (via CDN), HTML5, and Vanilla CSS3. We utilized the WebRTC and MediaDevices APIs to build a live, multi-device compatible camera scanner directly inside the browser, complete with local storage for persistent history tracking.
- Backend: Powered by Python and Flask, providing a lightweight, asynchronous REST API to securely route images, handle file compression, and proxy requests.
- AI & Machine Learning: We integrated the Google Gemini Vision API (Flash) to perform advanced multimodal Optical Character Recognition (OCR) and semantic analysis on the receipts. We utilized highly-tuned zero-shot prompt engineering to prevent hallucinations on blurry webcam captures and enforce structured JSON data extraction.
- DevOps: We implemented a robust Continuous Integration (CI) pipeline using GitHub Actions and
pytestto automatically test our backend logic on every code push.
Challenges we ran into
Integrating live web cameras across different hardware proved surprisingly difficult. We initially struggled with OverconstrainedErrors because different laptops and tablets have drastically varying camera resolutions and driver limitations. We overcame this by building a dynamic constraint fallback system in JavaScript that scales down from HD resolutions to basic video feeds to ensure compatibility on any device.
Additionally, webcam captures of laptop screens or crinkled paper receipts are notoriously blurry and prone to glare. Our initial AI tests resulted in heavy hallucinations—the model would guess random items when it couldn't read the text. We solved this by rigorously engineering the backend prompt to enforce strict output rules, ensuring the AI would cleanly reject illegible images rather than hallucinating fake data.
Accomplishments that we're proud of
- Zero-Friction UX: Getting the live camera feed to work flawlessly inside the browser without requiring users to download a native app.
- Fail-Safe Architecture: We built an automated "Demo Mode" fallback. If we hit an API rate limit during a live pitch, the app automatically switches to generating beautiful mock data instead of crashing, ensuring the presentation is always flawless.
- Instant Feedback Loop: Translating abstract carbon data (like "16,000g of CO2") into relatable, understandable metrics (like "Cost to offset: $0.32") fundamentally changes how users perceive their impact.
What we learned
We learned a massive amount about hardware integration in the browser (specifically the intricacies of the navigator.mediaDevices API). We also learned how to tame generative AI. Getting a multimodal model to read text is easy; getting it to return perfectly structured, deterministic JSON data under sub-optimal image conditions requires significant prompt engineering and architectural foresight.
What's next for Carbon Receipt Scanner
In the future, we plan to implement a real payment gateway (like Stripe) to allow users to actually execute the micro-transactions and donate those pennies directly to certified carbon offset funds. We also want to expand the AI's knowledge base to suggest greener alternative products (e.g., "Next time, try oat milk instead of almond milk to save 400g of CO2!").
Built With
- artificial-intelligence
- ci-cd
- css3
- flask
- gemini-api
- github-actions
- google-gemini
- html5
- javascript
- machine-learning
- pytest
- python
- react
- webrtc
Log in or sign up for Devpost to join the conversation.