Inspiration

One of the biggest issues I faced when working on my schoolwork was finding stuff. Sometimes teachers would make references to documents that I got two weeks ago or four months ago, and I simply could not find them. It was so frustrating, but I had no solution.

Then, when I was sitting there thinking about what to make for this hackathon, I came up with a really simple idea: what if there was an app that took care of that for you?

The problem with all scanner apps right now is that they're too much of a hassle to use. You have to manually take photo by photo, adjust all the points so it looks right, and sort them into folders. It's too much work for the off chance that you need a document a few months down the road.

However, it didn't need to be that way. What if you could simply take pictures of all your assignments and then forget about them? What if the app handled all that work, using AI to figure out what subject it was and what it was about, and automatically filed it away so whenever you need it later, you could simply pull it up without any hassle? That's how I created Snappy.

What it does

Snappy is a really frictionless, easy-to-use scanner app. All you have to do is open the app, press the camera icon, and snap a picture of whatever handout you are trying to save. There is no waiting for it to auto-detect where the page is or trying to manually correct keystones: simply press the shutter button, and it is done.

After you take the picture, it is sent back to Snappy's server, where an edge detection model automatically figures out where the page is and uses that to create an actual scan from the picture. That is then fed into an OCR model, which reads the text off the page, and sent to an LLM, which parses that text and categorizes the document.

This all happens in the background asynchronously on the server, so you are never sitting there waiting for anything. From your perspective, you simply snap the picture and forget about it; Snappy handles the rest.

How we built it

For Snappy to work, there are many moving parts that need to come together in conjunction.

The core of Snappy is the server, which is written in Python using FastAPI. Images are stored in an SQLite database on the server. When you snap an image, that image is processed by OpenCV, which uses a variety of edge detection techniques to figure out where the document is in the picture and isolate it. Then we adjust brightness, denoise the image, and make a few more adjustments so that the image itself looks clean and could even be reprinted later on to create a copy of the document if necessary. From there, it is fed into our OCR model, which extracts the text from the picture. We then use the Gemini API, along with the text extracted from the picture, to categorize the document into the correct categories.

On the front end, we have two platforms: a web front end created using SvelteKit and TypeScript, and an iOS front end created using SwiftUI so that it feels native.

Both front ends are functionally identical. However, the mobile front end is the preferred one, as it has the built-in camera feature that allows users to scan documents quickly.

One big thing that I focused on when creating Snappy was to make the user experience as frictionless as possible, which involved offloading most of the work to the server side. Therefore, after you upload an image, take a picture, or do any action in Snappy, you never have to wait for anything. All the processing and loading happens in the background on the server and updates once completed, so you don't have to sit around and deal with anything.

The idea is that the less time you spend in the app the better, because the more likely it is that you'll use Snappy to save all your documents. I started using it myself after creating it, and it's really convenient since I no longer have to worry about losing any of these sheets: I can simply snap a picture and forget about them.

Challenges we ran into

The biggest challenge with Snappy was handling document scanning. Even in 2026, scanning documents is really hit or miss: bad lighting and unsteady, less-than-ideal images can make it really difficult to turn a rough picture into a usable scan.

I spent a lot of time adjusting the models to hit a good balance between accuracy and reliability. While there were many more complex models I could have used, I ended up settling on a four-point system simply because it is user-adjustable. Most of the documents scanned in Snappy will never be looked at again anyway, so as long as the scanner works about 80% of the time, we will be solid.

Regardless of how the scan turns out, the OCR consistently identifies the text and categorizes the document; I made sure that part was as accurate and reliable as possible. On the rare chance that a user needs a document that was not scanned properly, they can simply hit the adjust button and fix the scan themselves before using it. In practice, however, the scan itself is fairly accurate and reliable, so I do not expect users to have to do this too often.

Accomplishments that we're proud of

The thing I'm most proud of in Snappy is the user experience. I think the UX is really well done: you only have to spend a few seconds in the app to get your scans in, and if you need to find them later on, it's incredibly easy. I'm really happy with how the scanning experience especially turned out, and how basically everything is asynchronous and handled by the server. It just makes the app super easy to use; I never have to actually go in and fix a bunch of things myself. I can just take the pictures and forget about them.

What we learned

I learned a ton about image processing and how to build server-centric user experiences.

In the past, I had only really made projects that were heavily client-side, mainly because I didn't have the infrastructure to host a server. While I don't have a server to host Snappy on for this project either, I decided to try something new and build something heavily server-based.

This came with a lot of challenges, such as:

  • How to keep the client updated on the status of server-side processes.
  • How to handle server-side issues gracefully.
  • How to prevent tasks from blocking the server from handling other requests.

I learned a lot, and overall, building Snappy was a really fun experience.

What's next for Snappy

Because I only have very limited resources, especially on the AI compute front, I wasn't able to make Snappy as smart as I could have.

In the future, I would like to use AI to make the process of finding, retrieving, and using information from documents even easier. For example:

  • Using AI to ask questions about any document stored in Snappy and getting a response about it.
  • Using AI to find specific items, such as due dates on a syllabus, and surfacing them for the user.

I think there's a lot of potential in this kind of interface for AI, as it works best when you give it a lot of sources and context. By scanning those documents, Snappy gets a ton of context, which means there's so much more the AI can do with knowledge of all those things to make things easier for students.

Built With

Share this project:

Updates

Submission history