Inspiration

When we sat down to brainstorm for this hackathon, we wanted to innovate and think out of the box, trying to think of solutions that could bring a real world impact. We kept coming back to a glaring vulnerability in the digital age: how easily everyday documents can be forged. With generative AI tools now widely accessible, altering a medical certificate, faking a graduation diploma, modifying a signed contract, or forging receipts for false claims is trivial and often undetectable to the naked eye. We realized that existing detectors only give a vague "AI generated" score for a whole image, which isn't helpful for documents where only a single date, name, or signature might be tampered with. We wanted to spend this hackathon building a concrete, technical solution that tackles this head-on, a tool that doesn't just guess if a document is fake, but proves exactly where it was altered.

What it does

DocDet is an AI forensic pipeline designed to confidently discern between authentic documents, human-altered forgeries, and entirely AI-generated fakes.

We recognized a fundamental flaw in standard image detectors: a tampered field (like a forged signature or altered total) often makes up less than 1% of a document. Downscaling a massive invoice to fit a standard model destroys the microscopic pixel evidence needed to catch the forgery. Instead of downscaling the whole image, DocDet acts as a digital magnifying glass. It scans documents using overlapping, high-resolution tiles, analyzing the microscopic gradient and compression artifacts left behind by generative AI or copy-move edits. It then aggregates these localized clues into a single, highly accurate page-level confidence score.

How we built it

We hand-coded and engineered DocDet entirely from scratch using the command line. Our stack relied on zsh and Python 3.12, running local compute on an Apple M1 via the PyTorch MPS backend.

The Architecture: We paired an ImageNet-pretrained ResNet18 visual backbone with a custom, differentiable ArtifactFeatures head. Instead of frequency transforms, this operates in the spatial domain, extracting 18 precise features—including grayscale residuals after 3×3 average pooling, gradient statistics, and per-channel means and standard deviations.

Training & Loss: The model acts as a binary classifier trained with BCE-with-logits, paired with a smooth-L1 consistency penalty between two independently augmented views (forcing robustness against real-world JPEG compression, blur, resize, and color jitter). We ensured class balance through paired positive/negative sampling. Masks were used strictly to place training crops, never as a prediction target.

Tile-Based Inference: To preserve forensic details, inference scans overlapping tiles (with a 25% overlap to avoid destroying seam evidence) at three page-relative scales. Tiles are floored at 160px and capped at 512px, then resized to our model's 224×224 input.

Page-Level Aggregation: We don't draw bounding boxes or reassemble images. Instead, a scikit-learn Logistic Regression model learns a single page verdict from the distribution of the tile scores using order statistics. The maximum score catches localized edits, while upper quantiles flag wholly generated pages.

The Interface: We built a lightweight web UI using Flask and server-rendered HTML to handle local document uploads and processing.

Challenges we ran into

Choosing to hand-code and operate entirely without a GUI IDE (like VSCode or Jupyter) pushed us to strictly manage our pipeline via CLI, which made debugging tensor shapes and aggregation logic challenging. Furthermore, training robustly on Apple Silicon (without CUDA) required careful optimization of our PyTorch workflows.

However, our biggest dataset challenge was the lack of reliable "clean" baselines. We utilized datasets from the Hugging Face Hub (Scam-AI/AIForge-Doc-v1, -v2, gpt4o-receipt), where the fakes were already generated upstream by state-of-the-art models (GPT-4o, Gemini Flash, Ideogram). To properly train our detector to spot anomalies, we desperately needed the original, untampered source pages to measure our false-positive rate against, but the project initially believed those authentic sources were lost or unavailable.

Accomplishments that we're proud of

Our proudest achievement was solving the missing baseline problem. We successfully recovered 1,973 authentic source pages that were thought to be lost by executing exact, outside-mask pixel matching against the tampered variants. This massive data-recovery effort was the sole reason we were able to measure a true page-level false-positive rate and calibrate our validation threshold to hold a strict 5% FPR.

We are also incredibly proud of the tile aggregation math. Moving away from standard image downscaling and successfully teaching a logistic regression model to interpret the statistical distribution of localized tile scores—proving that the maximum score accurately flags localized edits while upper quantiles catch whole-page fakes—was a major technical breakthrough for us.

What we learned

Building DocDet was a masterclass in spatial-domain analysis. We learned that the secret to catching modern generative AI isn't necessarily in complex frequency transforms, but in analyzing the micro-statistics of grayscale residuals and gradients. We also gained a deep understanding of order statistics, realizing that the distribution of localized scores tells a much more accurate story than attempting to spatially reassemble a heatmap. Finally, we proved to ourselves that we could hand-build and orchestrate a complex, mathematically dense machine learning pipeline entirely from a local command-line environment.

What's next for DetDoc

Our immediate next step is to update our Flask web interface to dynamically expose the page-level tile metrics (max, mean, std, quantiles) to the user, providing a more explainable breakdown of why the document was flagged. Moving forward, we want to expand our training data beyond financial receipts to include high-stakes legal and official documents, ultimately packaging DocDet as a lightweight, highly robust API for automated institutional verification.

Built With

+ 9 more
Share this project:

Updates

Submission history