Inspiration

AI-generated images are becoming increasingly realistic, but many detectors lose reliability after ordinary actions such as compression, resizing, cropping, blurring, or recoloring. We created ProofLens to explore whether AI-image detection can remain dependable under these realistic redistribution scenarios.

What it does

ProofLens analyzes an uploaded image exactly as received and estimates whether it is authentic or AI-generated. It presents calibrated probabilities, a clear decision, confidence information, and visible processing signals such as compression or blur.

Processing detection is explicitly presented as a heuristic estimate—not proof of an image’s editing history.

How we built it

We built ProofLens using a DINOv2-based binary image classifier with layer normalization and a classification head. Our staged experiments included head-only training, fine-tuning the final transformer blocks, transformed-image training, consistency objectives, and hard-transformation mining.

We combined SID Set data with generator-labelled AIGenImages2026 samples. Grouped splitting kept related images and duplicate clusters together, while generator families were held out to measure generalisation.

Training and evaluation covered JPEG compression, Gaussian blur, resizing, noise, color adjustment, and center cropping. Additional stress tests used WebP compression and screenshot-style redistribution.

The selected model was calibrated using validation data, exported to ONNX, and verified against PyTorch using a required 32-image numerical-parity test. A Gradio interface provides local, calibrated inference.

Challenges we ran into

Downloading and processing the large SID dataset reliably was difficult, so we created a resumable shard-by-shard acquisition process.

Preventing data leakage was another major challenge. Related images had to remain in the same split, and generator families needed separate holdout partitions.

ONNX export initially stopped because PyTorch tracing warnings were treated as fatal errors. We made the detector’s input validation tracing-safe and corrected the export process.

We also redesigned the interface after realizing that users should submit images that may already be transformed. The final demo therefore analyzes uploads directly instead of asking users to apply transformations.

Accomplishments that we're proud of

Our selected model achieved:

  • 0.9974 clean ROC AUC
  • 0.9971 macro robust ROC AUC
  • 0.9966 worst-condition ROC AUC
  • 0.9638 unseen-generator ROC AUC
  • 98.41% accuracy at the calibrated threshold
  • 96.23% recall

The model maintained ROC AUC above 0.9964 during separate WebP and screenshot stress tests.

We also completed validation-only model selection, calibration, held-out testing, ONNX export, numerical parity verification, automated testing, and a working local demonstration.

What we learned

We learned that clean-image accuracy alone is not enough. Robustness must be evaluated separately across transformation families, and checkpoint selection must remain independent of the test set.

We also learned that detecting AI generation and detecting image processing are different problems. Compression and blur may leave measurable evidence, while cropping, recoloring, and rescaling can be impossible to identify confidently from one final image.

Finally, reproducibility requires recording dataset revisions, grouped splits, hashes, configurations, calibration values, preprocessing versions, and model-export verification—not merely saving a checkpoint.

What's next for ProofLens

We plan to evaluate more cameras, platforms, generators, and compound transformation sequences. We also want to improve calibration under domain shift, reduce CPU inference time, strengthen processing-signal analysis, and provide clearer uncertainty and false-positive explanations.

ProofLens remains a research prototype rather than forensic proof. Its results should support human review and should never independently determine authorship, moderation, access, or penalties.

Built With

Share this project:

Updates

Submission history