Inspiration

What it does# Model B-Mix: Robust AI-Generated Image Detection

Inspiration

AI-generated images are becoming increasingly realistic, inexpensive and easy to distribute. However, images rarely remain unchanged after being uploaded online. They may be compressed, resized, cropped, blurred, colour-adjusted or screenshotted before reaching another user.

These transformations can damage the visual artifacts used by AI-image detectors, causing a model that performs well on clean benchmark data to fail in realistic conditions.

We therefore wanted to answer a harder question:

Can an AI-image detector retain useful performance after realistic transformations and when tested on images from a different dataset?

This led us to develop Model B-Mix, a transformation-aware, mixed-domain AI-image detector built using ResNet50.

What It Does

Model B-Mix performs image-level binary classification. Given an uploaded image, it estimates whether the image is:

  • AI-generated
  • An authentic photograph

Instead of returning only a label, the system displays:

  • AI-generated probability
  • Authentic-image probability
  • Final prediction
  • A warning when the result is close to the decision threshold

The probabilities are presented as estimates rather than definitive proof of an image's origin.

How We Built It

Stage 1: Establishing a Baseline

We began with an ImageNet-pretrained ResNet50 and replaced its final classification layer with two outputs:

  • Output index 0: AI-generated
  • Output index 1: Real

We trained the initial model using CIFAKE, which contains 120,000 balanced real and synthetic images.

We combined the original CIFAKE folders and created a stratified 70/15/15 split:

Split Real AI-generated Total
Training 42,000 42,000 84,000
Validation 9,000 9,000 18,000
Test 9,000 9,000 18,000

The baseline achieved a clean validation ROC AUC of 0.9955. However, its performance deteriorated severely under Gaussian noise:

Condition Baseline ROC AUC
Clean 0.9955
Noise (\sigma=0.02) 0.8821
Noise (\sigma=0.05) 0.6642
Noise (\sigma=0.10) 0.5434

This demonstrated that excellent clean performance did not guarantee real-world robustness.

Stage 2: Training for Transformations

We introduced realistic augmentations during training:

  • Gaussian noise
  • JPEG compression
  • Gaussian blur
  • Downscaling followed by upscaling
  • Centre cropping
  • Colour adjustments
  • Random resized cropping
  • Horizontal flipping

After robustness training, ROC AUC under Gaussian noise (\sigma=0.10) improved from 0.5434 to 0.9919, while clean AUC decreased by only 0.0022.

This showed that targeted data augmentation could substantially improve robustness without sacrificing much clean performance.

Stage 3: Mixed-Domain Rehearsal

We next evaluated the model using the higher-resolution saberzl/SID_Set_description dataset.

The CIFAKE-trained model achieved a SID AUC of 0.8114, but its fixed-threshold accuracy was only 0.505 because it classified nearly every SID image as AI-generated. This revealed a major domain-shift and calibration problem.

Fine-tuning exclusively on SID improved SID performance but reduced CIFAKE performance, demonstrating catastrophic forgetting.

To address this, we created Model B-Mix using mixed-domain rehearsal:

  • 900 real SID images
  • 900 fully synthetic SID images
  • 900 real CIFAKE replay images
  • 900 synthetic CIFAKE replay images

The final fine-tuning set therefore contained:

Source Real AI-generated Total
SID_Set_description 900 900 1,800
CIFAKE replay 900 900 1,800
Combined 1,800 1,800 3,600

The SID:CIFAKE sampling ratio was 1:1, and the overall real-to-fake ratio was also 1:1.

Model B-Mix was fine-tuned for three additional epochs with a batch size of 64. We selected the best checkpoint using the mean of SID and CIFAKE validation AUC:

$$

\text{Selection Score}

\frac{ AUC_{\text{SID}} + AUC_{\text{CIFAKE}} }{2} $$

This prevented CIFAKE's much larger validation set from dominating model selection.

Results

Model B-Mix achieved:

Dataset Clean ROC AUC
CIFAKE 0.9912
SID_Set 0.9057

We also evaluated the model under seven severe SID transformations:

Transformation ROC AUC
Gaussian noise (\sigma=0.10) 0.8310
JPEG quality 30 0.8635
Gaussian blur (\sigma=2.0) 0.8540
Downscale 0.25x and upscale 0.8497
Centre crop retaining 80% 0.8906
Colour intensity -20% 0.8719
Colour intensity +20% 0.9005

The mean AUC across these transformations was approximately 0.8659.

Using equal weighting for clean and transformed performance:

$$

S

0.5 \times AUC_{\text{clean}} + 0.5 \times AUC_{\text{robust}} $$

$$

S

0.5(0.9057) + 0.5(0.8659) \approx 0.8858 $$

Challenges We Faced

Clean Performance Was Misleading

Our first model appeared extremely strong because it achieved an AUC above 0.99 on clean CIFAKE data. Robustness testing revealed that it depended on fragile signals that Gaussian noise could destroy.

Transformation Severity Depended on Resolution

CIFAKE images have a native resolution of only 32x32 pixels. Applying blur before resizing caused AUC to fall to 0.6085, while applying the same blur after resizing to 224x224 produced an AUC of 0.9957.

This taught us that transformation order and image resolution must be defined carefully for a reproducible evaluation.

Cross-Dataset Generalization Was Difficult

A detector can learn shortcuts associated with a particular dataset instead of learning universal evidence of AI generation. SID_Set exposed generalization and calibration problems that were invisible during CIFAKE evaluation.

Fine-Tuning Caused Catastrophic Forgetting

Fine-tuning only on SID improved SID performance but reduced CIFAKE performance. Mixed-domain rehearsal allowed us to improve SID AUC while retaining a CIFAKE AUC above 0.99.

Consistency Did Not Guarantee Correctness

We tested transformation-consensus inference by classifying seven transformed views of each image. The views agreed 90.48% of the time, but the final performance did not improve.

This showed that consistency measures stability, not correctness. A model can make the same incorrect prediction across every transformation.

Limited Compute

We developed the project using a free Google Colab GPU runtime. Runtime resets repeatedly removed variables, extracted datasets and model state. Reproducible notebook cells and frequent checkpoint downloads became essential to completing the project.

What We Learned

The most important lesson was that AI-image detection is not solved by achieving a high score on one clean dataset.

We learned that:

  • ROC AUC and fixed-threshold accuracy measure different behaviours.
  • Strong clean performance can hide transformation weaknesses.
  • Data diversity can matter more than architectural complexity.
  • Robust augmentation should simulate realistic redistribution pipelines.
  • Cross-dataset evaluation is necessary to expose dataset shortcuts.
  • Threshold adjustment cannot repair poor ranking performance.
  • High confidence does not guarantee a correct prediction.
  • A responsible detector should communicate uncertainty.

Limitations

Model B-Mix is not a universal AI-image detector. Its performance can vary across generator families, image sources, resolutions and processing histories.

A small manual evaluation using five real and four newly generated images revealed two false positives and two false negatives. Because this test contained only nine images, we used it as qualitative error analysis rather than a formal benchmark.

This result highlighted an important distinction:

Transformation robustness does not automatically provide unseen-generator generalization.

What's Next

Given more time, we would:

  • Train using a wider range of generator families
  • Reserve complete generators for unseen-generator evaluation
  • Add more independent authentic-image sources
  • Calibrate probabilities using a separate validation set
  • Return an uncertain result for borderline predictions
  • Evaluate screenshots and combined transformations
  • Investigate a complementary semantic model branch
  • Build a larger generator-labelled real-world test suite

Model B-Mix is therefore both a working detector and an investigation into what robust AI-image detection actually requires: transformation-aware training, cross-domain evaluation and honest uncertainty reporting.

Built With

  • Python
  • PyTorch
  • torchvision
  • ResNet50
  • Hugging Face Datasets
  • scikit-learn
  • pandas
  • NumPy
  • Pillow
  • Google Colab
  • Jupyter
  • VS Code
  • GitHub
  • Gradio

The final detector runs locally and does not require a paid external inference API.

Datasets and References

Built With

Share this project:

Updates

Submission history