Inspiration

AI images are getting harder and harder to be spotted. I keep scrolling on TikTok recently and didn't know that the photos were AI generated until I clicked into their profile. Hence, seeing this challenge on question 5, I know immediately what I have to tackle on.

What it does

Our model can differentiate AI generated images from real images, even after augmentations like jpeg compression, resize, blurring etc.

How it addresses the problem statement

  • Robustness to real post-processing → probabilistic train-time augmentation + deterministic per-transform evaluation grid.
  • Generalization to unseen generators → family-disjoint train/test splits; entire generators held out.
  • Honest, reproducible evaluation → ROC-AUC, balanced accuracy, precision/recall/F1, FPR per transform; validation-tuned threshold.
  • Trade-off awareness → robustness-weighted selection score, explicit FPR discussion, documented limitations.

How we built it

Firstly, we aggressively augmented our training data with random transformation on our images dataset. (https://arxiv.org/abs/1912.11035) Then, we trained a few variant with the same dataset but different hyperparameters, which will be used for our model soup (https://arxiv.org/abs/2203.05482). We tried two sets of datasets, both from GenImage and contain various families of AI generators. The entire GenImage set contain ADM, BigGAN, GLIDE, Stable DIffusion v1.5, Midjourney, VQDM, and Wukong. For the first set, we held out Midjourney and Wukong entirely as testing set. For the second set, we held out ADM and GLIDE.

Development tools used

  • VSCode — fast code edits while training runs
  • Google Colab — live shared evaluation so multiple teammates can view at once
  • Hugging Face Hub — model weight storage & sharing
  • Git + GitHub, uv, NVIDIA RTX 3050

Models or APIs used

  • torchvision.resnet50 (ImageNet-pretrained) with a binary head — no external inference APIs; fully self-hosted.

Libraries & frameworks used

  • PyTorch 2.11, torchvision 0.26, Pillow, NumPy, pandas, Streamlit, Python 3.11 (uv)

Datasets & assets used

  • GenImage benchmark (real images paired with AI images per generator family).
  • Training/validation: Midjourney, Wukong, GLIDE, Stable Diffusion v1.5, VQDM.
  • Unseen-generator test: ADM, GLIDE.

Challenges we ran into

Firstly, the lack of computing power is a problem for us. Time is also a big issue. Hence, we borrowed our friends' computers to run sufficient trainings. Secondly, the model did not perform well on the testing set against Midjourney. We looked into the dataset and realised Midjourney generated images were so good that we humans could not differentiate the AIGC from the real content. Hence, we decided to train the second set. We also did a lot of research to find out techniques to increase our models' robustness but due to time constraint, we were not able to test out all the strategies. One such strategy is patch-based learning.

Accomplishments that we're proud of

I think as non-ML students, all of us were studying on ML on the first day and reading up lots of paper just to understand what is going on in the modern model training scene. The learning process and eventually making our own very first models is truly satisfying and we are proud of it.

What we learned

Lots.

What's next for PikaPic

More experiment with different techniques! Let's keep model as pets.

Share this project:

Updates

Submission history