Inspiration

We kept coming back to one thing in the Track 5 brief: the photo will get JPEG’d, blurred, resized, cropped, or colour-shifted before anyone scores it. Winning on a clean PNG is the easy version of the problem.

The other thing that stuck with us is how AIGC detectors actually fail. Forensic tricks (SRM residuals, FFT) look great on untouched pixels and then fall over after compression. CLIP-style semantics survive a repost better, but then you don’t really know if the model saw a generator, or just “this looks like a picture.” So we built RIFT as two opinions and a gate that has to pick who to trust, plus one cutoff $\tau$ chosen on clean SID_Set val at ~5% FPR and left alone after that. If you retune $\tau$ for every JPEG quality, you are not measuring robustness.

What it does

RIFT answers a simple question: is this photo AI-generated, or not? It returns a score between 0 and 1 (probability the image is AI-generated) and is built to keep that call stable after the image has been compressed, blurred, resized, cropped, or otherwise reposted.

It does that by combining two views: frozen CLIP for broad visual structure, and a forensic stream on residuals and frequency patterns. We only train a small set of heads on top of CLIP, not the ViT itself. Scoring uses one low-FPR threshold, frozen across every official transform — not a fresh cutoff for clean images and another for JPEG-30. The WildFake holdout never enters training. The point is a detector that still works on a social-media repost, not just a laboratory PNG.

How we built it

We froze CLIP ViT-B/32 and only trained a small projection on top of it. Next to that, a CNN looks at SRM high-pass residuals, log-FFT magnitude, and FFT phase (phase holds up under JPEG a bit better than magnitude). A softmax gate mixes the two. We added an auxiliary forensic head and a gate-entropy term because we knew the gate would try to ignore the forensic stream. It still did.

Training uses the official transform set, plus a consistency loss so a clean view and a mashed view of the same photo should agree. AdamW, cosine LR, mixed precision, early stopping. All of this ran on Kaggle; we wrote the code in VS Code / Cursor, kept it on GitHub, and drove train/eval from the CLI. No paid API.

Data is SID_Set from Hugging Face: real → REAL, full synthetic → FAKE, tampered dropped (it isn’t a binary AIGC label). Images cached, longest side 512, official train/val shards kept apart. The checkpoint we are submitting was trained on 20k images, not the 40k-per-class config. CIFAKE was how we checked the pipeline still ran. The WildFake holdout (COCO val2017 + DALL·E Advanced) is eval-only — we never train on it, and we do not quote transform numbers that aren’t in the repo.

Challenges we ran into

SID_Set is huge. We cached what we could, froze CLIP, and stayed at 20k. Getting the tampered class out of training mattered; stuffing it into FAKE would have been a silent bug. The entropy/aux tricks did not stop the gate collapsing. We also shipped a requirements.txt that forgot Hub / Datasets / PyArrow until late.

Accomplishments that we're proud of

We're proud that RIFT is a full pipeline, not a clean-image classifier with a lucky accuracy number. From one repo you can stream and validate data, train a frozen-CLIP detector, apply the official robustness augs, score images with an AIGC probability, freeze one threshold across 15 evaluation conditions, watch the CLIP vs forensic gate, dump robustness tables, list the actual errors, and keep the holdout out of training.

On the SID_Set real-vs-full-synthetic split, the checkpoint we validated hit AUROC 0.9996. At threshold ≈ 0.0317 (set for ~5% FPR) that is 97.6% accuracy and 99.9% recall. We are calling that what it is: excellent in-domain separation. It is not a promise about unseen generators or the official WildFake holdout.

What we learned

A frozen threshold is the whole game. CLIP is very good at ranking SID_Set in-domain; that is not the same as “works on a generator we have never seen.” Our gate put almost all its weight on CLIP, so this model is basically a frozen CLIP probe with a forensic branch we paid for and barely used. And at 5% FPR, the failures you will actually see are false accusations — compressed photos, screenshots, weirdly smooth textures — which we should pull from the prediction file, not invent from the table.

What's next for RIFT

  • Run the isolated WildFake demonstration holdout
  • Try generators and authentic sources that are not in SID_Set
  • Ablate CLIP-only vs forensic-only vs the gated hybrid
  • Keep threshold calibration off the final test set
  • Temperature-scale the probabilities
  • Hard-negative tests: screenshots, edits, memes, heavily compressed real photos
  • Train on a larger SID_Set slice
  • Put every dataset dependency in the install config

Built With

+ 11 more
Share this project:

Updates

Submission history