CrossGuard
CrossGuard detects AI-generated images and keeps working after the image has been compressed, cropped, blurred, or shrunk to a thumbnail.
How our solution addresses the problem statement
A detector that only works on pristine images does not work in practice. By the time a synthetic image is worth catching, it has been screenshotted, re-encoded by a messaging app, cropped to a profile frame, and reposted at reduced scale. Each of those steps strips away the fine-grained generation artefacts that detectors rely on.
CrossGuard makes robustness the training objective rather than an evaluation afterthought. Every training step sees each image twice: once clean, once through a randomly sampled distortion drawn from the same transform families the problem statement names. The loss has three terms: a focal classification loss on both views, a KL term tying the distorted view's prediction to the clean view's, and a feature-space MSE, applied through a small residual correction network, that pulls the two representations together. The model is penalised directly for changing its answer when an image is degraded.
Across the 14-cell transform grid on the held-out test split, CrossGuard holds a macro AUROC of 0.9865 against a clean baseline of 0.9921. The largest single drop is 0.0192, at 0.25x downscaling.
| AUROC | |
|---|---|
| Clean | 0.9921 |
| Macro across all 14 transform cells | 0.9865 |
Worst cell (resize_0.25x) |
0.9728 |
Calibrated output with thresholds fixed in advance. CrossGuard returns a calibrated probability, not a raw score. Temperature and operating thresholds were fitted on the validation split and applied unchanged to test; no threshold was derived from the final evaluation. At the deployed 1% false-positive operating point (threshold 0.9491) the model reaches TPR 0.8667 at FPR 0.0055.
Designed around the cost of a false accusation. A missed synthetic image stays up and remains catchable by other signals. A false positive tells a real creator that their own work is machine-made. CrossGuard ships the stricter threshold: 45 false positives across 8,116 real test images, or 1 in 180, at the cost of missing 13.3% of synthetic images. Scores fall into three triage bands so uncertain cases route to human review instead of automated action.
Generalisation to unseen generators. Five generators were withheld from training entirely. On those generators CrossGuard scores 0.9913 AUROC overall: 0.9963 on unseen diffusion models, 0.9831 on unseen GANs, 0.9978 on other architectures.
Development tools
- VSCode for all coding work.
- Modal for all GPU work. Dataset build, training, calibration, and final scoring ran as detached cloud jobs against a shared network volume.
- Git and GitHub for source control; Hugging Face Hub as private transport for dataset shards and checkpoints between team members.
- Claude Code and ChatGPT Codex as AI pair-programmers during development.
- NVIDIA H100 80GB for training. It provided the best performance to cost ratio for us. (Modal sometimes upgraded us to a H200 at no additional cost whenever there was spare capacity!) We tested a L40S but found it too slow. We also tested a B200 but were unable to get the parameters right to maximum training throughput. (Our final conclusion was that the B200 was network-bound)
Models and APIs
- DINOv2 ViT-L/14 (
vit_large_patch14_dinov2.lvd142m, Apache-2.0) as the backbone, at 448×448 input. - A single-logit classification head over globally average-pooled final patch tokens, plus a residual correction network used only by the training-time consistency loss.
- LoRA (rank 32, applied to attention and MLP projections) for the fine-tuning stage, after an initial linear-probe stage on frozen features.
- 306,114,561 parameters in total, within the 2B limit.
- No external inference APIs. The model runs offline from a single checkpoint.
Libraries and frameworks
- PyTorch and torchvision: training and inference
- timm: backbone architectures and pretrained weights
- peft: LoRA fine-tuning
- scikit-learn: AUROC, average precision, bootstrap confidence intervals
- Pillow: image decoding and the distortion grid
- NumPy, pandas and PyArrow: dataset manifest and sharded storage
- imagehash: perceptual-hash screening
- huggingface-hub: artefact transport
- tqdm: progress reporting
- Modal: cloud orchestration
Datasets and assets
| Asset | Licence | Use |
|---|---|---|
| SID_Set | CC BY 4.0 | Training and evaluation: real images, FLUX.1-dev synthetics |
| WildFake (GAN / Diffusion / Other) | Apache 2.0 per uploader | Training, and the held-out generator axes |
| CIFAKE | MIT, from CIFAR-10 and SD-1.4 | Evaluation only, thumbnail-regime row |
| COCO train2017 | CC BY 4.0 annotations, per-image Flickr terms | Enters the build through WildFake's real slices |
| DINOv2 ViT-L/14 | Apache-2.0 | Backbone weights |
The final build is 327,311 images: 257,433 training, 28,149 validation, 41,729 test, split so that no generator or source appears in more than one split.
Assumption on the restricted WildFake slices
The rules ban exactly two WildFake slices from training: the COCO val2017 reals and the DALL·E Advanced fakes that form the validation benchmark. We train on other WildFake slices. Both banned slices are excluded structurally, by path, at the point the archives are read:
coco2017/val2017/for the reals, andDALLE/Advanced/for the DALL·E fakes, which share an archive with the DALLE2 images we do train on. A perceptual-hash screen over COCO val2017 runs alongside the path exclusion as a second check, covering the case of those images reappearing elsewhere under different filenames. The banned slices are used only as the demo-benchmark evaluation row and never touch training, model selection, calibration, or threshold choice.
Limitations and next steps
CrossGuard ships a single model branch. Two additional branches were built and tested but never trained on the final dataset, so branch fusion is unmeasured.
The false-positive classes that matter most in deployment, namely screenshots, CGI and renders, heavily filtered photographs and AI-upscaled real images, could not be quantified, because no corpus covering them cleared licence review before the project's data cutoff. They are characterised by mechanism in the error analysis rather than by measurement.
The transform grid covers accidental degradation from ordinary redistribution. It is not an adversarial evaluation: no attacker optimised against this model. With more time, the priorities are a licence-cleared hard-real evaluation slice, per-cell false-positive rates alongside the AUROC figures, and an adversarial pass covering re-compression chains and deliberate thumbnail evasion.
CrossGuard Error Analysis Note
Test split, n = 41,729 (33,613 AI-generated, 8,116 real). Counts are taken at the deployed operating point: threshold 0.9491, fitted on validation and applied unchanged to test.
At that threshold the model produces 4,482 false negatives and 45 false positives.
False positives
45 real images out of 8,116, a rate of 1 in 180.
The image-level categories these fall into could not be measured: no corpus of screenshots, renders, or filtered photographs cleared licence review before the project's data cutoff, so no such slice exists in the evaluation build. The categories below follow from what the model keys on, which is camera-level detail such as sensor noise, optical imperfection, and compression history.
- Screenshots, UI-heavy images. Flat regions, hard synthetic edges, no sensor noise. Frequency: high on a social platform.
- Heavily filtered photographs. Beauty filters and auto-enhance suppress sensor-level detail. Frequency: high.
- CGI, renders, game captures. Genuinely rendered, so the "real" label is itself contestable. Frequency: moderate.
- AI-upscaled real photographs. A real capture whose pixels were partly generated. Frequency: rising.
The last two categories expose a limit of the task definition rather than of the model. A game capture is not a photograph, and an AI-upscaled holiday photo is neither wholly real nor wholly generated, so a binary label forces an answer that the image does not support.
False negatives
4,482 AI-generated images out of 33,613, a miss rate of 13.3%.
Misses concentrate where image degradation is severe. The four weakest cells of the 14-cell transform grid:
- resize 0.25x: AUROC 0.9728, down 0.0192 from clean
- noise sigma 0.10: AUROC 0.9740, down 0.0180 from clean
- blur sigma 2.0: AUROC 0.9775, down 0.0146 from clean
- JPEG q30: AUROC 0.9847, down 0.0074 from clean
These are one failure mode, not four. Generation artefacts live in high-frequency detail; quarter-scale downscaling, heavy noise, and strong blur each destroy that band before the image reaches the 448 px input. The practical consequence is that a thumbnail is the cheapest available evasion, and any deployment should score the highest-resolution copy of an image it can obtain.
Generator architecture is the second axis. Across five generators withheld entirely from training:
- Unseen diffusion (Imagen, VQDM): AUROC 0.9963
- Unseen other (MAGE): AUROC 0.9978
- Unseen GAN (GigaGAN, starGAN): AUROC 0.9831
GAN transfer is measurably weaker than diffusion transfer, tracking the training mix, which contains more diffusion exposure. Architectures further from that mix should be expected to perform closer to the GAN figure than the diffusion one.
Trade-offs
Recall traded for precision. The three available operating points produce very different error profiles:
- 1% FPR (deployed), threshold 0.9491: 4,482 false negatives, 45 false positives, 4,527 total errors
- Balanced accuracy, threshold 0.8976: 3,573 false negatives, 72 false positives, 3,645 total errors
- 5% FPR, threshold 0.3779: 1,693 false negatives, 271 false positives, 1,964 total errors
The deployed point has the highest total error count of the three. Moving from the 5% point to the 1% point accepts 2,789 additional missed images to avoid 226 false accusations, a ratio of roughly twelve to one.
That exchange rate is justified by the asymmetry in what each error costs. A missed synthetic image remains on the platform and stays detectable by provenance signals, reporting, and later rescoring. A false positive tells a creator their own work is machine-generated, and at platform scale produces appeals, reputational damage, and lost trust. The twelve-to-one ratio is the price paid for a false-accusation rate of 1 in 180.
Discrimination traded for a single model. CrossGuard ships one branch. Two further branches were built and tested but never trained on the final dataset, so the fusion gain is unmeasured.
Robustness measured, adversaries not. The transform grid covers accidental degradation from ordinary redistribution. No attacker optimised against this model, and the thumbnail result indicates where a motivated one would begin.
Calibration usable, not exact. Brier score 0.0404, expected calibration error 0.0445. Scores are sound for ranking and thresholding; a score of 0.90 corresponds to roughly 0.86 to 0.94 empirical precision, so individual values should not be presented to users as precise probabilities.
Consequences for deployment
The error profile supports triage rather than enforcement. Scores fall into three bands: above 0.9491 for prioritised review, 0.3779 to 0.9491 for review with additional context, below 0.3779 for lower risk.
Three measurements would materially change this analysis: a licence-cleared hard-real slice, which would convert the false-positive categories above from mechanism into counts; per-cell false-positive rates, which scripts/score_test.py computes but the shipped report predates; and an adversarial evaluation covering re-compression chains and deliberate thumbnail evasion.
Built With
- adversarial-robustnesss
- ai-generated-image-detection
- aigc-detection
- deepfake-detection
- dinov2
- image-forensics
- lora
- modal
- pytorch
Log in or sign up for Devpost to join the conversation.