Inspiration
Every image a moderation system actually sees has already been through an upload pipeline. It has been re-encoded as JPEG, resized for a thumbnail, maybe blurred or cropped or run through a filter.
That matters more than it sounds. The sharpest evidence that an image was generated—periodic upsampling ripples in the Fourier spectrum, a missing sensor-noise floor, and over-smooth micro-texture—all live in the high-frequency band. And the high-frequency band is precisely what compression and downscaling destroy first.
So a social platform's redistribution pipeline is, entirely by accident, an almost optimal attack on frequency-domain forensics. A detector tuned on pristine images can report 99% accuracy in a notebook and then collapse to near coin-flip on the same images once they have been posted once. We wanted to build the version that does not do that.
What it does
Given an image, it returns a calibrated probability that the image is fully AI-generated, plus a plain-language verdict. It is deliberately a binary question: real versus fully synthetic.
One product decision worth stating up front: the training data also contains tampered images—real photographs with an edited region. We label those real. A person still took the photograph, and "is this a real photo?" is the question a viewer is actually asking.
The design goal is not peak accuracy on clean images — that is easy. It is accuracy that holds after the image has been through a real pipeline.
How we built it
Three ideas, stacked.
1 · Two evidence branches. A frozen CLIP ViT-B/16 embedding gives a semantic read—coarse, but it barely moves under compression. Alongside it, a 129-dimensional hand-designed forensic vector (a 48-bin radial FFT log-power profile plus slope and curvature, 64 per-position DCT log-magnitudes, high-pass residual moments per channel, and cross-channel residual correlations) gives a precise read on clean images that degrades fast. Every one of those features is a sentence you can say out loud over a chart, which matters when a detector's output is an accusation.
2 · A degradation-aware gate. Rather than blending the branches with fixed weights, a small head estimates how damaged the image is, and that estimate sets the fusion weights. A pristine PNG leans on frequency evidence; a q30 re-encode leans on semantics. Nothing tells the model which case it is at test time. The estimator trains on free supervision—our augmentation pipeline applied the damage, so it already knows the answer.
3 · A consistency loss. Training pairs each clean view with a degraded one and adds a symmetric-KL term pushing the two predictions together. Ordinary augmentation asks the model to also get the damaged copy right. Consistency asks for the same answer as the clean copy—a strictly stronger constraint and exactly the property the robustness evaluation measures.
Only the head trains: 563,724 parameters on top of an 86M frozen backbone, 23× under the 2 billion cap. Features are extracted once and cached, so a full retrain after a data change takes seconds rather than hours, which is what made it possible to iterate honestly on a single 8 GB laptop GPU.
Challenges we ran into
Four blind spots we could not close The honest headline: We know exactly where this model fails because we went looking. All four are measured, reproducible, and still open.
Hyperrealistic real photos trigger false alarms. Studio-lit product shots, polished travel photography, and photographs of artwork get flagged. The model over-indexes on the ultra-clean, low-noise look that high-end AI art shares with professional photography — a 6.0% false-alarm rate on genuine photos in the benchmark's hardest real-image config, up from 1.7% before we added DALL·E 3 data. Plain snapshots stay under 1%. This is the cost of the capability we bought, and it is the limitation that most constrains where the model can be deployed.
Heavy noise works as a shield. Intense compression or additive noise buries the high-frequency fingerprint the forensic branch depends on, and a degraded synthetic starts to look like a messy real snapshot. At σ=0.1 noise—the worst cell in our grid—recall at a 1% false-positive operating point falls from 99.7% to 89.6%. Ranking survives (cell AUC 0.993), but roughly one AI image in ten slips past a strict threshold once it is grainy enough. Adding noise is not a clever attack, and it partly works.
GigaGAN slips through. DALL·E 3 (93% recall) and Midjourney v5 (87%) are caught reliably. GigaGAN sits at a 4% detection rate. We added ProGAN to training specifically to fix this, and it did not transfer—a 2018 category GAN and a 2023 text-to-image GAN leave different traces. Architecture was not the missing piece; the right training data were.
Localized edits are invisible. A real photo with a small AI-edited region scores as ~100% real. Partly by design—a tampered photo was still taken by a person—but it means inpainting, face swaps, and object removal all pass unflagged. Only whole-image generation is in scope.
Accomplishments that we're proud of
Robustness that actually holds. Across 16 transform cells—JPEG down to q30, blur to σ=2.0, downscale to 0.25×, noise to σ=0.1, ±20% color jitter, 80% crop—the clean AUC is 0.999, and no single cell falls below 0.993. Mean AUC drop under transformation: +0.0017.
It transfers. On a reference benchmark drawn from a completely different corpus, which the model has never trained on and which is perceptual-hash-excluded from our training data, it reaches 0.989 AUC on the leak-free configuration.
We measured our own weaknesses instead of hiding them. The ablation study is a null result, and we published it. The false-positive cost of our final data pass is in the README, the limitations section, and the commit message.
What we learned
A benchmark number without its provenance is worth nothing. The same model scores 0.999 on one config of the same benchmark and 0.808 on another. Which one you quote is a claim about your own honesty.
Calibration degrades exactly where coverage does. Expected calibration error is 0.008 on clean held-out images and stays ≤0.026 across all sixteen transform cells. On the transfer benchmark it holds up on the configs we cover well (0.022–0.036) and blows out to 0.301 on cross_generator—the one config dominated by a generator family we barely detect. Calibration turned out not to be a separate problem from coverage; it is a symptom of it and a useful early warning that the model is out of its depth.
Architecture is not always the lever. Our ablation showed the frozen CLIP branch alone matches the full model within ±0.001 AUC on this corpus. The frequency branch, the gate, and the consistency loss are a wash here—because after our data passed, the training set became largely semantically separable, and CLIP is already transform-robust. We kept the full architecture (it costs ≈564k parameters and never regresses anything), but we are not going to claim credit it did not earn.
Every capability has a price, and you should name it. Adding DALL·E 3 images took its recall from 0.72 to 0.93 and carried Midjourney up with it—and pushed false positives on polished real photography from 1.7% to 6.0%. We took that trade deliberately, and we say so.
What's next for AI-IMAGE-DETECTOR
One line per blind spot, in the order we would actually tackle them.
- Modern GANs (blind spot 3). ProGAN went into training and did not transfer to GigaGAN. StyleGAN-3 / GigaGAN-class data — each image paired with resolution- and aspect-matched real controls, the way our second attempt was built — is the clearest next win.
- Recover the false-positive headroom (blind spot 1) with hard-negative mining specifically on studio, product and stock photography. The same targeted pass took false positives from 1.2% to 0.4% earlier in the project.
- Train against noise as an attack (blind spot 2), not just incidental degradation — noise-injected synthetics, and a check on whether the gate learns to distrust the forensic branch harder at high σ.
- A localisation head (blind spot 4) beside the whole-image classifier, so inpainting and face swaps stop passing unflagged.
- Per-domain calibration — re-fit temperature on deployment traffic rather than shipping one global scalar.
- An abstain option. Extend the gate to emit a reliability estimate so the system can decline on images too degraded to judge, instead of guessing.
- Unfreeze the top transformer blocks and see whether the semantic branch can be pushed further than th## Inspiration
Every image a moderation system actually sees has already been through an upload pipeline. It has been re-encoded as JPEG, resized for a thumbnail, maybe blurred or cropped or run through a filter.
That matters more than it sounds. The sharpest evidence that an image was generated — periodic upsampling ripples in the Fourier spectrum, a missing sensor-noise floor, over-smooth micro-texture — all live in the high-frequency band. And the high-frequency band is precisely what compression and downscaling destroy first.
So a social platform's redistribution pipeline is, entirely by accident, an almost optimal attack on frequency-domain forensics. A detector tuned on pristine images can report 99% accuracy in a notebook and then collapse to near-coin-flip on the same images once they have been posted once. We wanted to build the version that does not do that.
What it does
Give it an image, it returns a calibrated probability that the image is fully AI-generated, plus a plain-language verdict. It is deliberately a binary question: real versus fully synthetic.
One product decision worth stating up front: the training data also contains tampered images — real photographs with an AI-edited region. We label those real. A person still took the photograph, and "is this a real photo?" is the question a viewer is actually asking.
The design goal is not peak accuracy on clean images — that is easy. It is accuracy that holds after the image has been through a real pipeline.
How we built it
Three ideas, stacked.
1 · Two evidence branches. A frozen CLIP ViT-B/16 embedding gives a semantic read — coarse, but it barely moves under compression. Alongside it, a 129-dimensional hand-designed forensic vector (a 48-bin radial FFT log-power profile plus slope and curvature, 64 per-position DCT log-magnitudes, high-pass residual moments per channel, and cross-channel residual correlations) gives a precise read on clean images that degrades fast. Every one of those features is a sentence you can say out loud over a chart, which matters when a detector's output is an accusation.
2 · A degradation-aware gate. Rather than blending the branches with fixed weights, a small head estimates how damaged the image is, and that estimate sets the fusion weights. A pristine PNG leans on frequency evidence; a q30 re-encode leans on semantics. Nothing tells the model which case it is at test time. The estimator trains on free supervision — our augmentation pipeline applied the damage, so it already knows the answer.
3 · A consistency loss. Training pairs each clean view with a degraded one and adds a symmetric-KL term pushing the two predictions together. Ordinary augmentation asks the model to also get the damaged copy right. Consistency asks for the same answer as the clean copy — a strictly stronger constraint, and exactly the property the robustness evaluation measures.
Only the head trains: 563,724 parameters on top of an 86M frozen backbone, 23× under the 2-billion cap. Features are extracted once and cached, so a full retrain after a data change takes seconds rather than hours — which is what made it possible to iterate honestly on a single 8 GB laptop GPU.
Challenges we ran into
Four blind spots we could not close
The honest headline: we know exactly where this model fails, because we went looking. All four are measured, reproducible, and still open.
Hyperrealistic real photos trigger false alarms. Studio-lit product shots, polished travel photography and photographs of artwork get flagged. The model over-indexes on the ultra-clean, low-noise look that high-end AI art shares with professional photography — a 6.0% false-alarm rate on genuine photos in the benchmark's hardest real-image config, up from 1.7% before we added DALL·E 3 data. Plain snapshots stay under 1%. This is the cost of the capability we bought, and it is the limitation that most constrains where the model can be deployed.
Heavy noise works as a shield. Intense compression or additive noise buries the high-frequency fingerprint the forensic branch depends on, and a degraded synthetic starts to look like a messy real snapshot. At σ=0.1 noise — the worst cell in our grid — recall at a 1%-false-positive operating point falls from 99.7% to 89.6%. Ranking survives (cell AUC 0.993), but roughly one AI image in ten slips past a strict threshold once it is grainy enough. Adding noise is not a clever attack, and it partly works.
GigaGAN slips through. DALL·E 3 (93% recall) and Midjourney v5 (87%) are caught reliably. GigaGAN sits at a 4% detection rate. We added ProGAN to training specifically to fix this and it did not transfer — a 2018 category GAN and a 2023 text-to-image GAN leave different traces. Architecture was not the missing piece; the right training data is.
Localised edits are invisible. A real photo with a small AI-edited region scores as ~100% real. Partly by design — a tampered photo was still taken by a person — but it means inpainting, face swaps and object removal all pass unflagged. Only whole-image generation is in scope.
Accomplishments that we're proud of
Robustness that actually holds. Across 16 transform cells — JPEG down to q30, blur to σ=2.0, downscale to 0.25×, noise to σ=0.1, ±20% colour jitter, 80% crop — clean AUC is 0.999 and no single cell falls below 0.993. Mean AUC drop under transformation: +0.0017.
It transfers. On a reference benchmark drawn from a completely different corpus, which the model has never trained on and which is perceptual-hash excluded from our training data, it reaches 0.989 AUC on the leak-free configuration.
We measured our own weaknesses instead of hiding them. The ablation study is a null result and we published it. The false-positive cost of our final data pass is in the README, the limitations section, and the commit message.
What we learned
A benchmark number without its provenance is worth nothing. The same model scores 0.999 on one config of the same benchmark and 0.808 on another. Which one you quote is a claim about your own honesty.
Calibration degrades exactly where coverage does. Expected calibration
error is 0.008 on clean held-out images and stays ≤0.026 across all sixteen
transform cells. On the transfer benchmark it holds up on the configs we
cover well (0.022–0.036) and blows out to 0.301 on cross_generator —
the one config dominated by a generator family we barely detect. Calibration
turned out not to be a separate problem from coverage; it is a symptom of it,
and a useful early warning that the model is out of its depth.
Architecture is not always the lever. Our ablation showed the frozen CLIP branch alone matches the full model within ±0.001 AUC on this corpus. The frequency branch, the gate, and the consistency loss are a wash here — because after our data passes the training set became largely semantically separable, and CLIP is already transform-robust. We kept the full architecture (it costs ≈564k parameters and never regresses anything) but we are not going to claim credit it did not earn.
Every capability has a price, and you should name it. Adding DALL·E 3 images took its recall from 0.72 to 0.93 and carried Midjourney up with it — and pushed false positives on polished real photography from 1.7% to 6.0%. We took that trade deliberately, and we say so.
What's next
One line per blind spot, in the order we would actually tackle them.
- Modern GANs (blind spot 3). ProGAN went into training and did not transfer to GigaGAN. StyleGAN-3 / GigaGAN-class data — each image paired with resolution- and aspect-matched real controls, the way our second attempt was built — is the clearest next win.
- Recover the false-positive headroom (blind spot 1) with hard-negative mining specifically on studio, product and stock photography. The same targeted pass took false positives from 1.2% to 0.4% earlier in the project.
- Train against noise as an attack (blind spot 2), not just incidental degradation — noise-injected synthetics, and a check on whether the gate learns to distrust the forensic branch harder at high σ.
- A localisation head (blind spot 4) beside the whole-image classifier, so inpainting and face swaps stop passing unflagged.
- Per-domain calibration — re-fit temperature on deployment traffic rather than shipping one global scalar.
- An abstain option. Extend the gate to emit a reliability estimate so the system can decline on images too degraded to judge, instead of guessing.
- Unfreeze the top transformer blocks and see whether the semantic branch can be pushed further.e whole-image classifier, so inpainting and face swaps stop passing unflagged.
- Per-domain calibration — re-fit temperature on deployment traffic rather than shipping one global scalar.
- An abstain option. Extend the gate to emit a reliability estimate so the system can decline on images too degraded to judge, instead of guessing.
- Unfreeze the top transformer blocks and see whether the semantic branch can be pushed further.
Log in or sign up for Devpost to join the conversation.