Inspiration
AI-generated image detectors are often evaluated on clean files, but real images rarely stay clean. They are reposted, recompressed, resized, blurred, cropped, color-adjusted, or contaminated by noise. These transformations can weaken the subtle evidence detectors rely on.
We wanted to build more than a detector that reports one impressive benchmark number. Our goal was a reproducible forensic system whose behavior under common image transformations could be inspected, tested, and honestly documented.
What it does
Adaptive AIGC Forensics recursively scans an image directory and produces a deterministic JSON file containing an AI-generated-image confidence score for every supported image:
[
{"image_path": "relative/path/to/image.png", "pred": 0.73}
]
The submitted system combines two complementary experts:
- A frozen RGB expert that learns visual forensic evidence directly from normalized image pixels.
- A compact signal expert that measures explicit Fourier-energy, neighbouring-pixel, and residual statistics.
- Learned static fusion, which combines their calibrated logits using a fixed allocation of 0.677 RGB and 0.323 signal for every image.
Both experts receive the same decoded, orientation-corrected RGB observation. This ensures that differences between their predictions come from how they analyze the image, rather than from inconsistent decoding or preprocessing.
## How we built it
The RGB branch uses PyTorch, torchvision, timm, and a revision-pinned Community Forensics checkpoint. The signal branch converts each image into a deterministic 26-value signal representation and passes it through a compact 26→16→1 tanh MLP with 449 trainable scalar parameters.
We created a reproducible corruption harness for six common transformation families:
- JPEG recompression
- Gaussian blur
- Resizing
- Additive noise
- Color adjustment
- Cropping
Our data is divided at the source-image level, before transformed variants are created. Every clean and transformed observation derived from one source therefore remains in the same partition. The pipeline also checks for exact and perceptual overlap across partitions.
The source allocation contains 8,000 expert-training images, 2,000 fusion-training images, 2,000 internal-validation images, and 2,000 sealed-internal-test images. The organizer-provided demonstration data remains evaluation-only and cannot influence training, calibration, model selection, fusion weights, or thresholds.
For reproducibility, model files, manifests, results, and completion receipts are bound with SHA-256 checksums. The inference command validates the checkpoint, signal model, calibrators, preprocessing configuration, and fusion bindings before constructing either expert.
## Challenges we ran into
Preventing data leakage
Every source image can generate many transformed observations. Splitting those observations independently would allow near-identical versions of one image to appear in both training and validation. We therefore split by source identity first and keep every derived observation attached to its source partition.
Proving complementarity
A second expert is only useful if it corrects mistakes rather than duplicating the first expert. We evaluated which calibrated RGB errors were corrected by the signal expert and measured the fused system against calibrated RGB-only performance.
Keeping artifacts reproducible
It was not enough for two runs to produce visually similar results. We needed stable ordering, deterministic transformations, strict prediction schemas, atomic output publication, pinned model revisions, and byte-level artifact verification.
## Accomplishments that we're proud of
Our multi-expert architecture genuinely resulted in a marked improvement in accuracy.
Log in or sign up for Devpost to join the conversation.