-
-
Architecture: Frozen CLIP semantic features and FFT frequency signals are combined using weighted probability fusion (α = 0.55).
-
Model comparison: Controlled experiments showed that probability fusion contributed more to performance than changing the visual backbone.
-
Robustness: TrueFrame achieved a 0.9796 combined validation score across 12 conditions, all exceeding 0.91 ROC-AUC.
-
Practicality: TrueFrame uses frozen features and lightweight classifiers for fast inference without fine-tuning or paid APIs.
TrueFrame: Robust Detection of AI-Generated Images
Inspiration
AI-generated images are becoming increasingly realistic, creating new risks involving misinformation, impersonation, fraud, and declining trust in online content. However, detecting these images in a laboratory setting is only part of the problem.
Images uploaded to real platforms are rarely preserved in their original form. They are compressed, resized, cropped, blurred, filtered, or repeatedly reposted. These transformations can erase or distort the subtle artifacts on which many AI-image detectors depend.
This inspired us to build TrueFrame: a lightweight detector designed not only to distinguish authentic images from AI-generated ones, but also to remain reliable after realistic transformations.
What it does
TrueFrame accepts a directory of images and produces a continuous confidence score between 0 and 1 for each image, representing the likelihood that it is AI-generated.
Instead of relying on one type of evidence, TrueFrame analyses every image from two complementary perspectives:
- Semantic and visual evidence: a frozen CLIP ViT-B/32 model captures the image’s objects, composition, texture, and overall visual meaning.
- Frequency-domain evidence: a Fast Fourier Transform exposes spectral patterns and generator artifacts that may not be visible in normal pixel space.
Separate logistic-regression classifiers convert these representations into probabilities. We combine them using scalar-gated probability fusion:
$$ P_{\mathrm{AIGC}} = \alpha P_{\mathrm{CLIP}} + (1-\alpha)P_{\mathrm{FFT}}, \quad \alpha = 0.55 $$
The resulting score can be used directly or converted into a REAL/AIGC decision using an appropriate threshold.
How we built it
Our spatial branch uses the frozen openai/clip-vit-base-patch32 model to extract a 512-dimensional embedding from each image. Freezing CLIP kept training efficient while retaining its strong pretrained visual representation.
For the frequency branch, we convert each RGB channel into the frequency domain using a two-dimensional FFT. We calculate its log-magnitude spectrum and average it into 32 radial frequency bins per channel, producing a compact 96-dimensional feature vector.
Each branch uses a StandardScaler followed by logistic regression. We selected the CLIP probe’s regularisation and class-weight settings using three-fold cross-validation, while the fusion weight was selected through a validation sweep from 0 to 1 in steps of 0.05.
Our primary dataset was saberzl/SID_Set. Under our hackathon compute budget, we sampled 500 images from each of its three original classes for training and 200 per class for validation, giving us 1,500 training and 600 validation images.
We evaluated the final detector across 12 conditions:
- Clean images
- JPEG quality 90, 50, and 30
- Gaussian blur
- Resize to 0.5× and 0.25× before upscaling
- Gaussian noise at two strengths
- Brightness and contrast adjustment
- Centre crop to 80%
Final combined score = 0.5 × clean ROC-AUC + 0.5 × average transformed ROC-AUC
TrueFrame achieved a clean ROC-AUC of 0.9861, an average transformed ROC-AUC of 0.9732, and a final combined validation score of 0.9796. Every evaluated condition remained above 0.91 ROC-AUC.
We packaged the final system as a command-line inference script that loads the trained classifiers, automatically uses CUDA when available, falls back to CPU, and writes the required image_path and pred fields to JSON.
Challenges we faced
Combining semantic and frequency evidence
Our first CLIP+FFT design concatenated both feature sets into one classifier. Although it improved some results, its performance on cropped images deteriorated significantly. FFT features capture global image structure, and cropping can alter that structure dramatically.
We learned that complementary features should not always be merged at the feature level. Training independent classifiers and combining their probabilities allowed the semantic branch to compensate when frequency evidence became unreliable.
Identifying what caused the improvement
We also tested DINOv2 as an alternative visual backbone. It produced a competitive combined score, but CLIP with weighted probability fusion achieved slightly better overall results and more stable fixed-threshold accuracy.
This taught us that the fusion strategy contributed more to the improvement than simply replacing one large visual backbone with another.
Domain shift and threshold calibration
A small exploratory test using independently sourced internet images exposed an important limitation. At our default threshold, polished professional photographs were sometimes classified as AI-generated.
On that small sample, increasing the threshold improved observed accuracy from 60% to 90%, but the same threshold reduced performance on our SID sample. We therefore retained the SID-validated default and continued outputting continuous confidence scores rather than pretending that one threshold works for every deployment domain.
This showed us that strong ROC-AUC does not automatically guarantee a universally calibrated decision boundary.
Limited time and compute
We worked with a hackathon-scale subset rather than the complete SID_Set. This required us to prioritise efficient experiments and controlled ablations. Frozen backbones and lightweight classifiers allowed us to compare several ideas without expensive end-to-end fine-tuning.
What we learned
The most important lesson was that robustness must be evaluated explicitly. Strong clean-image performance says little about how a detector behaves after compression, cropping, noise, or resizing.
We also learned that:
- A carefully designed lightweight fusion method can matter more than choosing a larger backbone.
- Semantic and frequency evidence fail in different ways, making them valuable complements.
- ROC-AUC and fixed-threshold accuracy reveal different aspects of performance.
- Error analysis is essential for identifying dataset bias and domain shift.
- Simple classifiers on frozen features can provide strong results while remaining fast, reproducible, and practical.
We additionally explored PRNU-style sensor-noise features, multi-scale FFT, azimuthal FFT averaging, and contrastive pretraining. None surpassed our final configuration within the available data and compute budget, but each experiment helped us understand which signals were genuinely useful.
What we are proud of
TrueFrame is more than a notebook experiment. It includes trained artifacts, a reproducible inference script, a robustness evaluation covering realistic transformations, and an honest analysis of false positives and cross-domain limitations.
The model contains approximately 151 million parameters, requires no paid API, and uses lightweight logistic-regression classifiers that train in under a second once features have been extracted.
What is next
Given more time, we would:
- Train and evaluate on a larger portion of SID_Set.
- Test on multi-generator datasets such as WildFake.
- Include more professional photography in the REAL training distribution.
- Develop domain-aware probability calibration.
- Investigate diffusion-reconstruction signals as a third complementary branch.
- Expand the external evaluation into a properly sampled cross-dataset benchmark.
Our biggest takeaway is simple: reliable AI-image detection is not just about achieving high clean accuracy. It is about understanding how the detector fails when images leave the laboratory and enter the real world.
Log in or sign up for Devpost to join the conversation.