Professional System Specification: Robust AI-Generated Content (AIGC) Detection Model

Target Application: Enterprise-grade AI-Generated Content Detection & Robustness Engine

Core Objective: High-precision, zero-leakage classification of authentic vs. AI-generated visual media under severe real-world redistribution degradations (e.g., social media re-compression, aggressive rescaling, and additive sensor noise).


Technical Performance & Empirical Benchmark

Primary Model Classification Metrics

Evaluated across internal validation (val_internal) and out-of-distribution transfer test splits (test_transfer). The deployed dual-tower ViT-L system delivers high operational accuracy and low False Positive Rates (FPR).

Split AUC TPR@1%FPR Dataset size (Real / Fake) Model Families & Counts
train — — 350,564 (175,670 / 174,894) SD (5), Diffusion/Pixel (5), GAN (6), PixArt (3), SD-Adaptor (3), SD-Img2Img (3), Kandinsky (3), OmniGen (2), Infinity (2), FLUX (1), Z-Image (1), Open-Weight Singletons (5)
val 0.9654 0.7348 53,600 (26,800 / 26,800) VQ/Autoencoder (4), SD (2), FLUX (2), Klein (2), Open-Weight Singletons (3), Commercial Closed-Weight (2)
test 0.9963 0.9825 37,295 (18,647 / 18,648) Open-Weight Singletons (4), SD (1), FLUX (1), Z-Image (1), Commercial Closed-Weight (11)
demo 0.9989 0.9758 — Coco val 2017 (1), Dall E Advanced

Robustness & Stress-Test Degradation Matrix

Evaluated against a baseline reference performance ($\text{AUC} = 0.9551$). Operational invariance is sustained across typical degradation pipelines, with catastrophic degradation observed only at extreme compression thresholds ($\text{JPEG } q=10$).

Applied Augmentation Pipeline Metric (AUC) Differential Score (Δ vs Clean Baseline) Operational Impact
Clean Baseline Reference 0.9551 — Baseline Nominal
JPEG Compression ($q = 90$) 0.9645 $+0.0094$ High Stability
JPEG Compression ($q = 70$) 0.9645 $+0.0094$ High Stability (Unseen Severity)
JPEG Compression ($q = 50$) 0.9550 $-0.0001$ Negligible Shift
JPEG Compression ($q = 10$) 0.8949 $-0.0601$ High Degradation
Gaussian Blur ($\sigma = 0.5$) 0.9673 $+0.0122$ Performance Gain
Gaussian Blur ($\sigma = 1.0$) 0.9585 $+0.0035$ High Stability (Unseen Severity)
Gaussian Blur ($\sigma = 2.0$) 0.9500 $-0.0051$ Nominal Shift
Spatial Downscale ($0.5\times$) 0.9600 $+0.0049$ High Stability
Spatial Downscale ($0.25\times$) 0.9353 $-0.0198$ Moderate Shift
Additive Noise ($\sigma = 0.02$) 0.9584 $+0.0033$ Nominal Shift
Additive Noise ($\sigma = 0.05$) 0.9565 $+0.0015$ Nominal Shift
Additive Noise ($\sigma = 0.10$) 0.9411 $-0.0140$ Moderate Shift
Color Jitter (Brightness $\pm 20\%$) 0.9488 $-0.0062$ Nominal Shift
Color Jitter (Contrast $\pm 20\%$) 0.9617 $+0.0066$ High Stability
Color Jitter (Saturation $\pm 20\%$) 0.9572 $+0.0021$ Nominal Shift
Spatial Center Crop ($80\%$) 0.9595 $+0.0045$ High Stability

Corpus Curation & Out-of-Distribution Splitting

The training corpus leverages external benchmark distributions combined with targeted internal synthetic image generations. It strictly isolates zero-day commercial architectures and unseen autoencoder lineages to prevent data leakage.

  • External Benchmarks: Baseline assets derived from NTIRE 2026 and WildFake datasets, capturing legacy GANs, foundational diffusion pipelines, and baseline organizer distributions.
  • In-House Targeted Generations: Synthetic visual suites anchored to real portrait photographs from Open Images V7. Designed specifically to capture modern mobile visual trends (e.g., social media portrait layouts), spanning both Text-to-Image (T2I) and Image-to-Image (I2I) generation paradigms.
  • Aspect Ratio Balancing: Open Images V7 portrait-oriented photographs ($18.0\%$ of sampled thumbnails $<0.8$ aspect ratio) are integrated to decouple aspect ratio structural shortcuts from the authentic class label.
  • OOD Generalization Strategy: Held-out validation and testing splits strictly isolate unseen open-weight architectures (e.g., FLUX variants, SANA) and closed-weight commercial APIs (e.g., GPT Image 2, Seedream 4.5, Ideogram 4.0). Splits are enforced at the decoder/autoencoder lineage level rather than model checkpoint names (e.g., grouping all flux2_vae lineages simultaneously).
                       Total Evaluated Corpus (N = 455,300)
                                        │
         ┌──────────────────────────────┼──────────────────────────────┐
         ▼                              ▼                              ▼
  Train Set (77.0%)             Val Set (11.8%)                Test Set (8.2%)
 Real: 175,670 | Fake: 174,894   Real: 26,800 | Fake: 26,800     Real: 18,647 | Fake: 18,648
 Total: 350,564                 Total: 53,600                  Total: 37,295

Quantitative Split Breakdown

Dataset Partition Authentic (Real) Synthetic (Fake) Aggregate Total Generative Families & Architectural Diversity
Train 175,670 174,894 350,564 SD (5), Diffusion/Pixel (5), GAN (6), PixArt (3), SD-Adaptor (3), SD-Img2Img (3), Kandinsky (3), OmniGen (2), Infinity (2), FLUX (1), Z-Image (1), Singletons (5)
Validation 26,800 26,800 53,600 VQ/Autoencoder (4), SD (2), FLUX (2), Klein (2), Open-Weight Singletons (3), Commercial Closed-Weight (2)
Test 18,647 18,648 37,295 Commercial Closed-Weight (11), Open-Weight Singletons (4), SD (1), FLUX (1), Z-Image (1)
Total 226,115 229,185 455,300 Full Architectural Cross-Section

Architectural Representation by Lineage

Generator Lineage Family Train Partition Validation Partition (Internal) Test Partition (Out-of-Distribution)
Stable Diffusion (SD) SD 1.4, SD 1.5 T2I, SD 2.1, SDXL 1.0, SDXL T2I SDXL Turbo, SD 3 Medium SD 3.5 Large
SD-Adaptor SD Adaptor ControlNet, SD Adaptor LoRA, SD Adaptor LyCORIS — —
SD-Img2Img SD 1.5 Img2Img, SDXL Self-Cond, SDXL Img2Img — —
FLUX FLUX.1 Kontext Dev FLUX.1 Kontext Dev, FLUX.1 Dev FLUX.1 Schnell
Generative Adversarial Nets DF-GAN, GALIP, StarGAN, StyleGAN, GigaGAN, BigGAN — —
Diffusion / Pixel UNet ADM, DDIM, DDPM, VQDM, Imagen — —
PixArt PixArt-512, PixArt-$\alpha$, PixArt-$\Sigma$ — —
Kandinsky Kandinsky 2, Kandinsky 3, Kandinsky 2.2 T2I — —
OmniGen OmniGen, OmniGen 2 — —
Infinity Infinity 2B, Infinity 8B — —
Klein — Klein 4B T2I, Klein 4B Ref Image —
VQ / Autoencoder — VQ-VAE, VQ-GAN, MAGE, MAE —
Z-Image Z-Image T2I — Z-Image Turbo
Open-Weight Singletons YOSO, Kolors, Janus Pro 7B, Ovis Image, DeepFloyd IF Playground v2.5, Lumina Image 2.0, Qwen Image Würstchen T2I, Sana 1600M T2I, HiDream, Ideogram 4.0 Turbo
Commercial Closed-Weight — Ideogram v3 Turbo†, ImageGen-4 Fast† Nano Banana Pro†, FLUX-2 Max†, ImageGen-4 Ultra†, Seedream 5 Lite†, Groq Imagine Image†, Google Nano Banana 2 Lite, ByteDance Seedream 4.5, OpenAI GPT Image 2, Meta Muse Image

Note: Models marked with (†) indicate preliminary proprietary evaluation suites.


Core System Architecture & Operational Pipeline

Dataset Ingest
     │
     ▼
Canonicalisation (200px Crop, strip short-side shortcuts)
     │
     ▼
Augmentation (11 Stochastic View Pipelines, 1-3 operations chained)
     │
     ▼
Dual Parallel Towers (DINO ViT-L @ 224px, 608M combined parameters)
     │
     ▼
Concatenation Head (2048-d Feature Vector -> 512-d MLP)
     │
     ▼
Loss Regularization (Class BCE + Aux Degradation + TeleAI KL/MSE Consistency)
     │
     ▼
Calibrated Probability Output & EQI Dynamic Temperature Scaling

1. Canonicalisation & Resolution Shortcut Stripping

Short-side pixel dimensions leak class labels due to source dataset encoding biases (e.g., COCO real photographs exhibiting short-side dimensions of 200px versus high-resolution synthetic renders). Canonicalisation enforces strict input standardization prior to feature extraction.

  • Window Dimensions: CANON_CROP_SIDE = 200 native pixel window, matching the bandwidth floor of authentic benchmark classes. Detail above 200px is omitted to prevent the network from learning high-frequency resolution artifacts as binary class indicators.
  • Crop vs. Band Resampling: Native cropping preserves high-frequency generator artifacts and pixel-grid structures. It provides measured gains for DINO-lineage ViTs (+0.079 on dinov2l) without introducing resampling grid interference.

2. Stochastic Multi-View Augmentation Pipeline

Training views ($N_{\text{views}} = 11$; 1 clean, 10 degraded) apply continuous, non-linear augmentation chains.

  • Compound Recipes: Each degraded view applies 1 to 3 chained operations drawn from 6 distinct transform families (JPEG compression, Gaussian blur, spatial resizing, additive noise, color jitter, and center cropping).
  • Withheld Severity Bands: To evaluate generalization against unseen degradation severities, specific operational parameters (e.g., JPEG quality $q \in [65, 75]$ and Gaussian blur $\sigma \in [0.85, 1.15]$) are strictly withheld during training.

3. Dual-Tower Vision Backbone

The primary extraction pipeline comprises two parallel 300M-parameter DINO-lineage Vision Transformer (ViT-L/14) towers operating at 224px native input resolution.

  • Parameter Topology: Complete unfreezing across all 24 Transformer blocks (--depth 24). Master weights are maintained in float32 precision with bfloat16 mixed-precision execution.
  • Optimization Differential: Tower learning rates are decoupled from downstream head parameters ($\text{LR}{\text{tower}} = 1\times 10^{-5}$ vs. $\text{LR}{\text{head}} = 1\times 10^{-3}$) to preserve foundational spatial representations.

4. Classification Head & Multi-Task Consistency Loss

Embeddings from both backbones are concatenated into a 2048-dimensional joint vector and fed into a 1M-parameter MLP head (Linear → GELU → LayerNorm(512) → Linear → GELU → Linear(1)). Training is regularized via a tri-part multi-task loss structure:

$$\mathcal{L}{\text{total}} = \mathcal{L}{\text{BCE}} + 0.3 \cdot \mathcal{L}{\text{degradation}} + 1.0 \cdot \text{KL}\left(P{\text{clean}} \parallel P_{\text{degraded}}\right) + 1.0 \cdot \text{MSE}\left(h_{\text{clean}}, h_{\text{degraded}}\right)$$

  • Asymmetric Reference Target: The clean feature branch is detached during backpropagation, establishing a fixed reference state to pull degraded feature representations toward without risking representation collapse.
  • Auxiliary Degradation Head: Predicts degradation presence and severity parameters to output a continuous degradation embedding vector ($d$).

5. Calibration & Abstention Policy

Inference delivers calibrated probability scores $P(\text{AI-generated}) \in [0, 1]$ mapped to a strict target operational threshold ($\text{FPR} = 0.01$).

  • Dynamic Temperature Scaling: Incorporates predicted degradation vectors ($d$) and handcrafted proxy metrics ($h$) to dynamically adjust logit scaling:

$$T(d, h) = \text{softplus}\left(\text{Linear}([d; h])\right)$$

Under severe input corruption, confidence scores automatically soften toward $0.5$.

  • Evidence Quality Index (EQI): Estimates residual forensic evidence within input images to support a three-way decision output: Clear (Authentic), Flag (Synthetic), or Review (High Uncertainty / Excessively Degraded).

Key Systemic Takeaways & Architectural Decisions

System Metric / Component Selection Rationale Empirical Justification & Evidence
Lineage-Based Data Splitting Grouped by generator family and VAE decoder lineage rather than model name. Prevents model memorization of shared latent spaces; holding out klein4b_t2i alongside klein4b_ref_image measures actual structural generalizability (splits.py).
200px Native Spatial Crop Standardizes input resolution prior to transform execution. Resolves resolution shortcuts where short-side dimensions yielded standard synthetic classification AUC of 1.0000 on baseline distributions (docs/resolution_shortcut.md).
Multi-Task Loss Regularization Integrates KL prediction divergence and MSE hidden-state consistency. Reduces calibration drift from 1.24 pts down to 0.37 pts (a 3.4× reduction), maintaining threshold stability across degraded deployment environments (docs/robustness_table.md).
Dynamic Temperature Scaling ($T$) Rescales raw classification logits using auxiliary degradation vectors ($d$). Mitigates overconfident false alarms on corrupted real images by softening probability outputs toward maximum uncertainty under high noise levels.
Dual ViT-L Parallel Architecture Concatenates two 300M DINO ViT-L towers at 224px input resolution. Provides high out-of-distribution transfer metrics (Transfer Test $\text{AUC} = 0.9961$; $\text{TPR @ 1\% FPR} = 0.9823$). Note: Evaluation warrants fine-tuned comparison against single-tower multi-policy baselines.

Built With

  • dinosv2
  • huggingface
Share this project:

Updates

Submission history