Problem Statement
Platforms like TikTok already auto-label AI-generated content via C2PA Content Credentials and invisible watermarks, but that provenance metadata is stripped the moment content is screenshotted, re-encoded, or reposted off-platform. Once that happens, there's no metadata left to check.
Katsu is a content-based fallback layer for exactly that scenario: it doesn't rely on any embedded signal, only the pixels themselves, and it's trained specifically to keep working after the kinds of transformations that strip metadata in the first place.
How Katsu solves the Problem
Katsu is a lightweight hybrid model for detecting AI-generated images (AIGC) using visual evidence contained directly in the image pixels. Link to Demo: https://katsu-demogit-ix9aua9mfr9fv4akep7dku.streamlit.app/
The model is designed specifically for real-world redistribution, where images may be recompressed, resized, cropped, or affected by noise. Rather than relying on metadata, watermarks, or provenance signals, Katsu combines semantic visual representations with complementary image-artifact and texture features. Robustness Evaluation Summary and Error Analysis Note are both included in GitHub (Robustness Evaluation.md).
How we built it
The system uses a frozen DINOv2 ViT-S/14 backbone together with two lightweight artifact-detection branches:
- NPR residual branch — captures image reconstruction and resampling artifacts.
- Texture-statistics branch — extracts LBP and GLCM texture features.
- Importance-weighted fusion head — combines the complementary feature representations.
Only 33,842 trainable parameters are used on top of the frozen backbone.
Accomplishments that we're proud of
Strong generalization on unseen data
- Achieved 82% accuracy / 91% AUROC on self-built validation set & 84% accuracy / 91% AUROC on the reserved WildFake subset
- Matching performance on unseen data signals Katsu is learning generalisable features, not overfitting to our dataset's quirks
Robust under real-world transformations
- Stress-tested against common image transformations (cropping, resizing, blurring, and noise)
- Performance held consistent across all transformations
- Even under blur and noise, which typically degrade detector performance significantly, 87% AUROC and 78% accuracy was maintained
Extreme data efficiency
- Achieves all of the above while training on just 0.6% of the data used by standard published detectors
- Demonstrates that strong, robust deepfake detection doesn't require massive datasets
What's next
Introduce CLIP ViT-L/14 as a second backbone, or explore mitigations such as gradient checkpointing, mixed precision, or sequential backbone forward passes to improve model's ability to catch semantically implausible generations
Selectively unfreezing DINOv2's last block (small LR) as an ablation to compare results & evaluation
Resources used
Models & APIs
- DINOv2 ViT-S/14: Frozen backbone
- ResidualEncoder: Trained branch (encodes NPR residual signal)
- texture_proj: Trained branch(handcrafted LBP/GLCM texture features)
- ImportanceWeightedFusion: Trained fusion head combining all three branches\
- Kaggle (CIFAKE & self-transformed datasets)
- Hugging Face Hub (SID_set)
- ModelScope Hub (WildFake dataset)
- PyTorch Hub
Libraries & frameworks
- PyTorch, torchvision: model, training loop, transforms
- NumPy, pandas: data handling
- scikit-image (skimage.feature): LBP and GLCM texture feature extraction
- scikit-learn: evaluation metrics (accuracy, ROC-AUC, F1, confusion matrix), train/test split
- matplotlib: plotting results
- Pillow (PIL): image loading/transforms (blur, noise, brightness, JPEG-compression augmentation)
- pyarrow: reading parquet-format data
- OpenCV (cv2): demo app, for heatmap colormap/overlay rendering
- Streamlit: interactive demo UI
- huggingface_hub — demo apps' optional checkpoint-download fallback
Datasets & assets
- AIGC Detection Dataset (self-transformed dataset)
- CIFAKE
- SID_Set
- WildFake
- hybrid_checkpoint.pt: trained checkpoint (output of training, input to inference/demo)
Built With
- colab
- python


Log in or sign up for Devpost to join the conversation.