About the Project Inspiration We'd all run into a photo online at some point that made us stop and think "wait, is that even real?" But the thing that actually got us excited about this specific problem was smaller and more annoying: almost every AI-image detector we looked at only worked on perfect, untouched images. The second you screenshot something, recompress it, or crop it into a profile picture, which is basically what happens to every image the moment it goes online, a lot of these tools quietly stop working. We wanted to build something that couldn't cheat its way to a good number by only getting tested on ideal conditions. What It Does Given a folder of images, our model outputs a confidence score for each one and how likely it is to be AI-generated. The part we actually care about isn't the base detection accuracy, which plenty of models can nail on clean data. It's what happens after: we trained and tested against the exact kinds of damage images take on their way around the internet, such as recompression, blur, cropping and separately tested whether the model's knowledge holds up against an entirely new generator it never saw during training. That second test is the one most detectors quietly skip. How We Built It Centaurus is a team of four. Before we wrote a single line of code, we planned out the entire weekend hour by hour: who's doing what, when things get handed off, and what happens if something's late. It sounds a bit much for a hackathon, but this meant nobody spent Saturday night wondering what everyone else was doing. Jeny (A) handled the data side: merging CIFAKE with a subset of SID_Set for training, and deliberately keeping WildFake completely separate as a held-out test set for generalization to a generator the model had never seen. We used a frozen CLIP backbone for the model itself, Camelia (B), since CLIP already has strong general visual features baked in from massive pretraining. Instead of training a whole classifier from scratch, we only needed to train a small head on top, which was a lot more realistic given our timeframe. Training was done on Colab for GPU access, and honestly, half the early frustration wasn't the model at all but it was realizing the Colab runtime wasn't actually connected to a GPU, and separately, fighting with getting files to load correctly from the GitHub repo. C, Darsini, built the augmentation pipeline using Albumentations: resized to 224×224 for CLIP, then an 80% chance of applying exactly one distortion (JPEG compression, blur, noise, color jitter, or crop), followed by CLIP's exact normalization values. The "one distortion at a time" choice was deliberate: real-world images usually pick up one kind of degradation, not five stacked on top of each other, so testing that way felt more honest. Pranathi (D) built the shared scoring infrastructure everyone else's results ran through — one function computing accuracy and AUC the exact same way for every test, so nobody had to wonder if two numbers were even comparable. D also worked with C the night before testing to lock in the exact conditions above, then spent the back half of the weekend turning everyone's raw numbers into the actual findings you're reading now — including digging into why the model was confidently wrong on WildFake, not just that it was. The three fixed conditions we tested against: JPEG quality 30, Gaussian blur at σ=2.0, and an 80% center crop, got locked in the night before any testing started, specifically so nobody was making that call under deadline pressure the next day. What We Learned Training loss for the augmented model dropped steadily and cleanly, from 0.63 down to 0.06, a clear sign the model was actually learning something useful from the augmented data, not just memorizing noise. Once our final numbers came in, the comparison was better than we expected walking in: the augmented model beat the plain model on every single condition we tested, in both accuracy and AUC — a consistent 1-4 point gain, including on clean images. Training-time augmentation wasn't a trade-off between clean performance and robustness; it just made the model better, full stop. The real story, though, came from digging into why the model was wrong when it was wrong. On our standard test conditions, the model's mistakes were rare and mostly forgivable. On WildFake, a generator neither model had seen during training, accuracy collapsed to essentially a coin flip (50-53%). What made that genuinely unsettling, rather than just disappointing, was the confidence behind those wrong answers: over half of WildFake's errors were made with high confidence, not genuine uncertainty. The model wasn't hedging or shrugging when it saw something unfamiliar but it was making a confident, wrong call, most often labeling real human faces as AI-generated. That taught us something we hadn't appreciated going in: robustness to image transformations and robustness to unseen generators are completely different problems, and a model can ace one while quietly failing the other — and failing it with total confidence is worse than failing it with doubt. One of us also went in with a pretty different mental model of how these detectors even work such as expecting them to catch obvious tells like six fingers or blurry edges, rather than the much more statistical, less human-intuitive signals a model like this actually relies on. Challenges We Faced Most of our real challenges weren't about the model itself but they were about everything around it actually working. Colab silently not being connected to a GPU cost real time before anyone noticed. Getting files to load correctly from GitHub took longer than expected too as raw image datasets don't travel through git the way code does, so paths that looked correct on one machine didn't exist on another. On top of that, one of our early training scripts had a chain of small bugs stacked on each other: a dataset class that hadn't been imported yet, a model that was defined but never actually instantiated, images sent to GPU while the model itself stayed on CPU, and a missing image preprocessing step. None of these were hard problems on their own, but untangling five separate small breaks from one confusing error message took real patience. By the time the code was actually correct, we hit the real issue underneath it all that is the CIFAKE image files themselves weren't sitting where the code expected them to be, since datasets that size just don't fit in a GitHub repo. It was a good, if slightly humbling, reminder that in a project spanning data, models, and infrastructure, having correct code doesn't mean anything runs but the environment and the data have to agree with it too. What's Next The clearest next step is closing the gap WildFake exposed. Our model is genuinely robust to the kind of damage images take when they're shared online but near-chance, overconfident performance on an unseen generator means it hasn't learned a truly general sense of "AI-generated," just the specific generators it trained on. Training on a much broader and continuously updated set of generators is the obvious fix, but so is a more fundamental one: teaching the model to recognize when it's out of its depth. A detector that says "I'm not sure" on unfamiliar content is far more deployable than one that confidently mislabels a real person's photo, so calibration methods aimed specifically at out-of-distribution inputs feel like the highest- leverage next step, not just more training data. Beyond that, we'd want to test against stacked, compounding transformations (a screenshot that's also been recompressed and filtered, the way images actually degrade after multiple reshares), and explore whether frequency-domain features where GAN and diffusion artifacts often show up more clearly than in raw pixels help the model generalize across generators in a way pure visual features haven't managed to yet.

Built With

+ 9 more
Share this project:

Updates

Submission history