Inspiration
Generative AI is making it easier than ever to create realistic synthetic images. This creates growing risks for online platforms, including misinformation, impersonation, fraud, and a loss of trust in digital content.
The problem becomes even harder when images are shared online. Compression, cropping, resizing, filters, and reposting can remove or distort the signals that traditional AI-image detectors rely on.
This inspired us to build TrueSight — a robust and explainable AI-image detection system designed for the way images are actually shared online.
Instead of relying on a single detection method, TrueSight combines multiple independent sources of evidence to answer a simple but increasingly important question:
Can I trust what I am seeing?
What it does
TrueSight uses a three-tier detection architecture to determine whether an image is authentic, AI-generated, or AI-modified.
Tier 1 — Provenance and Forensics
TrueSight first examines the original uploaded image without modifying it.
It checks for signals such as:
- C2PA Content Credentials
- AI provenance and watermark information
- Embedded metadata
- Image forensic anomalies
Verified AI provenance can provide strong evidence immediately. However, missing metadata or provenance is not automatically treated as proof that an image is real.
Tier 2 — Semantic Visual Analysis
If provenance alone is not enough, TrueSight uses a vision-language model to analyse the image for semantic inconsistencies.
It looks for clues such as:
- Unnatural lighting and shadows
- Incorrect reflections
- Anatomical inconsistencies
- Distorted text
- Repeated patterns
- Impossible geometry
- Background inconsistencies
- Perspective errors
This provides a different type of evidence from traditional pixel-based image classification.
Tier 3 — ConvNeXt Visual Classifier
TrueSight also uses an independently trained ConvNeXt-Tiny image classifier.
The model estimates:
$$ P(\text{AI-generated} \mid \text{image}) $$
The classifier operates independently from the provenance and semantic analysis layers.
This is important because TrueSight does not depend on any single detection technique.
The outputs from the different tiers are combined to produce the final TrueSight verdict, together with supporting evidence that helps users understand why the system reached its conclusion.
How we built it
TrueSight was developed primarily using Python, PyTorch, and TorchVision.
For our visual detection model, we used a pretrained ConvNeXt-Tiny architecture and replaced its original ImageNet classification layer with a binary classifier:
0— Real / Authentic1— AI-generated / AI-manipulated
Rather than training the entire network from scratch, we used transfer learning.
Training was performed in two stages:
- Classifier warm-up — the pretrained ConvNeXt backbone was initially frozen while the new binary classifier was trained.
- Selective fine-tuning — later ConvNeXt layers were unfrozen so that the model could adapt its visual features specifically for AI-image detection.
This allowed us to take advantage of pretrained visual knowledge while keeping the model practical for hackathon-scale computing resources.
Making TrueSight robust
A major focus of our project was ensuring that TrueSight could still work after images had been redistributed online.
An image uploaded to social media is rarely preserved exactly as it was originally created. Platforms may automatically resize or compress it, while users may crop, edit, screenshot, or apply filters before reposting it.
To simulate these real-world scenarios, our training pipeline includes transformations such as:
- JPEG compression
- Gaussian blur
- Image resizing
- Gaussian noise
- Colour adjustment
- Cropping
Our goal is for the model to maintain a similar prediction even after realistic transformations are applied.
Conceptually:
$$ f(x) \approx f(T(x)) $$
where:
- (x) represents the original image
- (T(x)) represents the same image after a transformation, such as compression, cropping, or resizing
- (f(\cdot)) represents the TrueSight AI-image classifier
Rather than learning fragile signals that disappear after compression or resizing, we want TrueSight to learn features that remain useful after redistribution.
Explainability
AI detection should not simply return a number without explaining where it came from.
TrueSight therefore incorporates Grad-CAM to visualise which regions of an image influenced the ConvNeXt model's prediction.
The resulting heatmap helps users understand what the classifier was focusing on when making its decision.
However, we deliberately avoid presenting Grad-CAM as definitive proof that a highlighted region is AI-generated. Instead, it serves as supporting evidence alongside provenance information and semantic analysis.
This makes TrueSight more transparent than a simple black-box AI detector.
Challenges we faced
Real-world robustness
One of our biggest challenges was ensuring that detection performance did not collapse after an image was compressed, resized, blurred, or cropped.
A detector that performs well only on clean images would have limited usefulness on real social-media platforms.
This motivated us to incorporate realistic transformations directly into our training and evaluation process.
Generalisation
Another major challenge was avoiding dataset-specific shortcuts.
An AI-image detector can achieve high accuracy on one dataset while performing poorly on images generated by models it has never encountered before.
Because of this, we focused not only on accuracy but also on:
- Balanced accuracy
- Precision
- Recall
- F1-score
- ROC-AUC
- False positives
- False negatives
- Performance after image transformations
- Cross-dataset generalisation
Combining different types of evidence
Provenance, semantic reasoning, and visual classification each have different strengths and weaknesses.
Metadata can be removed.
Provenance may be unavailable.
Visual classifiers can produce false positives.
Vision-language models can make incorrect interpretations.
Instead of allowing one weak signal to determine the entire result, we designed TrueSight so that each tier operates independently before their results are combined.
Avoiding false confidence
Perhaps the most important challenge was recognising that AI-image detection is not perfect.
TrueSight therefore avoids presenting its prediction as absolute proof of an image's origin.
Instead, it provides multiple pieces of evidence that help users make a more informed decision.
What we learned
One of the biggest lessons from building TrueSight was that AI-image detection is not simply an image-classification problem.
A reliable system needs to consider different questions:
- Where did the image come from?
- Can its provenance be verified?
- Are there semantic inconsistencies?
- Does a trained visual model detect AI-related patterns?
- Does the prediction remain stable after the image is transformed?
- What evidence supports the final result?
We also learned that high accuracy alone is not enough.
For a real-world detection system, false positives, false negatives, robustness, generalisation, calibration, and explainability are equally important.
Most importantly, we learned that AI detection should be treated as evidence aggregation rather than absolute proof.
What's next
TrueSight is currently a hackathon-scale prototype, but there are many directions we would like to explore next.
Future improvements include:
- Training on images from more AI generators
- Expanding authentic-image datasets
- Testing against additional social-media processing pipelines
- Improving cross-dataset generalisation
- Improving probability calibration
- Reducing false positives
- Improving model explainability
- Extending the system beyond images to other forms of AI-generated media
In the longer term, we envision TrueSight becoming more than an AI-image detector.
Our goal is to develop it into a digital-content trust layer that does not just tell users whether something may be AI-generated, but also explains:
What evidence supports this conclusion, and how much should I trust it?
Log in or sign up for Devpost to join the conversation.