💡 Inspiration In an era dominated by rapid advancements in generative AI, hyper-realistic deepfake videos and synthetic diffusion images are proliferating across social media, political campaigns, and news outlets. Traditional media verification tools rely on single binary classification models or superficial metadata checks—both of which easily fail against modern generative models like Midjourney, FLUX, and Sora.

We were inspired to build TriNetra AI ("The Three Eyes") to restore trust in digital media through multimodal verification and biological proof of life. We realized that while generative AI can mimic superficial surface pixels, it cannot simulate microscopic human biological signals like cardiac blood flow variations (rPPG), nor can it eliminate spatial diffusion noise residuals or match perceptual fingerprint histories across the web.

👁️ What it does TriNetra AI is an evidence-backed, zero-simulation digital media forensic platform that analyzes uploaded images and videos to determine authentic vs. AI-generated/manipulated media.

Our engine operates through The Three Intelligence Eyes:

EYE 1 — Visual AI Detection: Runs real PyTorch deep learning inference (ResNet18 for facial image crops and EfficientNet for video frames) alongside deterministic Error Level Analysis (ELA) and Diffusion Noise Residual heatmaps to detect spatial frequency grid anomalies. EYE 2 — Physiological rPPG Analysis: Extracts microscopic cardiac blood volume pulse signals from facial skin Regions of Interest (forehead and cheeks). Using Plane-Orthogonal-to-Skin (POS) and CHROM algorithms, Eye 2 measures estimated Heart Rate (BPM) and Pulse SNR. AI deepfakes fail to reproduce coherent biological blood flow! EYE 3 — Media Fingerprinting & Traceability: Computes streaming SHA-256 checksums, DCT perceptual hashes, and 512-dimensional PyTorch visual embeddings to trace child-parent media propagation trees ($A \rightarrow B \rightarrow C$). All evidence is mathematically combined using calibrated Log-Odds fusion into an exact, reproducible confidence score, complete with an interactive TRINETRA Voice Assistant that reads out spoken forensic reports.

🛠️ How we built it Backend & ML Inference: Built with Python & FastAPI, running real PyTorch model forward passes, OpenCV image analysis, SciPy/NumPy signal processing (Butterworth filtering, FFT Power Spectral Density), and SQLAlchemy database storage. Frontend & UI/UX: Built with React, TypeScript, and Vite, styled with Vanilla CSS featuring a futuristic cyber grid background, glowing high-contrast action buttons, atmospheric background text watermarks, and Lucide React icons. Voice Narrative Assistant: Integrated the browser's native Web Speech API (SpeechSynthesis) for real-time speech report narration with animated audio equalizer visualizers. Testing & Quality Assurance: 100% passing automated test coverage using PyTest covering end-to-end tensor inference, quality gating, and manifest verification. 🚨 Challenges we ran into rPPG Signal Noise & Lighting Sensitivity: Extracting sub-visual cardiac pulse signals from facial video skin ROIs required filtering out ambient lighting changes and head motion noise. We solved this by implementing temporal Butterworth bandpass filtering (0.75 Hz - 3.5 Hz) and multi-ROI spatial averaging across forehead and cheek patches. Calibrating Multimodal Evidence Fusion: Combining probabilistic deep learning scores, deterministic signal SNR values, and compression error levels into a single coherent score without overestimating confidence was tricky. We implemented Log-Odds calibration with a strict Technical Quality Gate Filter (Laplacian blur variance and face coverage limits) to prevent false positives on blurry media. Zero-Simulation Real-Time Execution: Ensuring that PyTorch tensor forward passes and OpenCV signal extractions ran fast enough for interactive live web execution required optimizing tensor normalization and frame extraction pipelines.

🏆 Accomplishments that we're proud of Strict No-Simulation Guarantee: Every single confidence score and metric displayed in TriNetra AI is derived live from actual PyTorch model forward passes and mathematical signal processing—ZERO fake or mock data exists in the codebase. Biological Proof of Life: Successfully demonstrated that biological heart-rate extraction (rPPG) can serve as a robust physical barrier against facial deepfakes. Interactive Voice Forensic Assistant: Built a hands-free voice synthesis reporter that reads complex forensic verdicts and probabilities out loud. 100% Automated Test Passing Rate: All backend inference, model registry, and dataset manifest pipelines pass clean automated unit tests.

📚 What we learned Multimodal beats single-model AI: Relying on one deep learning model is insufficient; combining deep learning logits with physical biological signals (rPPG) and spatial noise residuals provides vastly superior forensic resilience. Quality Gating is Essential: High-accuracy models can give misleading results on low-resolution or heavily compressed images. Implementing quality checks (blur index, contrast, face coverage) is essential before passing data to ML models. Signal Processing + Deep Learning Synergy: Combining classical DSP (Butterworth filters, FFT spectral analysis) with modern PyTorch neural networks bridges the gap between interpretability and raw predictive power.

🚀 What's next for TriNetra AI Audio Deepfake & Voiceprint Analysis: Expanding Eye 2 to analyze audio tracks in videos, detecting synthetic speech artifacts using mel-spectrogram transformers and acoustic feature extraction. Browser Extension for Real-Time Fact-Checking: Developing a Chrome/Firefox extension that lets users right-click any image or video online to run instant TriNetra forensic analysis. Decentralized Forensic Verification Ledger: Anchoring perceptual hash fingerprints onto an immutable blockchain ledger to create tamper-proof media provenance chains for news organizations and legal authorities. Real-time Video Stream Forensics: Optimizing the rPPG and frame inference pipeline to support live video stream analysis (e.g., Zoom/Meet video call verification).

Share this project:

Updates

Submission history