Deepfakes have made it easy for a face swap, a cloned voice or a pre-recorded clip to pass for a real person on a live video call. Remote hiring and onboarding depend on trusting who is on the other side of the screen, and interviewers can't reliably tell. Defending against this today means juggling separate tools for video, voice and liveness, so we built Sapien: one API that answers one question: Is this person real?
Sapien checks whether a live video and voice session is human or synthetic. The candidate takes short live challenges throughout the call, following prompts on camera, while Sapien checks whether their face and head movements look natural, whether the video frames show signs of a deepfake, and whether their voice sounds human and matches the words they were asked to say. A decision engine combines those checks into a single real-or-synthetic answer, shown in an operator console in real time.
Under the hood, the frontend is a React app that runs MediaPipe Face Landmarker in the browser via WebAssembly to track the face, and records the audio and sampled video frames. The backend is a FastAPI service that routes each input to its own check behind one API. Frames are cropped to the face with OpenCV and scored by a pretrained Vision Transformer classifier, and we average the fake probability over all frames in the challenge window so one bad frame can't decide the verdict. Voice authenticity uses a model built on wav2vec2 audio embeddings, and Whisper transcribes the audio to verify the candidate said the prompted words. Every model is pretrained and each check is kept separate from the API routes so any model can be swapped without rewriting the system.
The hardest part was that not all detectors are equal. We tested 12 pretrained deepfake models across three datasets, and some scored almost perfectly on one while being no better than guessing on another. Video was harder than photos too: on compressed video frames the classifier called everything real until we cropped each frame to the face, but cropping then hurt photos, so no single setting worked everywhere. We also had only a handful of our own real and fake samples, and the face-cropped setup flagged some real photos as fake. Finally, because of resource constraints, the hosted version of Sapien is limited in capabilities, since free hosting couldn't run all of these models at once, and the full system runs locally.
We're proud of the working end-to-end demo, and of the fair comparison of 12 models, where we held out data before picking anything so we wouldn't fool ourselves with a lucky score. We also wrote a reusable test script that measures accuracy, false alarms and misses on any folder of real and fake images, and we were honest about what we haven't proven instead of claiming an accuracy number we can't back up. The biggest lesson is that a detector that looks near-perfect on a benchmark can fall apart on your own data, and that a deepfake check works best as one signal among several, not the only judge.
Next, we'll record real webcam sessions with real users and simulated attacks and set the decision thresholds from that data. We also want to train our own detector on footage that matches live video, add image and text authenticity checks, and pilot Sapien with a few platforms that hire or onboard remotely, using the single API as the integration point.
Built With
- fastapi
- huggingface
- mediapipe
- react
- tanstack
Log in or sign up for Devpost to join the conversation.