Inspiration
A resume has never been cheaper to write, and never been worth less as evidence. Anyone can now produce a fluent, specific-sounding account of work they did not do - and the entire screening stack downstream rewards exactly that, because keyword matching cannot tell the difference between having done something and being able to describe it.
One asymmetry survived: writing the claim became free, and having been there did not. The person who actually ran the pipeline remembers what broke in week one, which page went to the wrong on-call rotation, the number they can't fully defend. That detail is expensive to fake and free to recall. We built an interviewer that goes looking for it.
What it does
VeriScreen is an AI video interviewer that checks whether a resume is true, then ranks candidates on what they demonstrated instead of what they claimed.
- A job description becomes a weighted requirement checklist — capabilities, not keywords ("runs streaming systems in production," not "Kafka").
- A resume is split into atomic, falsifiable claims, each rated for materiality (does this matter for this job) and checkability (can talking to someone verify it). "Strong communicator" scores 1 on checkability and is never worth a question.
- The candidate answers five spoken questions on video, about four minutes. The interviewer asks about mechanism, decisions and failures — never for a recital of the resume, because anyone can read one back.
- A separate audit pass — different model, different prompt, run after the call — rules on each claim and must quote the transcript verbatim to justify it.
- Three numbers come out: veracity (of what we checked, how much held up), coverage (how much of the resume five questions reached), and verified fit (requirement coverage, gated by whether the claim behind it survived).
The load-bearing move is that last one: fit is relevance multiplied by the veracity of the backing claim. Claiming more cannot lift your score — it only creates more for the interview to dig into. The second is that "we never asked" is excluded from the average entirely rather than scored as a zero. Conflating we didn't check with they failed is where this kind of score stops being honest.
On top of that: ranking with a veracity floor, and pool search across both interview answers and resume claims, always labelled which source a hit came from.
How we built it
FastAPI + MongoDB + React/Vite, the whole stack in Docker Compose behind one make demo.
- Claude Sonnet 5 drives each interview turn — warm and fast, in the room.
- Claude Opus 5 runs the audit — slow and strict, after the fact. Deliberately a different model and prompt: the thing that has to be pleasant in conversation should not be the thing deciding whether someone is telling the truth.
- Anthropic structured outputs into Pydantic models, so every turn returns a typed object rather than parsed prose.
- Deepgram nova-3 for speech-to-text and aura-2 for the spoken questions.
- Voyage AI voyage-3.5 embeddings and rerank-2.5 for requirement-to-evidence matching and pool search.
scoring.pyis plain Python — a lookup table and three weighted averages. No model ever emits a score. That makes the whole thing reproducible and re-runnable without re-interviewing anyone.
Challenges we ran into
The interview kept eating itself. With a cap of three questions per claim, one stubborn claim consumed 60% of a live test interview and coverage collapsed to 17%. Dropping the cap to two guarantees at least three distinct claims get touched. We also had to add an explicit no-repeat rule — the model would re-ask the same question in new words, which reads as interrogation and yields nothing new.
The first live audit died on a rate limit. Reranking once per requirement tripped Voyage's free-tier limit of three requests a minute on a seven-requirement JD, and took the entire audit down with it. We rewrote fit scoring as cosine similarity over vectors embedded once at job creation and once per audit — scoring a candidate now costs zero extra API calls — and added a retry plus a graceful fallback so a throttle degrades the score instead of destroying it.
Relevance doesn't start at zero. The reranker returns ~0.35 for genuinely unrelated pairs, which would hand every requirement a third of its weight for evidence that does not exist. We map the useful band onto 0–1 before it reaches the score.
Making the demo work with no API keys at all. The seed writes claims, turns and verdicts directly, so every screen is live before the agent is — and a full reset takes ten seconds. That mattered more than it sounds like on a hackathon clock.
Accomplishments that we're proud of
Two near-identical candidates for the same role, one lying about one thing:
| Priya | Daniel | |
|---|---|---|
| Veracity | 0.82 | 0.18 |
| Verified fit | 0.60 | 0.11 |
Daniel's resume is the stronger of the two, and his raw relevance scores are as high as Priya's - his answers genuinely do address the requirements. It is the veracity gate that empties them, which is precisely what a keyword matcher cannot do.
The contradiction that sinks him is never mentioned in the room. Question 1: "I led the design end to end." Question 4, approached from a completely different angle: "I actually found out about it at the architecture review — the platform team had already evaluated both." The interviewer just asks, and the audit finds it. Tipping the candidate off would destroy the signal.
And every number traces back to a sentence — click a claim on the scorecard and the evidence panel jumps to the answer that decided it. There is no unexplained figure anywhere in the product.
What we learned
The model was not the hard part. The honesty of the metric was. Three decisions did most of the work: excluding "never asked" from the average, making coverage a first-class number and abstaining entirely below 25% rather than implying we checked more than we did, and keeping all arithmetic in code so that changing the rubric re-ranks the archive instead of invalidating it.
The finding that surprised us most: fluency is not evidence. Confidence, structure, enthusiasm and correct textbook knowledge are exactly what a well-read person produces without having done the work — it is the single most common way this judgement goes wrong, so the auditor is instructed against each of them by name. In the same breath it's told not to penalise nerves, filler words, short sentences or non-native phrasing, which are not signal either.
What's next for VeriScreen
- Confidence intervals. Five fixed verdict points are a stand-in for self-consistency sampling and proper variance on the score.
- Integrity signals. Every turn already stores
asked_at,answered_atand speech duration although nothing reads them yet — a long silence followed by fluent, dense speech is the second-screen signature, and none of it can be backfilled later. - Auth and tenancy on the admin API, plus
$vectorSearchon Atlas in place of brute-force numpy. - Calibration against real hiring outcomes. The number that ultimately matters is whether veracity predicts anything downstream — and that takes a real pipeline and real hires to find out.
A few notes on what I did and didn't claim:
- The "60% of the interview / coverage fell to 17%" and the Voyage 3-req/min failure are both real events from this build, documented in config.py:26 and the README. Those specifics are what make a Devpost story read as built rather than pitched — I'd keep them.
- I left out the /health proxy bug and the secure-context camera limitation, since neither is a story beat.
- One thing changed under me while I worked: scoring.py now has REL_FLOOR/REL_CEIL rescaling that wasn't there when I read it for the artifact earlier. It doesn't affect the headline demo numbers, which still match the README, but the scoring artifact doesn't mention the rescaling step. Say the word and I'll add it.
Built With
- fastapi
- mongodb
- react/vite
Log in or sign up for Devpost to join the conversation.