Inspiration

Modern recruitment is broken. When a fast-growing startup posts a tech role, recruiters are flooded with hundreds of resumes within hours. But there is a silent crisis nobody talks about: 10% to 20% of submitted resumes contain inflated experience, date overlaps, prompt injections, or ghost companies. Traditional Keyword ATS systems actually amplify this problem—they assign top scores to keyword-stuffed, exaggerated resumes while burying honest, qualified candidates.

We built RecruitShield AI to solve this exact problem: an autonomous, fraud-resistant candidate screening and ranking co-pilot that neutralizes honeypot resumes and elevates genuine talent.


What it does

RecruitShield AI transforms the hiring workflow through four core capabilities:

  1. 5-Point Fraud Firewall (Anomaly Detection): Automatically audits candidate timelines, career progression, and date overlaps to detect experience inflation, impossible career duration claims, ghost companies, and hidden prompt injections.
  2. Dense Vector Semantic Ranking: Uses 768-dimensional dense vector embeddings to match candidates against job descriptions based on true skill semantics rather than surface-level keyword matching.
  3. AI Recruiter Co-Pilot (Gemini 2.5 Flash): An interactive assistant that allows recruiters to query candidate background evidence, verify skill flags, and inspect candidate availability in natural language.
  4. Agent Telemetry & Audit Console: Displays transparent execution logs (/agent_logs) showing exactly how anomaly penalties were calculated and why a candidate was ranked.
  5. One-Click HR Export: Seamlessly exports verified candidate shortlists to Excel/CSV.

How we built it

We engineered RecruitShield AI as a decoupled, high-performance full-stack application:

  • Frontend: React 19, TypeScript, Vite, and Tailwind CSS for a modern, responsive co-pilot dashboard (deployed on Vercel).
  • Backend API: Python 3.11 with FastAPI and Uvicorn for asynchronous microservice endpoints (deployed on Render).
  • Semantic Embedding Engine: Powered by BAAI/bge-base-en-v1.5 (sentence-transformers) generating 768-dim embeddings for cosine similarity calculations.
  • LLM Intelligence: Google Gemini 2.5 Flash API for real-time natural language recruiter query handling with smart local fallback safeguards.

Challenges we faced

  • Embedding Model Cold Starts: The BAAI/bge-base-en-v1.5 sentence-transformer model took 30–60 seconds to load on Render's free tier, causing API cold starts to time out. We fixed this by refactoring model initialization into a non-blocking background daemon thread upon server startup.
  • Score Inflation Bug: Early iterations of our hybrid scoring algorithm allowed skill scores to exceed 1.0 (100%) because proficiency multipliers stacked on top of each other. We resolved this by mathematical capping and normalizing sub-scores within a strict $[0.0, 1.0]$ range.
  • Building the Honeypot Firewall: Writing anomaly rules that actually work meant researching real founding years for 40+ real tech companies to flag impossible employment timelines (e.g., claiming 10 years of Kubernetes experience at a company founded 3 years ago).
  • LLM Rate Limit Resilience: Safeguarding our recruiter co-pilot chat against 429 rate limits by engineering an autonomous Python fallback mechanism to keep candidate querying 100% functional.

Accomplishments that we're proud of

  • 100% Fraud Neutralization: Built a 5-point anomaly firewall that successfully unmasks synthetic candidate resumes, hidden prompt injections, and inflated career timelines without penalizing legitimate top talent.
  • Blazing Fast Vector Search (<300ms): Achieved sub-second semantic search, candidate ranking, and anomaly auditing across complex applicant datasets by optimizing matrix embedding operations in NumPy.
  • Resilient Recruiter AI Co-Pilot: Engineered a Gemini-powered conversational co-pilot with automatic local fallback protection, ensuring zero service disruption even under heavy LLM API rate limits.
  • Full Transparency & Auditability: Built a dedicated Agent Telemetry Console (/agent_logs) that exposes every calculation, penalty, and reasoning step behind AI candidate rankings.
  • Production-Grade Live Deployment: Shipped a clean, responsive full-stack platform live on Vercel and Render complete with interactive OpenAPI documentation.

What we learned

  • Why Deterministic Rules Beat Pure LLMs: LLMs can hallucinate, but rule-based firewalls never do. Our honeypot detection works 100% reliably because it relies on deterministic math and real timeline validation rather than AI guessing.
  • How Semantic Embeddings Actually Work: Learned firsthand why 768-dimensional cosine similarity using BAAI/bge-base-en-v1.5 beats traditional keyword matching—it understands true skill context and meaning, not just word overlap.
  • FastAPI + React Full-Stack Integration: Built our first full-stack AI product from scratch—handling CORS, async server startup, binary vector array caching, and live REST API integration between React 19 and Python FastAPI.
  • Designing Explainable AI for HR: Discovered that recruiters don't want a "black box" AI. Exposing telemetry audit logs (/agent_logs) builds user trust by showing exact score breakdowns and anomaly flags.

What's next for RecruitShield AI

  • Document OCR Parsing: Direct PDF/DOCX resume ingestion engine.
  • Automated Interview Generator: Dynamically generating interview questions tailored to investigate specific anomaly flags found in a resume.
  • ATS Integrations: Native webhooks for Greenhouse, Lever, and Workday.

Built With

Share this project:

Updates

Submission history