Inspiration
Most AI tools aimed at scams and social engineering do the same thing: classify a message as suspicious and block it. That's useful, but it's also where every other AI/ML hackathon submission in this space stops — a binary detector bolted onto an LLM call. Coming from a background in offensive security and cybersecurity competitions, the more interesting question felt like: what if, instead of just blocking the attacker, you let them think they're winning? A honeypot that talks back, stalls for time, and quietly pulls information out of the conversation turns every scam attempt from a dead end into usable threat intelligence — and that's a problem AI is genuinely well-suited to, not just applied to for novelty's sake.
What it does
MIRAGE deploys AI "employee" personas that engage incoming social-engineering attempts in realistic conversation — staying in character, stalling, and asking questions designed to draw out identifying details, without ever revealing they're AI. Every conversation is run through a second extraction pass that converts it into a structured incident record: the MITRE ATT&CK technique used, any indicators of compromise, and the manipulation tactics employed (urgency, authority, scarcity). Each record is embedded and compared against every past incident, so MIRAGE can recognize when two superficially different attacks are actually the same attacker or the same script reused across targets — and surface that connection on a live dashboard. A built-in Red-Team Simulator lets a second AI play a scripted attacker, so the full pipeline runs end to end live, without needing a real scammer in the room.
How we built it
The backend is Python/FastAPI with a Postgres database for personas, attacker scripts, incidents, messages, and extracted profiles. The Claude API drives three distinct jobs in the pipeline: generating the in-character persona replies, structured extraction of the MITRE technique/IOCs/tactics from a transcript (constrained to fixed vocabularies and validated against a strict schema), and producing embeddings used for similarity clustering across incidents. A React dashboard shows live incidents, the extracted profile and linked attacker clusters, and one-click incident report generation. The frontend is deployed on Vercel; the backend and Postgres instance run on Render.
Challenges we ran into
Keeping the persona believably "in character" while still getting clean, structured data out of the conversation turned out to be two competing goals — a persona optimized purely for realism tends to produce messy, inconsistent transcripts that are hard to extract from reliably. We addressed this by splitting the two jobs into separate LLM calls with very different settings: a higher-temperature, loosely constrained call for the conversational persona, and a low-temperature, schema-constrained call for extraction. The other real challenge was resisting the urge to claim a high accuracy number for the extraction pipeline without evidence — there's no single "correct" answer for open-ended extraction from conversation, so instead of an unverifiable accuracy claim, we built a small hand-labeled eval set and a fallback path (needs_review) for when the model's output doesn't pass validation, rather than silently guessing.
Accomplishments that we're proud of
Getting a three-stage AI pipeline — generation, structured extraction, and embedding-based clustering — working end to end inside a hackathon timeline, instead of stopping at a single LLM call wrapped in a chat UI. We're especially proud of the decision to report a real, measured eval score alongside the system rather than an inflated accuracy claim, and to build a visible "flagged for review" state instead of ever fabricating a confident-looking wrong answer.
What we learned
Building reliable, structured output on top of a generative model is a genuinely different design problem from building a good conversational agent — the two pull in different directions, and the fix is architectural (separate calls, separate settings, separate validation) rather than a better prompt. We also came away with a sharper sense of how to talk about AI system "accuracy" honestly: a small, disclosed eval score with a visible fallback for low-confidence cases is more credible, and more useful, than any single headline percentage.
What's next for MIRAGE
A voice mode using Whisper and ElevenLabs so personas can handle live vishing (voice phishing) calls, not just text. Direct integration with a real phishing-reporting inbox, so a forwarded scam email automatically launches a decoy engagement instead of just getting marked as spam. Expanding MITRE technique coverage and moving from in-memory cosine similarity to a proper vector store as incident volume grows. Longer term, the goal is a deployable deception layer that security teams can plug into their existing reporting workflow — every scam attempt becomes a live decoy engagement and a piece of connected threat intelligence instead of a dead end.
Built With
- anthropic-claude-api
- fastapi
- git
- github
- openai
- postgresql
- pydantic
- python
- react
- render
- sentence-transformers
- sqlalchemy
- tailwind-css
- typescript
- vercel
- vite
Log in or sign up for Devpost to join the conversation.