Inspiration

The best way to learn is by doing. Most security training is a slideshow you click through and forget. Social-engineering attacks such as vishing (voice phishing) don't target your computer. They target your instincts: a friendly "IT person" with an urgent problem and a reason you have to act now. We wanted people to experience that pressure in a safe setting, so they recognize it when a real one happens.

What it does

How we built it

We built this platform by combining a Python backend with a clean HTML/CSS frontend, utilizing Supabase for database management and user authentication. To power the simulated attacks, we integrated the Twilio API for automated communication and ElevenLabs for realistic AI voice generation.

  • Phone calls: Twilio sends live call audio over WebSockets to our Python (FastAPI) server.
  • Speech-to-text: Deepgram streaming transcription, which also detects when the person has finished talking.
  • The brain: an OpenAI GPT-4.1-mini agent. Each turn it gets the scenario (who it's playing, the rules, which tactics to try) and the live conversation state. It streams its reply and calls tools to log red flags and recommend training.
  • Voice: ElevenLabs streaming text-to-speech. Each sentence starts playing as soon as it's ready, so replies feel quick. Interruptions: if you talk over the agent, it stops immediately and remembers only what you actually heard. Training report: a scoring module turns the tactics you reacted to into a risk level and a short training plan.
  • Front end: a JavaScript web page plus a live dashboard showing the transcript, tactics, and results.
  • 60+ automated tests, a fake-call replay mode for rehearsing without phone credits, and Docker/Fly.io deploy scripts. ## Challenges we ran into
  • working out when someone has really finished speaking, handling pauses mid-sentence, and stopping the agent cleanly when it gets interrupted. -chaining speech-to-text, the LLM, and text-to-speech fast enough to feel like a real phone call.
  • making the simulation convincing while guaranteeing it never stores real passwords or codes, and always discloses the drill.
  • consent requirements, Twilio trial limits, and SMS carrier registration (A2P 10DLC) that takes days to approve.
  • wiring several paid APIs together under hackathon time pressure.

Accomplishments that we're proud of

  • A real phone-based AI agent that holds an unscripted conversation, not a menu of canned lines.
  • It turns a stressful "gotcha" into a teachable moment with specific, personal feedback.
  • Ethical guardrails built into the code: consent required, a clear disclosure at the end, no real secrets ever stored, and an opt-out at any time.
  • A working, tested end-to-end pipeline built in under 48 hours.

What we learned

  • How social engineering works on people: authority, urgency, and fear beat technical defenses.
  • How real-time voice AI works: streaming audio, turn detection, interruptions, and keeping replies fast.
  • That responsible security tools need consent, transparency, and data minimization from day one.
  • That the most effective training is an experience you remember.

What's next for HackMe

More scenarios: fake bank fraud alerts, CEO gift-card scams, and package-delivery texts. Text (smishing) simulations once SMS registration is approved. An organization dashboard to run consented campaigns and track improvement over time. Multi-language support and different voices for each scenario.

Built With

Share this project:

Updates

Submission history