-
The game sets you working late hours at a ledger for 6 days
-
Approve or hold transactions based on the given info
-
Incorrect approvals or holds can cost you and your company money
-
Chat real time with NPC customers moving money, powered by AI
-
Dont give away account info EVER
-
After 6 days you get an overview tailored to what you did wrong and how you can work to not fall for scams
Inspiration
Every school in the country now has to teach digital safety. They run the assembly, they hand out the slideshow, everyone nods — and then nobody has the faintest idea whether a single student would actually catch a scam if one landed in their inbox on a Tuesday night.
That gap bothered us. Awareness training measures attendance, not resistance. You can sit through the whole thing and still hand over a 2FA code to someone who sounds important and impatient, because knowing that "phishing exists" is not the same skill as noticing that the person messaging you right now is applying pressure.
So we stopped trying to teach the lesson and built something that measures it instead.
What it does
FLAGGED puts you at a desk on the night shift as a payments analyst at Meridian. Cases come in through two channels — an INBOX of email and a LEDGER of pending transactions — and each one asks you for a single decision: approve or hold.
You have tools. SENTRY gives you session evidence: device IDs, login locations, file access history, timestamps. RELAY is your phone, where the person behind the request will happily talk to you — and where you can ask them anything you want.
That is the actual game. You interrogate them. You ask the contractor what city they are in, then check SENTRY and find their session came from somewhere else. You ask the "new person in legal ops" a question only a real employee could answer. You ask why it has to happen tonight. Some of them are legitimate and you are wasting a real vendor's time. Some of them are not, and the contradiction is sitting right there in the evidence panel if you thought to look.
Approve a fraudulent case and money leaves. Hold a legitimate one and you have blocked your own company. Six days, and the balance follows you the whole way.
The part that matters comes at the end. Every decision you made is tagged with the manipulation tactic that was being used on you — not just whether you got it right, but what specifically got past you. Urgency works on you but authority doesn't. You verify identity but never question the timeline. That per-tactic breakdown is the real output: a vulnerability map, for one specific person, generated from behaviour instead of a quiz.
How we built it
Rules are data, and the generator verifies its own work. Cases are not hand-written scripts. Each one is assembled from a rule set, and the generator runs a solvability check on every case before it can ship — it confirms that the evidence available to the player is actually sufficient to reach the correct verdict. If a case cannot be solved from what the player can see, it never gets served. No unwinnable cases, no guessing, no "gotcha" difficulty.
NPCs hold their story under adversarial questioning. The characters on RELAY are conversational, which means players will absolutely try to break them — and teenage playtesters try hard. They stay in character. They do not leak their instructions, do not confess when accused, and do not fall out of the fiction when someone types something hostile or absurd at them. An attacker who folds the moment you ask "are you a scammer?" teaches the wrong lesson entirely, because that is not how it works in real life.
Tactic balance was measured, not guessed. We ran roughly 500 simulated playthroughs to check the distribution of outcomes across tactics, found the ones that were disproportionately easy or hard to catch, and tuned them. This matters more than it sounds: if one tactic is trivially detectable, the vulnerability map at the end reports a strength the player does not actually have. The measurement is only worth something if the difficulty is even.
The interface is the art. Everything is drawn in a locked retro palette — dark purples and greys, amber for the interface, red for danger, green for money, cyan for evidence. Hard shadows, no gradients, no rounded corners, CRT scanlines. The aesthetic is not decoration; the deliberately cramped, text-dense terminal is what makes reading the evidence feel like work, which is exactly what it feels like at 11pm when someone is rushing you.
Challenges we ran into
Making the wrong answer feel reasonable. A scam that looks like a scam teaches nothing. The hard design work was making the fraudulent cases genuinely defensible in the moment — a plausible name, a real-looking invoice, a reason the timing is tight — so that approving one feels like a judgment call rather than a mistake. Several of our early cases were too obvious and had to be rebuilt from the pressure outward.
Conversational characters that don't break. Getting an NPC to answer freely while never stepping outside its role, never revealing its own setup, and never conceding under direct accusation took far more iteration than we expected. Every playtester found a new angle of attack.
Balancing measurement against playability. For the per-tactic breakdown to mean anything, each tactic needs enough decisions behind it to be a real signal — but the run also has to stay under the length where a player will actually finish it. Getting coverage and pacing to coexist drove the six-day structure.
Verifying solvability automatically. Writing a checker that can prove a generated case is solvable from the visible evidence was fiddly, and it caught real bugs — cases where the contradiction we thought we had planted was not actually reachable by the player.
Accomplishments that we're proud of The output is a diagnostic, not a score. Most safety games tell you that you got 7 out of 10. Ours tells you which manipulation levers work on you. That is the thing a teacher or a parent could actually act on. Every case is provably fair. The generator will not ship a case the player cannot solve. That guarantee is enforced by code, not by us remembering to check. The NPCs survived teenagers. They stayed in character through every jailbreak attempt our playtesters threw at them. We measured our own balance across 500 runs rather than shipping on vibes. It looks like a real thing. The pixel-art desk, the CRT terminal, the phone that you physically put down to use the monitor — the world sells the tension before a single line of dialogue does. What we learned
The biggest lesson was that the skill being taught is not recognition, it is resistance to pressure. Our playtesters could all define phishing. They still approved things, because the message was urgent and the sender sounded senior and holding it felt rude. Knowing the concept and holding the line under social pressure are different abilities, and only the second one protects you.
We also learned how much the interrogation changes engagement. Reading a suspicious email is passive. Getting to ask the person a question and watching their story fail to line up with the evidence is the moment it clicks — that is when playtesters visibly changed how they were playing.
And on the technical side: making a generator check its own output was the single highest-leverage thing we built. It turned a category of bug we would have shipped into a category of bug that cannot exist.
What's next for FLAGGED
Put the vulnerability map in a teacher's hands. Aggregate an anonymised class-level view: which tactics a whole cohort is weakest against, so a lesson can be aimed at the actual gap instead of the general topic. More channels. Voice notes, QR codes, and the specific shape of scams that target teenagers — marketplace deals, game-account recovery, "your parcel is held." A pre/post measurement mode, so a school can run FLAGGED before and after their existing safety unit and finally see whether it moved anything. Longer-horizon cases that span multiple days, where a relationship is built early and cashed in later — because that is how the expensive ones actually work. Accessibility pass: screen-reader support and a high-contrast alternative to the locked palette.
Built With
- css
- education
- game-design
- javascript
- llm
- pixel-art
- procedural-generation
- react
- security
- vercel
- vite
Log in or sign up for Devpost to join the conversation.