Inspiration

Scam messages cost people money every day, and every consumer-facing "is this a scam?" tool we could find has the same untested assumption: that one prompt can be simultaneously aggressive enough to catch fraud and calm enough not to flag your own bank. We didn't believe that assumption, so instead of building around it we measured it first. A strong, carefully written single prompt — given the exact same task, output schema, and safety instructions our final workflow node gets — caught every scam in our test set (11 of 11) but confidently cleared only 2 of 9 genuine messages. One of the six it hedged on was an ordinary appointment reminder. One it actively mislabelled — at 90/100 risk — was a real bank fraud alert. A detector that flags the message designed to stop fraud on your account is worse than having no detector at all: it teaches you to ignore the one alert you actually need to trust.

What it does

Tripwire takes any message — SMS, email, or WhatsApp text, in English or Bengali — and decides scam, legitimate, or genuinely uncertain, with a calibrated risk score and specific, plain-language reasons. Unlike a single-prompt classifier, it never lets urgency, a large dollar amount, or a request to reply be treated as guilt on their own — because real fraud alerts have all three, and a detector that can't tell them apart from scams isn't safe to ship.

How we built it

Five nodes instead of one prompt:

  1. Normalise (no model) — deterministic extraction of sender type, URLs, digit runs, urgency phrasing. Facts a regex can decide don't need a model call, and a fact extracted before any model sees the text can't be argued with by adversarial phrasing in the message itself.
  2. Threat evidence (fast tier, parallel) — builds the case for fraud only. Never allowed to reach a verdict.
  3. Legitimacy evidence (fast tier, parallel) — the fix for the measured failure. Asks the question a plain scam prompt never asks: what would a genuine message like this actually look like, and how much of that does this one have?
  4. Adversarial review (strong tier) — reads both evidence lists, names whichever conclusion is currently winning, and is required to argue the opposite as strongly as the facts honestly allow. This is the node aimed directly at the real-fraud-alert failure: forced to consider that a genuine alert looks exactly like an urgent, high-value message asking for a reply, instead of treating those as proof of guilt.
  5. Reconcile (strong tier) — weighs the assembled record and decides. Abstention is reserved for cases meeting two explicit conditions, not used as a default hedge on anything with a deadline in it.

Built on Featherless, which reserves concurrency rather than billing tokens — so evidence nodes 2 and 3 run on a small model (Qwen2.5-7B, 1 of 4 account slots) genuinely in parallel, while the adversarial/reconciliation nodes use a larger one (DeepSeek-V3, all 4 slots). The provider client is swappable via one environment variable.

Challenges we ran into

  • A real bug, found and fixed mid-project, not glossed over. Early on, Node 5 would correctly name the decisive fraud tell in its own stated reasoning and then still hedge into "uncertain" anyway — a rule-application bug, not a detection failure. Found by reading the model's own reasoning trace rather than just the wrong verdict, and fixed with an explicit meta-rule: if your own reasons already name a decisive tell, a calm tone doesn't override it.
  • An honest negative result we chose to publish anyway. We validated against a public, out-of-distribution dataset (the UCI SMS Spam Collection) neither system was tuned against. There, the workflow does not beat the baseline — a specific, identified blind spot where premium-rate SMS marketing mimics the same legitimacy signals (shortcode, visible T&Cs, no request for secrets) our legitimacy-evidence node is tuned to reward. We disclosed this in the documentation instead of quietly dropping the check, because a project built on "measured, not claimed" doesn't get to publish only the number that flatters it.

Accomplishments that we're proud of

  • False positives on genuine messages: 11.1% → 0.0%, with zero cost to scam recall (0% missed scams, both before and after).
  • Verified to generalise to Bengali-script messages with no language-specific prompting — including the hardest case in the whole set, a polite "I accidentally sent you extra money, please send it back" scam with zero threatening language.
  • An ablation study, not just an endpoint comparison: removing the adversarial node alone makes missed scams jump from 0% to 13%, direct evidence that node earns its complexity rather than just adding latency.

What we learned

Decomposing a judgment task into nodes that each commit to one side of the argument — instead of one pass that free-associates toward its first impression — turns "the model might have considered X" into "the model demonstrably did consider X, here's its written case for it." That's the difference between a claim and a measurement, and it's the whole reason this approach beat the single prompt on the metric that actually determines whether a scam detector is safe to hand to a real person.

What's next

Extending Node 3's legitimacy heuristics to cover the premium-rate/ subscription-trap fraud category our out-of-distribution check surfaced as a blind spot, and wrapping the workflow in a web front end.

Built With

  • adversarial-reasoning
  • ai-safety
  • bengali
  • concurrent-programming
  • cybersecurity
  • deepseek
  • evaluation-harness
  • featherless-ai
  • fraud-detection
  • json
  • llm
  • multi-agent
  • multilingual-nlp
  • natural-language-processing
  • prompt-engineering
  • python
  • qwen
  • regex
  • rest-api
  • scam-detection
Share this project:

Updates

Submission history