Inspiration

Finding a place to rent in Vancouver is stressful enough. Then you add scams. A lot of us have seen, or know someone who's fallen for, a listing that looked perfect: great price, nice photos, a landlord who's "working overseas" and just needs a deposit to hold it. By the time you figure out the apartment was never available, your money's gone.

The RCMP publishes a list of rental scam warning signs, but nobody reads that list while they're scrolling Marketplace at 1am. We wanted those warning signs to show up right on the listing, at the moment you're about to message someone or send money.

What it does

RedFlag is a Chrome extension that checks rental listings on Facebook Marketplace, Craigslist and Furnished Finder. Open a listing and a small traffic light shows up in the corner:

  • 🟢 No red flags
  • 🟡 Check first: one warning sign
  • 🔴 High risk: two or more warning signs

Click it and you see exactly which checks the listing failed, like asking for a deposit before a viewing, requesting your SIN, or sending money to someone out of the country. It doesn't tell you a listing is a scam. It tells you what to look at before you pay.

How we built it

The extension reads the listing from the page you already have open (title, description, price, photos) and sends it to our backend.

The backend is a FastAPI service in Python. It runs two things:

  • NLP detectors built with spaCy. We wrote five detectors, one for each RCMP warning sign: deposit demands, e-transfer requests, requests for personal information, pressure to act fast, and payments going outside Canada. They don't just look for keywords; they look at how the sentence is built. "Send me a $500 deposit to hold it" gets flagged, but "the deposit is due when you sign the lease" doesn't. For countries, we use named-entity recognition, but only flag a country when it's connected to the payment. "Send it to my brother in Mexico" counts; "I grew up in Mexico" doesn't.
  • A decision-tree model built with scikit-learn. Each tree focuses on one behaviour (one warning sign, or the price), and a listing goes red when at least two of them agree.

The data. We collected 20 real Vancouver rentals from Marketplace and wrote 5 scam listings based on the RCMP's list. Then we built a generator to scale that up to 500 listings (400 real, 100 scams), written to look alike in every way except how they behave.

Challenges we ran into

  • Our first model was cheating. It scored well, but when we looked at what it had learned, it was splitting on words like "the", "bedroom" and "fictional". The fake listings were just written differently from the real ones. We had to rebuild the dataset so real and fake listings matched in style, neighbourhoods and prices, and even add honest listings that mention deposits and e-transfers the normal way. Only then did it start learning actual scam behaviour.
  • All the trees agreed on one thing. Our bagged decision trees kept picking the same strongest feature (deposit requests), so they missed scams that asked for ID or a verification code instead. We restricted each tree to one behaviour, and that raised both accuracy and recall.
  • Price was misleading. Our generated prices started out higher than the real market, so the model thought normal listings were suspiciously cheap. Even after fixing that, cheap doesn't mean scam: rooms and studios are just cheaper. So price only counts as a red flag alongside another warning sign.
  • Facebook doesn't make scraping easy. Class names are scrambled and change often, so we read the listing from the page the user already has open, using text patterns instead of fragile selectors.

Accomplishments that we're proud of

  • The model never flagged a real listing in cross-validation (precision 1.00), and caught about 71% of scams.
  • On listings it had never seen, it caught 5 out of 5 of our hand-written scams and flagged 0 of 20 real ones.
  • The explanations make sense to a normal person. Instead of a mystery score, you see "Deposit demand" or "Payment sent abroad".

What we learned

  • A high accuracy score means nothing until you look at why the model is right. Checking what the trees split on caught problems that the numbers hid.
  • Getting the data right mattered more than the model itself.
  • For something that could label a real person's ad as "high risk", avoiding false alarms matters more than catching every single scam.

What's next for RedFlag

  • More real, labelled scam listings, so we're not relying as heavily on generated data
  • Price comparisons that understand the difference between a room, a studio and a full apartment
  • Reverse image search to catch photos stolen from other listings
  • Support for more rental sites and cities beyond Vancouver

Built With

Share this project:

Updates

Submission history