Inspiration
Phishing is still the #1 delivery mechanism for credential theft — and generative AI made it worse. Attackers can now spin up fluent, personalized lures and lookalike domains (paypa1, xn--pple-43d.com) at industrial scale, while the people most exposed — students, seniors, non-native speakers — are the least equipped to spot them.
We looked at existing anti-phishing tools and saw two failure modes: cloud-based scanners that upload everything you browse (a privacy problem of their own), and mute red/green verdicts that teach users nothing. We wanted a third option: protection that runs entirely on-device and explains itself.
What it does
PhishLens is a phishing-detection toolkit with two faces sharing one engine:
- A browser extension (Manifest V3, Chrome + Firefox): every page you visit gets a real-time risk score shown on the toolbar badge. Dangerous pages get an in-page warning banner listing exactly which signals fired — so the warning itself is a micro-lesson.
- A web analyzer (live link below): paste any URL, get an instant scored verdict with per-signal explanations. Same engine, zero install — judges can test it in 5 seconds.
Detection layers, all local:
- Typosquat & homoglyph radar — Levenshtein distance + a confusable-character folding table (Cyrillic/Greek lookalikes,
rn→m,1→l) catchespaypa1,g00gle, and punycode IDNs likexn--pple-43d.com - URL forensics —
user@hostdeception, raw-IP hosts, abused TLDs, URL shorteners, subdomain stuffing (paypal.com.secure.evil.top), credential-bait path keywords, domain entropy - Page-content signals — password forms posting cross-site or over plaintext HTTP, fields fishing for CVV/SSN/OTP/seed phrases, hidden iframes, and urgency/scare language ("verify within 24 hours")
- Explainable 0–100 score — every flag carries a plain-language reason; detection plus education
How we built it
The core is a single dependency-free engine.js (~400 lines of vanilla JS) that exposes analyzeUrl() and analyzePage(). The same file is loaded three ways — content script, popup, and the GitHub Pages web demo — so every surface gives identical verdicts.
The extension is standard MV3: a content script snapshots the DOM (form actions, input names, iframe visibility, link ratios, visible text), feeds it through the engine, reports the score to a service worker that paints the badge, and injects a Shadow-DOM warning banner when the score crosses the danger threshold. The popup adds a paste-any-link analyzer for links received in chat/email.
The web demo is a static GitHub Pages site with the analyzer plus two bait pages — a fake "PayPal suspended" page (cross-site HTTP credential form, CVV/SSN fields, urgency copy, hidden iframe → scores 100) and a clean control page — so anyone can watch the detector fire.
No frameworks, no build step, no backend: clone, load unpacked, done.
Challenges we ran into
- Calibrating weights. Typosquat detection vs. brand whitelisting pull against each other —
paypal.com.secure.evil.topmust flag whilepaypal.commust not. We resolved it with layered signals: exact-brand hosts get a trust bonus, brand-names-embedded-in-foreign-domains get a heavy penalty. - Hyphenated lookalikes.
paypa1-secure-loginas a whole never edit-distance-matchespaypal. Answer: fold confusables first, then compare — and let secondary signals (abused TLD + bait path keywords) compound the score. - Privacy constraint = no blocklists. No Google Safe Browsing, no threat feeds — every signal had to be derivable from the URL and page itself. That constraint became the product's identity.
- MV3 quirks. Service workers can't touch the DOM, content scripts can't read the badge — we settled on a clean message-passing split (content → background → popup).
Accomplishments that we're proud of
- A detector that catches page-content phishing signals, not just URLs — our hosted bait page scores 100 even on a legitimate github.io domain, because the credential form is the smoking gun.
- Fully functional in ~2 hours with zero dependencies, zero API keys, zero network calls.
- Every verdict is explainable — the tool teaches why, which is the actual defense against phishing.
What we learned
Phishing signals are composable — no single heuristic convicts, but weighted evidence does. We also learned how much of "looks legit" lives in page behavior (where the form posts) rather than appearance, and how Manifest V3's sandboxing shapes extension architecture.
What's next for PhishLens
- A compact on-device ML classifier layered on top of the heuristics
- Community-maintained brand/TLD registries + a one-click false-positive reporter
- Safari support (WebExtensions make the port nearly free)
- Warning-banner i18n for non-English-speaking users — the demographic phishers target most
Built With
- css
- cybersecurity
- html
- javascript
- privacy
Log in or sign up for Devpost to join the conversation.