Inspiration
As autonomous AI browsing agents (like Perplexity, ChatGPT Agent, Claude with Computer Use, and custom enterprise bots) become ubiquitous, they introduce an entirely new cyber attack surface: the information environment itself.
While human users view a visually rendered webpage, an AI agent parses the raw DOM, computed styles, hidden containers, accessibility tags, and metadata schemas. Traditional web security mechanisms like Content Security Policies (CSP), Web Application Firewalls (WAF), and XSS sanitizers only look for threats aimed at backend databases or human browsers. They are completely blind to natural language instructions designed to hijack or poison an AI model.
The real-world danger was recently exposed in security audits of browser assistants, where attackers embedded hidden CSS instructions on third-party pages that forced browsing agents to unknowingly exfiltrate private cookies and session credentials to external servers.
Inspired by Google DeepMind's landmark research on "AI Agent Traps" (Franklin et al., 2025), we created TrapScan — an automated, browser-native defense layer that inspects web content in real time and neutralizes adversarial traps before an agent can be hijacked.
What it does
TrapScan acts as a real-time security checkpoint operating in the browser, scanning pages across 6 core adversarial threat categories:
- Content Injection: Detects invisible CSS camouflage (
display:none,visibility:hidden, 0-pixel fonts, matching background/text colors), off-screen absolute positioning, hidden HTML comments, and manipulatedaria-labeltags. - Semantic Manipulation: Analyzes statistical term repetition and aggressive superlative density designed to skew an agent's reasoning and summary weights.
- Cognitive State Attacks: Flags fabricated JSON-LD schemas (
application/ld+json), conflicting metadata, and instruction-ladendata-*attributes that attempt to poison long-term memory or RAG vector storage. - Behavioural Control: Identifies direct jailbreak sequences (DAN mode, safety override commands) and unauthorized exfiltration payloads (
fetch(),XMLHttpRequest). - Systemic Traps: Detects coordinated Sybil-style repetition patterns engineered to manipulate multiple agents simultaneously.
- Human-in-the-Loop Attacks: Flags urgency triggers, countdown deception, and social engineering aimed at inducing approval fatigue in human supervisors.
Key Capabilities:
- Instant Risk Scoring (1–10): Categorizes pages into Safe (1–3), Suspicious (4–6), or Critical (7–10).
- Explainable Threat Cards: Pinpoints the exact element, attack vector, explanation, and potential damage in plain English.
- On-Page Warning Overlay: Pops up an urgent warning banner directly on pages with risk scores ≥ 5.
- Audit Reports: Generates downloadable, standalone HTML security audit reports.
- Dual-Mode Experience: Works both as an unpacked Chrome Manifest V3 extension and as a Zero-Install Web Demo.
How we built it
TrapScan is built using a modular, high-speed 3-tier architecture:
- Layer 1: Content Script (Client-Side Inspection): Injected on
document_idle. It captures a DOM snapshot and runs 6 heuristic and regex detection algorithms locally with near-zero latency without sending full webpage text over the network. - Layer 2: Service Worker & Gemma 4 AI Analysis: A Manifest V3 background service worker receives isolated suspicious fragments. When deeper semantic verification is required, it calls Gemma 4 (
gemma-4-26b-a4b-itvia Google AI Studio API) with strict JSON schema enforcement to validate intent and filter out false alarms. - Layer 3: React Popup & Dashboard: Built with React 18 and Vite, featuring interactive threat breakdowns, local IndexedDB scan history (
trapscan-db), and audit report generation. - Web Demo & Proxy: A Vercel serverless function (
api/scan.js) that fetches target URLs server-side to bypass CORS, allowing anyone to test live URLs in the web app instantly.
Challenges we ran into
- Accurate CSS Camouflage Detection: Catching text hidden through visual trickery required computing background colors against foreground RGB values and parsing complex CSS attributes without degrading browser performance.
- Eliminating False Positives in Developer Comments: Many benign sites contain standard build comments (e.g., webpack, vite, copyright). We engineered multi-stage filters to separate developer comments from adversarial prompt injections.
- Manifest V3 Async Message Passing: Chrome MV3 service workers terminate when idle. Ensuring reliable async message routing between content scripts, IndexedDB, and the Google AI Studio endpoint required disciplined event-driven state handling.
Accomplishments that we're proud of
- 100% Detection on Real-World Traps: Successfully detects all embedded attack vectors across synthetic adversarial test pages.
- Privacy-First Architecture: Sensitive page data never leaves the browser during local scans; only isolated suspicious candidate snippets are verified via the AI model.
- Resilient Fallback Engine: If an AI Studio API key is not supplied or network requests fail, TrapScan seamlessly defaults to local signature scoring so protection never stops.
What we learned
- The subtle nuances of indirect prompt injection and how semantic schemas (like JSON-LD) can be abused for stealth knowledge poisoning.
- How to write robust prompt schemas for LLMs that guarantee deterministic JSON responses for automated security pipelines.
- Modern Chrome Manifest V3 service worker lifecycle constraints and high-speed DOM parsing techniques.
What's next for TrapScan — AI Agent Trap Detector
- Framework Middleware: Releasing an SDK/middleware for LangChain, AutoGPT, and Playwright to scan URLs before browsing tools execute.
- Live DOM Mutation Observers: Continuous scanning for Single-Page Applications (SPAs) where content hydrates dynamically over time.
- Automated Shield Mode: Automatically stripping out or sanitizing malicious elements from the DOM before passing clean HTML to an agent.
- Enterprise Webhooks: Instant Slack and SIEM alert integrations for SOC teams monitoring enterprise agentic workflows.
Built With
- chrome
- css3
- cybersecurity
- gemma
- google-ai-studio
- html5
- indexeddb
- javascript
- manifest-v3
- prompt-injection
- react
- vite

Log in or sign up for Devpost to join the conversation.