Inspiration

Prompt injection caught my attention because its potential impact extends far beyond a single chatbot. As AI agents increasingly browse websites, retrieve information, and perform actions on behalf of users, untrusted web content can become a way to influence their behavior. I knew that a 24-hour hackathon would not be enough to solve such a broad security problem, and I did not want to pretend otherwise. Instead, I saw it as an opportunity to learn by building: to understand how indirect prompt injection works, explore how different AI agents encounter suspicious content, and turn that understanding into something concrete. InjectionLens grew out of that curiosity. Rather than attempting to build a universal defense, I focused on making hidden instructions visible, tracing where they came from, and explaining why their risks depend on what an AI agent can actually do. For me, the goal was not to solve the entire problem in one weekend, but to make a small, evidence-backed contribution while learning how to approach a much larger one.

What it does

InjectionLens analyzes web content from four different ingestion perspectives:

HTTP Source: What appears in the original HTML response.

Rendered DOM: What becomes available after the page is rendered.

Reader/Markdown: What survives text-oriented extraction.

Accessibility Tree: What is exposed through accessibility information.

The tool identifies AI-directed instructions, preserves their original locations and evidence, and explains their potential impact based on the selected agent's capabilities.

For example, a malicious payment instruction has different implications for an assistant that can only summarize content and an agent that can perform transactions. InjectionLens also supports: AI-summary-link auditing. URL-fragment inspection. Local User-Agent cloaking comparisons. Detection of hidden and obfuscated instructions. Source-backed local attack replicas. Reproducible evaluation and evidence reporting. Its goal is to help users understand what was detected, where it appeared, which ingestion paths exposed it, and why the selected agent's capabilities matter.

How we built it

InjectionLens uses a JavaScript-based architecture with a React frontend, a Node.js backend, and Playwright for browser-based analysis. The system extracts evidence from multiple representations of the same web page, preserves source and location information, and evaluates suspicious instructions through a capability-aware risk model. We also built a local replica suite inspired by documented prompt-injection techniques, including hidden instructions, Unicode obfuscation, AI-targeted cloaking, recommendation manipulation, and transaction-related attacks. A separate evaluation harness supports deterministic sampling, controlled payload transformations, multiple HTML insertion positions, pipeline-level measurements, and reproducible reporting. The final application includes a one-click demonstration designed to make these technical differences understandable without requiring judges to configure the project manually.

Challenges we ran into

One of the biggest challenges was that the same web page can expose different content through different ingestion pipelines. We had to preserve where each suspicious instruction appeared and which pipelines actually observed it, rather than merging everything into one anonymous text string. Another challenge was capability-aware risk assessment. A suspicious instruction should not automatically receive the same risk level regardless of what the receiving agent can do. Unicode obfuscation introduced additional complexity. During evaluation, we discovered a normalization inconsistency that reduced detection coverage for certain obfuscated inputs. We documented this limitation rather than claiming complete protection. We also discovered an evaluation error: eight benign control observations had been included in the denominator of our attack-detection measurements. We corrected the aggregation using the preserved raw observations and added regression tests. These challenges demonstrated that building a security tool requires not only detection logic, but also careful evidence preservation, meaningful metrics, and honest validation.

Accomplishments that we're proud of

By the end of the hackathon development process, InjectionLens had: Passed 254 automated tests. Passed all 8 local replica scenarios. Completed smoke testing, UI verification, and a successful client build. Produced a reproducible evaluation matrix and pipeline-level visualizations. Delivered a working local demo with capability-aware risk explanations. In our provisional, pre-P1 attack-matrix evaluation, 808 of 3,176 valid attack observations reached medium risk or above (25.4%), and 628 reached high risk or above (19.8%). These figures describe detection within our constructed evaluation matrix, not attack-prevention success or representative real-world detection accuracy. Real-world benign-page false-positive measurements were not performed. The evaluation also retains documented dataset-licensing and normalization limitations.

What we learned

This project changed how I understand prompt injection. Before building InjectionLens, I primarily thought of it as malicious instructions hidden in content. During development, I learned that the ingestion path, the location of the instruction, and the capabilities of the receiving agent can all affect the resulting risk. I also gained a deeper appreciation for security evaluation. Writing a detector is only one part of the work. Building reproducible fixtures, tracing evidence, handling obfuscated inputs, defining meaningful metrics, and checking the evaluation itself are equally important. Perhaps the most valuable lesson was learning to distinguish between detecting a potential attack and actually preventing one. Passing local tests does not establish real-world effectiveness, and a percentage is only meaningful when its population and methodology are clear. A 24-hour project cannot answer every question, but it can reveal which questions matter. InjectionLens gave me a practical starting point for continuing to explore AI-agent security.

What's next for InjectionLens

Future work could include improving Unicode normalization, evaluating more agent capabilities, measuring false positives on authorized real-world benign pages, and investigating more advanced detection methods. The long-term goal is to make indirect prompt-injection analysis more transparent, reproducible, and useful for developers working with AI agents.

Built With

Share this project:

Updates

Submission history