Inspiration
AI assistants can now browse the web and act for us: read email, fill forms, send messages. But there is a hidden danger. A web page can carry text that a human cannot see, hidden with CSS tricks, tiny fonts, or light-gray colors. A human reads a normal recipe. The AI reads the recipe plus a secret instruction.
We saw demos where a page quietly told an AI to leak a user's data or send them to a phishing site. The user never saw the trap. This attack is called indirect prompt injection, and it is one of the top security risks for AI agents (OWASP LLM01). We wanted to build a defense that is simple to add and easy to understand.
What it does
Our project is a shield that sits between an AI agent and the web. It protects the agent in three layers:
- Layer 1, Visibility gap: We open the page in a real browser and compare what a human sees with what the AI would read. Any text hidden from humans is stripped out.
- Layer 2, Judge: A second model checks anything suspicious and decides if it is an attack.
- Layer 3, Action guard: Before the AI does something risky, like emailing a stranger or sending a card number, the shield blocks it or asks the user first.
We also built a live dashboard that shows every step the AI took and what the shield decided, with a clear red HIJACKED or green PROTECTED result. You can replay saved runs with no internet, so the demo always works.
How we built it
- Agent: A Gemini-based agent that browses pages and uses tools (read email, send email, browse).
- Sandbox: A set of fake "evil" web pages, each hiding text a different way, plus normal pages to test for false alarms.
- Shield: The three-layer pipeline above, with a benchmark that compares us against an existing scanner.
- Dashboard: A Streamlit app that reads the run logs and shows the timeline, the attack chain, and a side-by-side "what the human saw vs. what the AI received" view.
We split into four roles (agent, sandbox, shield, and frontend/pitch) and worked on separate branches, then merged everything together and ran the full test suite before the deadline.
What we learned
- Hidden text is easy to make and hard to catch. There are many ways to hide it, so a single trick is not enough. We needed to render the page like a real browser to see the truth.
- One layer is not enough. Some attacks slip past the action guard but are caught by the visibility scan, and the other way around. Three layers together are much stronger than any one alone.
- Newer models resist better, but not always. Stronger models refused many attacks, but a determined page could still trick weaker ones. Defense at the tool layer matters no matter which model is used.
Challenges we ran into
- Making the AI fall for the attack. Newer models often refused, so we had to tune the payloads and test across models to get a clear, honest demo.
- Catching every hiding technique. Each new trick needed a new check, and we had to avoid blocking normal pages by mistake.
- Merging four people's work under time pressure. We kept everyone on their own branch, merged locally, fixed small mismatches, and confirmed all tests passed before touching main.
Model disclosure
Our demo runs on gemini-3.5-flash. We disclose this openly: newer models resisted some of our attacks, so we chose a model that shows the attack and the defense clearly. The shield works at the tool layer, independent of the model.
What's next
- Add more hiding techniques and more real-world test pages.
- Package the shield as a simple wrapper any developer can add in a few lines.
- Test against public prompt-injection datasets.

Log in or sign up for Devpost to join the conversation.