Inspiration Every year, elderly and disabled Americans leave over $30 billion in benefits unclaimed — not because they don't qualify, but because the application form itself is the barrier. We wanted to build an agent that doesn't just explain the form, but actually completes it — safely, for people who can't supervise an AI themselves.
What it does Proxy takes a plain-language task ("Complete my housing benefits application") and completes it on a real web form autonomously — reading the page, deciding its next action, and acting, with no fixed script. A split-screen interface shows the conversation with the agent on one side and a live view of the actual page it's working on on the other. Before any high-consequence action (like submitting), it pauses for human approval. For sensitive fields — SSN, date of birth, account numbers — it pauses and asks the human to type the value directly; the AI never sees or stores it.
How we built it Proxy runs a perceive → reason → gate → execute → observe loop: Playwright drives a real headless Chromium browser, an LLM decides the next action via tool-calling, and every single proposed action passes through a security gate before it's allowed to execute. The gate is a three-layer defense: a heuristic + LLM-judge scanner detects prompt injection hidden in page content (invisible text, white-on-white styling, malicious alt attributes), a strict allowlist contains what the agent can do even if detection fails, and anything uncertain escalates to a human. Backend is FastAPI streaming live events (including screenshots) over WebSocket to a split-screen frontend.
Challenges we ran into Indirect prompt injection is invisible by design, so a live screenshot of a blocked attack looks like nothing happened — we built a dedicated "Content Flagged" view that surfaces the exact injected text instead. We also found and fixed a real concurrency bug where synchronous API calls inside our async agent loop could freeze the WebSocket connection, and a browser-resource leak where completed tasks left orphaned Chromium processes running.
Accomplishments that we're proud of A working agent with a security model most hackathon projects (and some production agents) don't have — sensitive fields structurally never enter the LLM's context at all, not just "detected and blocked after the fact."
What we learned The hardest part of agent security isn't detecting attacks — it's designing containment that holds even when detection fails, and building trust UX for users who can't verify an AI's actions themselves.
What's next for Proxy Direct prompt injection defense, OAuth-scoped account integrations, voice input/narration for accessibility, and expansion beyond benefits forms to healthcare and appointment systems.
Built With
- chromium
- fastapi
- groq
- html
- javascript
- playwright
- python
- websocket
Log in or sign up for Devpost to join the conversation.