Inspiration

Every single day, we open dozens of websites, and before reading a single line of content, we encounter cookie banners. For the average user like Jia, these banners present confusing choices where clicking "Accept All" often feels like the default path. For power users and developers like Arjun, navigating through multi-layered preference settings introduces unnecessary friction.

The idea behind No, Thanks stemmed from exploring whether an autonomous approach could help manage cookie and privacy preferences without requiring manual navigation through deceptive interface patterns.

What it does

No, Thanks is an experimental multimodal AI agent built to interact with web privacy banners. Instead of relying solely on traditional scripts, it utilizes visual web understanding to interpret webpage elements, navigate consent menus, and assist in setting tracking preferences. It operates in two modes:

  • Headed Mode: A transparent, step-by-step view displaying the agent's visual processing and action sequence for debugging and verification.
  • Headless Mode: A faster, background execution mode designed to handle banner interactions with minimal UI overhead.

How I built it

No, Thanks was developed using a straightforward automation and AI stack:

  • Multimodal LLMs: Utilized to process visual page snapshots and reason through basic layout structures.
  • Browser Automation Framework: Built using Playwright to manage browser instances and coordinate event triggers.
  • TypeScript / Node.js: Powered the core control flow and script execution logic.

The basic loop involves capturing a snapshot of the browser view, passing it to the vision model to identify relevant elements, and executing corresponding automation commands based on the model's output.

Challenges I ran into

  1. DOM Variability: Standard scripts using rigid CSS selectors often break when website layouts change or use complex structures. Exploring a visual approach helped mitigate some of these layout-dependency issues, though edge cases still require refinement.
  2. Complex Preference Menus: Sites that lack a direct opt-out button require multi-step navigation through preference centers. Handling these deep trees reliably across different website designs proved to be a significant orchestration challenge.
  3. Balancing Speed and Feedback: Ensuring the agent's actions could be observed clearly without introducing excessive execution delays led to the separation of the headed debug view and the headless runtime.

Accomplishments that I'm proud of

  • Successfully implementing a working prototype that combines Playwright automation with multimodal visual reasoning to interact with live web elements.
  • Designing a dual-mode interface (Headed and Headless) that makes the agent's execution path observable during testing.
  • Building a complete project from scratch that addresses a practical web annoyance within a hackathon timeframe.

What I learned

Working on No, Thanks provided practical insights into the limitations and capabilities of current computer-use models and vision-based automation. I gained a better understanding of the trade-offs between deterministic script execution and probabilistic AI reasoning when dealing with unstructured web environments.

What's next for No, Thanks

  • Extension Packaging: Exploring ways to package the core logic into a more accessible browser extension format.
  • Rule Customization: Refining user controls to allow more specific preference parameters rather than binary choices.
  • Reliability Improvements: Continuing to test against a wider variety of edge-case layouts to improve overall execution stability.

Built With

Share this project:

Updates

Submission history