Inspiration
Most browser agents rely entirely on the DOM to know what's on a page — but real pages break, layouts shift, and sometimes the only reliable signal is what a human would actually see. We wanted an agent that could fall back to vision when structure alone isn't enough, and be honest about exactly how far that vision layer currently goes.
What it does
You describe a task in plain language, and Grasshopper splits it into steps, drives a real Chromium browser through them, and verifies each step before moving on. Alongside DOM-based understanding, it captures screenshots at each step and can detect visual change between them — useful for catching broken layouts, unexpected popups, or UI states that DOM selectors alone would miss. It stops to ask for human approval before anything risky, and saves successful runs as reusable playbooks so a repeated task costs nothing the second time.
How we built it
Python, FastAPI, and Playwright for the browser and screenshot capture layer. A small vision service that compares screenshots step-to-step to flag visual change and clickable regions, called over HTTP so it can run separately from the main agent process. A model router for the reasoning side, with a hard budget cap enforced in code. An MCP Streamable HTTP server exposes the agent to any compatible client.
Challenges we ran into
Being honest about scope was itself a challenge: our vision layer currently does change-detection and basic element localization from screenshots, not a full OpenCV 5 pipeline — we chose not to overstate what it does. Getting a real recorded browser video (not a staged one) took several rebuilds after an early attempt drew a fake browser frame over screenshots, which we scrapped once we realized it wouldn't hold up to scrutiny.
Accomplishments that we're proud of
Real benchmark numbers from actual browser runs across five scenarios — success rate, step count, and cost — published in our README, not just claimed. A dashboard panel showing before/after screenshot diffs so a reviewer can see exactly what the vision layer catches. A repeated task needs zero LLM calls the second time, replayed from a saved playbook.
What we learned
Vision-based fallbacks are genuinely useful even at a modest level of sophistication — simple change detection already caught cases our DOM logic missed. We also learned it's better to ship a smaller, clearly-scoped vision feature we can defend than to overclaim a deeper computer-vision pipeline we didn't have time to build and verify.
What's next for Grasshopper: Voice-Controlled Multi-Step Browser Agent
A deeper OpenCV-based layer for clickable-element detection on pages with no usable DOM structure, and running that vision service on AWS as a separate, independently scalable component.
Log in or sign up for Devpost to join the conversation.