Inspiration
We wanted a hands-on proof-of-concept that actually demonstrates good AI-agent practices — not just calling an LLM, but knowing when not to call it, capping what it can spend, and never letting it act on something risky without a human saying yes first.
What it does
You describe a task in plain language — "find the cheapest 4-star book" or "check the top headlines" — and Grasshopper splits it into steps, drives a real Chromium browser through them, verifies each step, and stops to ask before anything risky (payments, sharing, publishing). The first run reasons through an LLM; every identical run after that replays a saved playbook with zero model calls.
How we built it
Python, FastAPI, and Playwright for the browser layer. A model router that sends easy steps to a fast, cheap model and only escalates to a stronger one on failure. A hard, code-enforced budget cap — not just a documented limit — stops any run before it can overspend. An MCP Streamable HTTP server exposes the agent to any compatible client, and a learning store records successful playbooks so repeat tasks cost nothing.
Challenges we ran into
Proving a real LLM call happens — and costs exactly what we claim — meant building the budget-tracking and cost-logging before we ever touched a live API key, not after the fact. Getting a genuine recorded browser video (not a staged one) took a few rebuilds; our first attempt drew a fake browser frame over screenshots, and we scrapped it once we realized it wouldn't survive scrutiny.
Accomplishments that we're proud of
Real numbers from real runs: dozens of live LLM calls across five real-site scenarios for a few cents total, and the same scenarios repeated afterward at zero cost, replayed from a stored playbook. Every payment path enforces a daily and per-transaction limit before a human approval gate, and the full test suite runs with zero API keys required.
What we learned
The best-practice habits that mattered most weren't about the model at all — they were about designing the system to refuse, log, and ask before doing anything it couldn't undo. Building the guardrails first made everything built on top of them easier to trust.
What's next for Grasshopper: Voice-Controlled Multi-Step Browser Agent
Extending the playbook-learning approach to more real-world sites, and adding a lightweight vision fallback for pages where structure alone isn't a reliable signal.
Log in or sign up for Devpost to join the conversation.