Inspiration

We wanted a hands-on proof-of-concept that actually demonstrates good AI-agent practices — not just calling an LLM, but knowing when not to call it, capping what it can spend, and never letting it act on something risky without a human saying yes first.

What it does

You describe a task in plain language — "find the cheapest 4-star book" or "check the top headlines" — and Grasshopper splits it into steps, drives a real Chromium browser through them, verifies each step, and stops to ask before anything risky (payments, sharing, publishing). The first run reasons through an LLM; every identical run after that replays a saved playbook with zero model calls.

How we built it

Python, FastAPI, and Playwright for the browser layer. A model router that sends easy steps to a fast, cheap model and only escalates to a stronger one on failure. A hard, code-enforced budget cap — not just a documented limit — stops any run before it can overspend. An MCP Streamable HTTP server exposes the agent to any compatible client, and a learning store records successful playbooks so repeat tasks cost nothing.

Challenges we ran into

Proving a real LLM call happens — and costs exactly what we claim — meant building the budget-tracking and cost-logging before we ever touched a live API key, not after the fact. Getting a genuine recorded browser video (not a staged one) took a few rebuilds; our first attempt drew a fake browser frame over screenshots, and we scrapped it once we realized it wouldn't survive scrutiny.

Accomplishments that we're proud of

Real numbers from real runs: dozens of live LLM calls across five real-site scenarios for a few cents total, and the same scenarios repeated afterward at zero cost, replayed from a stored playbook. Every payment path enforces a daily and per-transaction limit before a human approval gate, and the full test suite runs with zero API keys required.

What we learned

The best-practice habits that mattered most weren't about the model at all — they were about designing the system to refuse, log, and ask before doing anything it couldn't undo. Building the guardrails first made everything built on top of them easier to trust.

What's next for Grasshopper: Voice-Controlled Multi-Step Browser Agent

Extending the playbook-learning approach to more real-world sites, and adding a lightweight vision fallback for pages where structure alone isn't a reliable signal.

Built With

Share this project:

Updates

Submission history