Inspiration

Every "AI agent" demo we'd seen either faked a browser or only worked inside a sandbox its creator controlled. We wanted the opposite: an assistant you talk to in plain language — through Alexa+, Telegram, or a CLI — that actually drives a real Chromium browser through a real multi-step job, the same way a person would, and that gets cheaper and faster every time it repeats a task.

What it does

You describe a job in a sentence or a voice note. Grasshopper splits it into steps, opens a real browser, walks the site, verifies each step, and stops to ask before it pays, shares, or publishes anything. The first time it does a job it reasons through an LLM; the second time, it replays a saved playbook with zero model calls. It exposes itself as an MCP Streamable HTTP server so Alexa+ (or any MCP client) can call it directly, and a browser-based Alexa+ simulator shows the same flow with voice in and voice out.

How we built it

Python, FastAPI, and Playwright for the browser layer; an MCP SDK 2.x server for the Alexa+ interface; a router that sends easy steps to a fast model (Nebius Token Factory / Nemotron) and only escalates to a stronger model on failure; a hard budget ledger that stops any run before it can overspend; and an allowlist/denylist policy so the agent only ever touches sites we explicitly cleared for the demo. A learning store records successful playbooks so repeat runs need no LLM at all.

Challenges we ran into

Getting a real recorded browser video instead of a staged one took several rebuilds — our first attempts drew a fake browser chrome over screenshots, which we scrapped once we realized it wouldn't hold up to scrutiny. Keeping cost at zero risk while still proving real LLM calls happen meant building a budget gate we could trust before ever touching a paid key. Balancing 500ms-class tool responses (for a voice assistant) against multi-minute browser tasks pushed us to an async start_task / get_task_status pattern instead of one blocking call.

Accomplishments that we're proud of

A first run on a real site takes dozens of LLM calls; the second identical run takes zero, replayed straight from a stored playbook. Every payment path enforces a daily and per-transaction limit before a human approval gate. The MCP server round-trips with an independent client, not just our own code. Nothing in the repo or its history contains a live secret — we ran a full history scrub before publishing.

What we learned

Trustworthy autonomy is mostly about what an agent refuses to do without asking — not how much it can do on its own. We also learned that "looks real" and "is real" are different bars, and that the second one is the only one worth building toward.

What's next for Grasshopper: Voice-Controlled Multi-Step Browser Agent

Wiring up real Alexa+ device access (beyond the browser simulator), deepening the Nebius routing story with live cost-per-task comparisons, and adding a computer-vision layer for sites that resist DOM-based automation.

Built With

Share this project:

Updates

Submission history