Inspiration
We kept seeing "AI agent" demos that either faked browser control entirely or only worked inside a sandbox the author fully controlled. We wanted to build something that actually drives a real browser through real multi-step tasks — and to do it on infrastructure we don't own, using an open model we can inspect, not a black-box API.
What it does
You describe a job in plain language — "find the cheapest 4-star book and summarize the author" — and Grasshopper splits it into steps, opens a real Chromium browser, walks the site, verifies each step, and stops to ask before anything risky (payments, sharing, publishing). The first run reasons through an LLM served by Nebius Token Factory; every identical run after that replays a saved playbook with zero model calls and zero added cost.
How we built it
Python, FastAPI, and Playwright for the browser layer. A model router that sends everyday steps to Nebius Token Factory's Nemotron models (our default fast tier) and only escalates to a stronger model when a step fails. A hard, code-enforced budget ledger caps spend per run and per day before any call goes out — not just a warning in the docs. An MCP Streamable HTTP server exposes the same agent to any MCP-compatible client. A learning store records successful playbooks so a repeated job costs nothing the second time.
Challenges we ran into
Nemotron Ultra's reasoning trace (1000+ tokens before it writes any answer) ate our entire completion budget on a max_tokens=1000 call and came back empty — we had to raise the limit and add resilient JSON extraction with a retry. Getting a real recorded browser video instead of a staged one took several rebuilds; our first attempts drew a fake browser chrome over screenshots, which we scrapped once we realized it wouldn't survive scrutiny. Proving a real LLM call happens — and costs what we say it costs — meant building the budget gate before we ever touched a live key, not after.
Accomplishments that we're proud of
A real Nebius Token Factory run: 41 live Nemotron decisions across five real-site scenarios for a total of $0.0238. The same scenarios repeated afterward: zero LLM calls, replayed straight from a stored playbook — an ~86% cost drop we can show, not just claim. Every payment path enforces a daily and per-transaction limit before a human approval gate, and nothing in the repo or its git history contains a live secret.
What we learned
Nebius Token Factory's OpenAI-compatible endpoint made swapping models trivial — the harder engineering problem was building a system honest enough to tell us when a call was mocked versus real, and cheap enough that we could run real benchmarks dozens of times without worrying about cost. We also learned that model-size tradeoffs (Nano vs. Ultra) matter as much for completion-budget planning as for raw capability.
What's next for Grasshopper: Voice-Controlled Multi-Step Browser Agent
Deeper per-task cost comparisons across Nebius model sizes, a Tavily-backed research-and-browse scenario, and extending the same router pattern to more open models available through Token Factory.
Log in or sign up for Devpost to join the conversation.