Inspiration

Every time I've needed data off a site with no public API, the answer has been the same: launch a headless browser, wait for the page to paint, dig through the DOM, and pray nobody renames a CSS class.

But that page isn't rendering data out of thin air. It's calling its own backend over clean JSON. The API is right there — it's just undocumented, and nobody hands you the client. So: use the site once, watch what it actually does, write the client from that, and throw the browser away.

What it does

You name a site and what you want from it. Beeline uses it once and hands you a typed client that calls its private API directly.

It works by using the flow three times with different inputs and diffing the network traffic. Anything that tracks your input is a parameter. Anything identical across all three is static. Anything that changes on its own is a session token — traced back to whatever issued it, whether that's a set-cookie, an earlier JSON field, or a value buried in a <script> tag.

Three runs is the minimum that works. With two, a rotating session cookie and a search query you happened to change look identical.

Then it goes looking for arguments the site never used. A model proposes candidates; each is tested against the live API and kept only if it provably narrows the result:

~ zai-org/GLM-5.3-Flash   guessed 12 parameters
                          page per_page limit offset sort order query
                          search filter min_schema_version author published_after
~ checked                 filter is real — narrowed 638 rows to 1

filter is in no documentation, and zed's own search box filters client-side so the site never sends it. Beeline found it by asking, then proving it.

How it works, end to end

Every number is from one real run against zed.dev/extensions.

   "get all extensions from zed.dev/extensions"
                    |
   1  USE IT 3x     |  provides = themes | languages | icon-themes
                    |  313 / 304 / 345 requests recorded
                    v
   2  DIFF THE RUNS |  moved with you  -> provides            PARAMETER
                    |  never changed   -> max_schema_version  STATIC
                    |  moved by itself -> csrf_token, session VOLATILE
                    v                     (traced to their source)
   3  PROBE         |  12 names proposed, 11 changed nothing
                    |  filter: 638 rows -> 1                  KEPT
                    v
   4  COMPILE       |  63 lines of TypeScript, zero dependencies
                    v
   5  VERIFY        |  88ms over HTTP  vs  3,212ms via browser = 36x
                    v
   6  REMEMBER      |  Cloudflare Worker + D1, re-checked every 15 min,
                       re-infers the schema when the site moves

The whole argument:

   WITHOUT                              WITH
   a headless browser, ~400 MB          fetch(), zero dependencies
   3,212 ms                             88 ms
   breaks when a CSS class is renamed   breaks when the API changes,
                                        and tells you within 15 minutes

How I built it

Browserbase + Playwright drive the capture over CDP, one cloud session per run — sharing one makes session cookies look static. A differential analyzer aligns the target request across runs, classifies every field, resolves volatiles to their origin, and infers a JSON schema. Baseten (GLM-5.3-Flash) proposes parameters and a verifier tests them live. OpenAI (gpt-5.4-mini) reads page structure when there's genuinely no endpoint. Cloudflare Workers + D1 store the specs, serve them at /call/:name, and run drift detection on a cron. The dashboard streams every stage over SSE — the live browser, the requests, the diff table, and each model guess with whether it survived.

Challenges I ran into

The two-run trap. My first version used two runs and confidently labelled session tokens as parameters. Three is the smallest number where "changed with my input" and "changed on its own" are separable.

Sites that lie about where their data is. zed.dev looked server-rendered, so Beeline fell back to scraping HTML and produced a client that worked but returned the wrong thing. The real cause: candidate requests were filtered by exact origin, and the endpoint lives on cloud.zed.dev. One fix turned a brittle scrape into a real parameterised API.

Making a model useful without letting it lie. Every LLM output here is checked against something real before it can reach the output.

Accomplishments that I'm proud of

The probing phase, specifically its failure mode. A suggestion must return fewer rows, at least one row, and only rows present in the original set. It's strict enough to reject parameters that genuinely exist — the right direction to be wrong in.

Registration that refuses to advertise a broken URL. The Worker bundles its own executor, so a spec newer than the deployment registers fine and then answers with nonsense. Registration now calls the endpoint once and checks; if the brain can't serve it, the entry is withdrawn.

Tracing a session token into a <script> tag. Steam's g_sessionID isn't a cookie or a header — it's assigned inside a JavaScript variable in the page HTML. Beeline finds it, works out which request produces it, and emits a client that bootstraps itself.

What I learned

Most of the modern web is a thin client over a fast JSON API nobody wrote down. Stop treating the page as the product, treat the traffic as the product, and a lot of scraping infrastructure turns out to be unnecessary.

And the thing I'd carry anywhere: let a model propose, and let something real decide. The model is allowed to be wrong, cheaply and often, because nothing it says reaches the output unverified. That's a better contract than prompting it into never being wrong.

What's next for Beeline

  • Stateful mutations — POSTs that change something, with a dry-run mode, since "run it three times to see what varies" is very different when the requests have consequences.
  • OpenAPI export, so a learned spec drops into Postman or a codegen pipeline.
  • Client-signed requests. Sites that compute a signature in JavaScript are out of reach today — the captured value works until it expires, and Beeline says so rather than pretending otherwise.

Built With

Share this project:

Updates

Submission history