Inspiration

Every CI pipeline already runs Lighthouse. It has been telling you the same three things for months.

That is the whole problem. A score is not a fix - it is a chore, assigned to nobody, competing with the roadmap, and it loses every sprint. The work waits for somebody with a free afternoon to guess which file is at fault, guess whether their change helped, and guess whether it broke something else. Usually nobody has the afternoon, and the page stays slow until a customer says something.

The gap is not knowledge. Nothing turns a measurement into a diff a reviewer can accept in thirty seconds.

The name came from the metaphor that fit. A pit stop is a few seconds of work, judged entirely on the clock. Nobody argues about whether the stop was good - the lap time settles it.

What it does

Pitstop is a digital pit stop for your website. Bring your site in. A crew of five AI agents measures it, finds what is slowing it down, writes the fix, proves it worked, and hands you a pull request. You wave it through.

On every push to your default branch, without anyone asking:

  1. Measure. Build, serve and time every route - five loads each, cold cache, identical hardware. The median counts, never the best lap.
  2. Blame. Trace the failing metric to the file behind it: the actual LCP element, the script that blocked, the stylesheet holding the background.
  3. Patch. Write a real diff - re-encoded images, <picture> wrappers, image-set() with the original kept as a fallback, dimensions, fetchpriority, defer, font-display.
  4. Re-measure. Run the entire measurement again against the patched build. Same routes, same samples, same machine.
  5. Judge. Six criteria, all of which must hold, or the patch is thrown away with the reason kept on the record.

Those six: the target metric improved by at least 10%, no other Core Web Vital degraded by more than 5%, the rendered layout did not move, the build stayed green, no new console errors or failed requests, and the five samples agreed closely enough to compare.

Most patches lose. Every rejected one is kept with the criterion it failed and by how much, because a tool that only shows its successes is a tool you have no way to check. Our own landing page shows a withdrawal beside two wins, on purpose.

How we built it

Five agents, one job each - built on the Google ADK, not a hand-rolled loop. A SequentialAgent runs Surveyor to Attributor to RepairLoop to Scribe, and the RepairLoop is a LoopAgent holding the Surgeon and the Referee, capped at three attempts per problem.

Agent Job
Surveyor Puts the site on the clock. Stops the run if the samples are too noisy to trust.
Attributor Finds the file behind the number. Refuses to name one it is not sure about.
Surgeon Writes the diff. One change at a time, and reads why the last attempt was rejected.
Referee Decides, on rebuilt and re-measured numbers only.
Scribe Writes the pull request a human wants to read.

The Referee is the design decision the whole thing rests on. It runs with ADK's include_contents="none", so the Surgeon's reasoning is not in its context at all. It is not instructed to ignore the argument for a patch - it cannot read it. An agent that cannot see the argument cannot be talked round by it. That is the honest answer to "why would I let an AI write into my repository?"

Everything runs on Google Cloud. A Cloud Run Service takes the webhook, checks the approval gate, mints an installation-scoped token and starts a Cloud Run Job - one execution per run, destroyed afterwards, running as a dedicated service account with four roles and no access to the App's private key. The agents call Gemini 3.7 Flash on Vertex AI. Records go to Firestore, screenshots to a private Cloud Storage bucket signed per-render through the IAM Credentials API, so no service-account key file exists anywhere.

The measurement is deliberately not AI: Lighthouse, Playwright and axe on fixed hardware. The agents decide what to try and whether it worked. The stopwatch is a stopwatch.

Challenges we ran into

A wrong key returns empty rather than raising. This was the theme of the entire build. Lighthouse 13 renamed the audit we depended on (image-optimization-insight became image-delivery-insight), and .get() on a missing key returns {} - so the parser found nothing, reported nothing, and every test still passed. It looked exactly like a fast site.

The same shape bit four more times: a sidecar forwarding 2 of 9 audit ids; an escape that a shell heredoc turned into a control byte, so a regex silently matched nothing; a colour that landed on index 0 of a glyph ramp and was never drawn at all; and a self-trigger guard comparing against a product name that a rename had changed, leaving the bot one push away from triggering itself. None of them raise. All of them look fine in review.

"Isolated" is a claim about identity, not about files. We moved the GitHub App private key out of the container that runs user build commands and called it fixed. Then we found the job ran as the default compute account, which held roles/editor and secretAccessor on that key - so a hostile build could have fetched it from the API whether or not it was mounted. The question was never what the container can see; it is what the container's identity can ask for.

Gemini 3.x is served only from the Vertex global endpoint. A regional value does not fail - it silently serves 2.5 instead. A silent downgrade is far harder to notice than an outage.

Measurement noise is a real adversary. Our most dramatic demo page loads in 22 seconds and genuinely varies between runs. Teaching the system to say inconclusive and do nothing was harder, and more important, than teaching it to find things.

Accomplishments that we're proud of

It opened real pull requests, by itself, with the numbers in them. LCP 10.1s to 3.9s on /about - 61% faster, three files, four fix classes, verified by a full re-measurement. All of it public in the demo repository.

It refuses. The withdrawal on our landing page is the feature we are proudest of. A patch improved CLS and made LCP 21% worse, so nothing shipped and the reason was recorded. Most agent demos show you only the wins.

The Referee's blindness is structural, not a prompt. "Ignore the reasoning" is a request. include_contents="none" is a fact about what is in the context window.

The front end practises what it sells. Server-rendered, one stylesheet, no framework, no CDN, no webfonts - 71 KB per page. The ASCII animations are drawn procedurally into a canvas rather than shipped as the 627 KB inlined video the usual generators produce. A page-speed product with a slow front end is an argument against the product.

679 tests, and the ones worth reading are the ones that pin a mistake already made.

What we learned

Test doubles that are more permissive than reality certify code reality rejects. Our fake Firestore accepted document ids containing /. The real one does not, and every run died at the first write. The double is stricter than the service now.

Refusing to answer is a feature, and it has to be designed in. Every layer needed an "I do not know" that is louder than a wrong answer: the Attributor declining to name a file, the run reporting inconclusive on noisy samples, each fix class declining when a guard trips.

Separation beats instruction. Anywhere we wanted an agent not to be influenced by something, the reliable move was to remove it from the context rather than ask nicely.

Verify by running, not by reading. Every claim we made from reading a config was wrong at least once. The IAM policy looked correct while the job could still fetch the key. An animation looked broken while the code was fine - a hidden browser tab pauses requestAnimationFrame, so every probe was reading a frozen frame.

What's next for Pitstop

Finish the isolation. The job still writes every tenant's records directly, which is why approval is manual today. Reporting results back through the web service instead removes the last reason to gate installs, and open self-serve signup follows immediately.

More fix classes, each guarded so it declines rather than guesses: render-blocking CSS, third-party script deferral, content-visibility, and responsive srcset generation.

Server-rendered frameworks. Today it needs a static build. Next.js and SvelteKit in server mode are the most-asked-for gap.

Budgets per route, and a regression watch. A metric that has been green for a month and starts drifting should open a pull request before anybody notices.

Learning from rejections across repositories. Every withdrawal is already stored with the criterion it failed. That is a signal for which fix classes actually work on which shapes of site, and nothing uses it yet.

Built With

Share this project:

Updates

Submission history