Inspiration

Picking which tests to run after a change always felt like a guess. The usual trick is to match on filenames, so if you touch payment/processor.py you run test_processor.py. It's fast, and it's wrong in a specific, dangerous way. The test that actually catches your bug is often two or three calls away, in a file whose name has nothing to do with what you changed. So the filename picker skips it, the run goes green, and the regression ships anyway. I wanted the selection to be based on what the change can reach, not on what its name happens to look like.

What it does

On a merge request, Blast Test Picker posts the smallest set of tests you actually need to run, and the ones you can safely skip.

It starts from the files you changed and walks the Orbit call graph outward to every test that reaches them, directly or through a chain of calls. Anything a test can reach might be affected by your change. Everything else is provably untouched, so you can skip it.

On my demo, changing charge_card comes back as 4 must-run tests out of 10, so 60 percent of the suite is safe to skip. Three of those four tests are two hops away (test_checkout reaches the change through orders.checkout, and so on), and their names share nothing with payment.processor. A filename picker runs one test and misses the other three. That is the gap this closes.

How we built it

Same shape as a couple of my other tools: a small Python program, standard library only, that runs in GitLab CI on merge request events and is published as a Duo Agent flow.

It pulls the changed files from the MR, then walks the Orbit graph. The cross module part goes through ImportedSymbol, the same as my revert tool. I query "who CALLS the ImportedSymbol for this module," map each caller back to its file and module, and keep walking up to a few levels deep. Any test file I reach goes in the must-run list, every other test in the repo is safe to skip, and then I post the list along with the pytest command to run just those.

A few things I was careful about. It only certifies a skip when it actually finished the walk: if an Orbit query fails, it does not quietly shrink the suite, it reports a degraded run, withholds the skip set, tells you to run everything, and exits red. A test picker that skips tests because a query timed out is exactly how a bug sneaks through. The comment is rendered from graph data only, and names are sanitized before they go in the table. And the graph read retries with backoff and jitter.

Challenges we ran into

The interesting tests are never the obvious ones. Getting the transitive walk right, so test_checkout shows up when you change payment.processor even though they are three hops apart, was the whole point, and it was the part that needed the most care with the ImportedSymbol indirection.

Deciding what "safe to skip" is allowed to mean. The honest version is "no path from the change reaches this test within the depth I searched." I had to be careful not to oversell that, and to treat a failed query or a truncated result as "I can't certify this," not "skip it."

The same false green trap as my revert tool. An empty result from a broken query looks identical to "nothing to run." I made the degraded path explicit and wrote tests for it.

Accomplishments that we're proud of

It catches the depth-2 test a filename match throws away, and it does it on a real MR with the reasoning shown, so you can see this test reaches the change through that chain.

It never skips a test because something broke. Degraded means run everything, not skip everything.

The win grows with the codebase. The must-run set tracks how far your change reaches, not how big the suite is, so 4 tests stays around 4 whether the suite has 10 tests or 5,000, and the skip rate climbs toward 99 percent on a big suite.

No dependencies, and it posts live on real MRs.

What we learned

The reach of a change is a graph property, full stop. You can't get it from filenames or folders, and that is exactly why filename based selection misses things.

The same backwards walk I wrote for revert safety turned out to be the engine for test selection too, just read in the other direction.

The savings argument that actually lands isn't the percentage on a toy repo, it's that the must-run set is bounded by blast radius, so it barely grows as the suite grows.

What's next for Blast Test Picker

Wiring the selected set into a real pipeline stage that runs only those tests, with the full suite as a nightly safety net, so the time saving is something you can watch happen. Flagging when a changed file has no test reaching it at all, which is its own kind of risk. Treating a truncated query (too many callers to page through) as degraded instead of a clean skip. And more languages, since the walk itself doesn't care which one you are in.

Built With

  • ai-catalog
  • gitlab
  • gitlab-ci-cd
  • gitlab-duo-agent-platform
  • gitlab-knowledge-graph
  • glab
  • md
  • mypy
  • orbit
  • pytest
  • python
  • rest-api
  • ruff
Share this project:

Updates

Submission history