The biggest software failures rarely come from the line someone was looking at. Knight Capital lost about $460M in 45 minutes when old code was woken up by a new deployment. CrowdStrike crashed 8.5 million Windows machines when a parser read one field past its input. A single regular expression took Cloudflare down for 27 minutes. In every case the dangerous code was somewhere else, and nobody could see it from the edit site. We wanted developers to never feel scared about pushing an update.
Magellan is a pre-commit check for Python. It compares your working tree to your last commit, finds exactly what changed, follows the change through the call graph to every caller it could break, and runs a checklist of 12 rules built from real software disasters. You get one verdict: ok, review or block.
Unlike an editor's error checking, which asks whether the code in front of you is wrong, Magellan asks what your change broke. It catches problems in files you never touched, ranks how far the damage spreads, and stays quiet about old issues your change didn't cause.
We built Magellan in Python by mapping a codebase from its syntax trees, giving every function, method, class and constant a stable name and a hash so that reformatting or moving code never counts as a change. We then diff that map against the last git commit to find what was added, removed, renamed, or had its signature or body changed, and walk the call graph backwards from each change, fading the score at every hop to rank the code it could break. A checklist of 12 rules, each taken from a real software failure, runs only on the code the change touched, and the result is a single verdict of ok, review or block. To test it, we rebuilt famous failures like CrowdStrike, Knight Capital and Cloudflare as short code histories and ran Magellan on them, and we wrapped everything in a command line tool, a local live server and a static site with an in-browser "Try it" section.
The hardest part was deciding what counts as a real change, so that reformatting and moved code didn't create noise, while still following the impact across files without flooding developers with "everything is connected to everything." We also had to turn real disasters into rules specific enough to be useful but quiet enough to trust, since a tool that cries wolf gets ignored. On top of that, we had to make the same engine work in the terminal and in the browser, and we spent time getting setup to work across Windows and WSL.
We're proud that all seven of the failures we replayed were caught by the real engine, and that Magellan can point at a file nobody edited and say "this will break," which is the whole idea in a single result. We also built a try-it-yourself page that needs no install and uploads nothing, and we kept the tool honest about its limits by stating clearly that it is static analysis and that a clean report is not proof.
We learned that the hard part of preventing outages isn't finding bugs, it's seeing the connection between an edit and the faraway code that depended on it. We also found that a few well-explained findings are far more useful than a long list of warnings, and that reading real post-mortems, like the SEC filing on Knight Capital and CrowdStrike's root cause analysis, gives much better rules than guessing at hypothetical bugs.
Next, we want to extend Magellan beyond Python to languages like TypeScript and Java, and add a GitHub Action that comments on pull requests with the verdict and blast radius. We also plan to add an agent mode so AI coding tools are blocked before they ship a breaking change, data-flow checks for unvalidated input, and more rules drawn from real incidents, along with a measured false-positive rate from replaying open-source history.
Log in or sign up for Devpost to join the conversation.