Inspiration
Every Python project I have worked on has the same graveyard: a stack of open Dependabot pull requests nobody has touched in months.
It is not laziness. It is that the PR does not actually help you decide. It tells you urllib3 went from 1.26 to 2.2, pastes in two hundred lines of release notes, and leaves you to work out whether any of it matters to you. Reading the changelog is not the hard part. Mapping it onto your own code is. So you either merge and hope, or you let it sit there while the security warnings pile up.
I wanted something that closed that gap. Not another bot that watches version numbers, but one that reads the code you actually wrote and tells you, line by line, whether this particular bump touches it.
What it does
You give driftwatch a repo and a version bump. It finds every place in your codebase that uses that library, pulls the library's real changelog, and returns a verdict per call site.
The output is deliberately not a binary safe or unsafe. For a real cryptography bump it will tell you that line 24 uses Blowfish, which moved to a different module, so look at it, and that line 30 uses AES in exactly the same shape of code and is completely unaffected. Same file, same library, same bump, different answers. That contrast is the whole product.
It can also scan a repo's requirements and show you what has drifted, so you can find the bumps worth auditing before anyone opens a PR at all.
How I built it
It is a Strands agent with six tools:
find_usagesdoes a local AST scan for every import and call siteget_changelogpulls real release notes from GitHub, falling back to the repo's own changelog file when the project does not publish Releasesget_file_contextwidens the view around one call site when a snippet is ambiguoussearch_github_issueslooks for real-world reports when the changelog says nothing usefullist_dependenciesreads requirements.txt and pyproject.toml and checks PyPI for driftpost_verdictwrites the audit back to the pull request
The most important decision was architectural: the repo never enters the model's context.
Finding where a library is used is a solved, deterministic problem, so Python's built in ast
module does it locally and exhaustively before the model is involved at all. Only the matched
call sites and the changelog text get sent to the LLM. The practical consequence is that cost
scales with how much you use one library, not with how big your repo is. I measured it over a
4,212 file corpus: 12 seconds, and a 500 file repo costs about the same as a 5 file one.
The second decision was to not hardcode the tool sequence. The model picks which tools to call and when it has enough evidence to stop. That is why different bumps produce genuinely different paths. A well documented CVE fix needs two tools. A library that publishes no changelog at all sends it hunting through GitHub issues. Watching that choice happen live in the UI is the part that makes it an agent rather than a script with a prompt bolted on.
What I learned
Changelogs lie, and only running the code tells you the truth.
urllib3's 2.0.0 changelog says getheader() will be removed in 2.1.0. I installed 2.2.1 to
check. It is still there. The removal never happened.
cryptography's changelog says Blowfish was removed from the cipher module in 46.0.0. I installed 46.0.0. The old import still works, because they left a compatibility alias behind that quietly hands you the class from its new home.
Both times my assumption was wrong and I only found out because I tested it. That turned into a hard rule in the agent's instructions: the word "removed" in a changelog does not mean your code crashes. Say the API moved, say it is discouraged, but never claim a runtime failure the evidence does not support. Being wrong in the alarming direction destroys trust in every other verdict you give.
A changelog can live in two places, and I was only looking in one.
pyca/cryptography publishes zero GitHub Releases. My tool came back with nothing and the agent honestly answered "not enough evidence" for every call site. It looked like a reasoning failure. It was not. There was a complete 131KB CHANGELOG.rst sitting in the repo root the whole time, just never published through the Releases feature, because using that feature is optional and plenty of serious projects skip it.
Challenges I ran into
The scanner missed exactly the code that mattered.
The first version of the AST scanner only matched the literal urllib3.foo() spelling. So it
caught urllib3.PoolManager() and requests.Session(), which are the two call sites that do not
matter for a bump, and completely missed resp.getheader() and session.get(), which are
exactly the ones that do. Real code goes through instance variables, not through the module name.
I had to teach it to follow the chain: http = urllib3.PoolManager() to http.request() to
resp.getheader(). Later I extended that across files, because almost every real codebase puts
its shared client in one module and imports it everywhere else.
Azure OpenAI and Strands did not fit together the documented way.
The docs suggest pointing the OpenAI provider at a custom base_url. I read the installed
source instead of trusting that, and found it always constructs a plain AsyncOpenAI client,
which does not send the api-version parameter Azure requires on every request. The fix was to
build a real AsyncAzureOpenAI client myself and inject it through the provider's client=
escape hatch. Fifteen minutes of reading source saved what would have been hours of confusing
failures.
Two bugs that only showed up when I pushed past the demo.
FastAPI resolved correctly but returned zero releases, because I was only fetching the first page of GitHub Releases and FastAPI has shipped more than a hundred versions since the range I asked about. And my first run over a 4,000 file corpus took 67 seconds, because both passes were parsing every single file including thousands that never mention the library. Filtering on raw source text before parsing is sound, since an import has to literally contain the module name, and it cut the time to 12 seconds with identical results.
Getting the verdicts to stop overclaiming.
Early runs would label a call site "affected" and then explain in the next sentence why it was fine. Others asserted code would crash when it demonstrably does not. Fixing that was not a code change, it was writing much stricter rules about what each verdict label means and when the evidence actually supports a claim of breakage.
What's next
Running the code against both library versions in isolated environments and diffing real behaviour, which would turn a well grounded judgement into actual evidence. Support for languages beyond Python. And a scope aware variable tracker, since the current one is a good heuristic rather than real type inference.
Built With
- ai-agents
- ast
- azure-openai
- css
- fastapi
- github-api
- gpt-4o-mini
- html
- javascript
- openai
- packaging
- playwright
- pypi-api
- python
- requests
- server-sent-events
- strands-agents
- uvicorn
Log in or sign up for Devpost to join the conversation.