-
-
An agent that reads every commit, notices when the docs stopped being true, and decides whether to fix it, ask about it, or step back.
-
Both commits change one numeric default. The right one sits beside a second wrong claim, so the agent asks instead of guessing.
-
The artifact behind the FIX route — pull request #4, opened by the agent and merged into main without changes.
-
Event-driven architecture on GCP. The agent runs asynch and routes each finding to one of three outcomes, based on how certain it is.
-
A cent per finding, and four fifths of it is the model deciding whether it is sure. The expensive failure was never the bill.
-
The raw source behind the routing: every verdict the agent has reached, filtered in Cloud Logging. All three routes across several runs.
Inspiration
I renamed a config parameter. Tests passed, CI was green, the pull request got approved in four minutes. Three weeks later someone new joined, followed the README exactly, and spent an afternoon debugging a setup that could not possibly work.
Nothing was broken. The documentation had simply stopped being true, and there is no test for that.
No existing tool catches this. Linters check code against itself. CI checks behaviour against expectations. Nothing checks prose against reality — because the answer is almost never a clean yes or no. Sometimes the fix is obvious. Sometimes you genuinely need to ask the person who wrote it. That ambiguity is why this has stayed a human chore, and why it is a good problem for an agent that is allowed to be uncertain.
What it does
Driftwood watches a repository. When a commit lands, it reads the diff, finds the documentation that talks about the symbols that changed, and decides how confident it is:
| Route | When | What it does |
|---|---|---|
| FIX | The docs are demonstrably wrong and the correction follows directly from the code — a renamed parameter, a changed default, a removed flag. | Opens a pull request with the corrected text. |
| ASK | Something no longer matches, but several corrections are plausible and the agent would have to guess intent. | Opens an issue asking the maintainer one specific question. |
| ESCALATE | The docs describe a feature that no longer exists. Deleting it is a product decision, not a text fix. | Notifies a human. Changes nothing. |
The routing is the point. An agent that always opens a pull request is a nuisance within a week, because the false positives train you to ignore it. An agent that knows the difference between "I can fix this" and "I should ask" is something you actually leave switched on.
How we built it
A GitHub webhook hits a Cloud Run receiver that does almost nothing: verify the HMAC signature, publish the event to Pub/Sub, return 200. In the deployed system that round trip is 59 milliseconds. GitHub expects an answer in seconds and the analysis takes closer to a minute, so the two are decoupled from the start.
From Pub/Sub, Eventarc routes the event to a Workflow, which calls the Cloud Run Admin API to start the agent job. That extra hop is not decoration — see below.
The agent is built with the Google Agent Development Kit on Gemini 3.5 Flash. It extracts the changed symbols from the diff, retrieves only the documentation sections that reference them, and produces a structured verdict: the drift found, a confidence class, and the proposed action. It then calls the GitHub API to execute whichever route the verdict selected.
Firestore holds per-repository state — a fingerprint of every drift already reported and the pull request or issue it produced — so a rerun does not duplicate work it has already done.
Stack: Gemini 3.5 Flash via Vertex AI, Google ADK, Cloud Run (service + job), Eventarc, Workflows, Pub/Sub, Firestore, Secret Manager, GitHub REST API.
Challenges we ran into
Eventarc cannot invoke a Cloud Run job. There is no --destination-run-job; only services expose an HTTP endpoint Eventarc can push to. Jobs are started through the Admin API, which Eventarc will not call for you. The fix is a Workflow sitting between the two, invoking jobs.run on the agent. It is a small piece of YAML and it is load-bearing: without it the pipeline does not run at all. This is the kind of thing you only find by deploying.
The agent triggered itself. Every FIX opens a branch and a commit on the target repository — which fires two more push webhooks at the very agent that caused them. The branch-creation event arrives with before: 0000000000000000000000000000000000000000, which repo.compare() cannot resolve, so the job crashed. Cloud Run retried it three times by default, so one FIX produced eight crashed containers. The first fix I reached for was catching the 404. That was the wrong layer: it stops the crash but still runs the whole pipeline twice per FIX for nothing. The real fix is four lines at the top of handle_event — skip any push to refs/heads/driftwood/* before the diff is ever fetched. An agent that acts on the world generates its own inputs, and nothing warns you about that.
The model wanted to be helpful. Asked plainly to classify, Gemini returned "confidently fixable" for nearly everything, because that is the answer that sounds most useful. Uncertainty only became reliable when I inverted the burden of proof in the prompt: ASK is the baseline, and FIX has to be justified against the code. The asymmetry is deliberate — a wrong ASK costs a maintainer thirty seconds, a wrong FIX costs trust in the whole tool.
Idempotency was harder than the analysis. "Have I already reported this?" is not a string comparison, because the same drift looks different after every intervening commit. The fingerprint therefore keys on the subject — the symbol plus the documentation location — not on the generated text, and that holds across reruns. It is not finished: two symbols mentioned in the same README section still fingerprint separately, so one correction can arrive as two pull requests. Keying on the finding within a location rather than on the symbol alone is the next change.
Accomplishments that we're proud of
Driftwood runs unattended and produces something a maintainer would actually merge. Pull request #4 on the testbed repository changed "a short code between 6 and 10 characters" to "between 6 and 12 characters", because the default had moved from 10 to 12 — written by the agent, merged without edits.
The routing distinguishes cases that look identical from the outside. Two commits in that repository each changed a single numeric default in the same function signature. One became a pull request. The other became an issue, because the paragraph it touched also claimed a wrong expiry time, and correcting one number while leaving the other wrong would have been worse than asking:
[FIX] shorten_url @ README.md#shorten_url(long_url: str) -> str
The default max_length parameter in shorten_url was changed from 10
to 12, making the documentation's claim of 'between 6 and 10
characters' incorrect. The number 10 should be mechanically updated.
[ASK] shorten_url @ README.md#Example
The documentation is outdated regarding both the code length and the
expiration time. A simple mechanical fix is not sufficient as we need
to clarify if both values should be corrected.
That is context sensitivity rather than pattern matching, and it is the behaviour the whole design is aimed at.
The architecture holds up under measurement: 59 ms to acknowledge a webhook, 48 seconds from push to classification, with the gap absorbed entirely by the queue rather than by GitHub.
The economics were the part that surprised me most
Symbol extraction and documentation search run locally, in plain Python. Gemini is only called once a changed symbol actually appears somewhere in the docs. The cost scales with findings, not with commits — the overwhelming majority of pushes touch nothing that is documented and cost exactly nothing.
When there is something to judge, this is the bill:
| Tokens | Cost | |
|---|---|---|
| Input | 2,614 | $0.004 |
| Output, of which the verdicts are 479 | 2,583 | $0.023 |
| One push, three classifications | 5,197 | $0.027 |
Under a cent per finding. Now set that against the thing it exists to prevent — the afternoon from the top of this page, one engineer following a README that could not possibly work:
| Cost | Per incident | |
|---|---|---|
| One Driftwood finding | $0.009 | — |
| Half a day of engineering time, fully loaded | ~$300 | one wasted onboarding |
| Reviewing a repo's docs by hand, quarterly | ~$600/yr | catches drift months late |
One prevented afternoon pays for roughly 33,000 findings. A repository would have to generate a hundred findings a week for a decade before the API bill approached the cost of the single incident that started this project. Driftwood does not need a good hit rate to be worth running — it needs to be right once.
And the cheap part is not the interesting part. The expensive failure mode was never the bill; it was an agent that opens enough wrong pull requests that people stop reading them. That failure costs nothing on the invoice and everything in practice: a muted agent has a hit rate of zero no matter how accurate it was, because nobody is looking. Every false FIX spends trust that cannot be bought back at $9 per million tokens.
Which is exactly where the money goes. Of the tokens billed as output, only 479 are the actual verdicts. The remaining ~2,100 are thinking tokens — four fifths of the bill is the model working out whether it is sure. The cent buys the hesitation, and the hesitation is the product.
What we learned
That the interesting part of an agent is not what it does — it is what it declines to do. The version that opened a pull request for everything was easier to build and would have been switched off in a week. Most of the work went into teaching it to stop, and that turned out to be the whole product.
And that an agent with write access is a participant in the system it observes, not an observer of it. I designed a pipeline that reacts to pushes, then gave it the ability to push. The feedback loop was obvious in hindsight and invisible in the architecture diagram until the logs showed eight crashed containers from a single successful fix.
What's next for Driftwood
Beyond READMEs: API references, tutorials, and comments that describe behaviour that has since changed. Beyond a single repository: a monorepo-wide view where drift in one package's docs is caused by a change in another.
And beyond documentation entirely — the confidence routing generalises to anything where an agent has to admit it is not sure. That pattern is the reusable part.
Built With
- cloud-run
- firestore
- gemini
- gemini-3.5-flash
- github-api
- google-adk
- google-cloud
- pub-sub
- python
- vertex-ai
- webhooks

Log in or sign up for Devpost to join the conversation.