Inspiration

I run CI/CD across a dozen plus personal repos, from Flutter apps to Node projects to Discord bots. Every red build means the same ritual: stop what I'm doing, open the logs, scroll past setup noise to find the actual error, decide whether it's a real bug, a dependency conflict, or just a flaky test, and only then figure out what to do. Most of that doesn't need me. It needs someone to read the log and make the obvious call. Sentinel exists so I only get pulled in for the decision that genuinely needs a human.

What it does

Sentinel watches GitHub Actions, either triggered automatically the moment a workflow fails or run on demand. When it sees a failure, it works through a fixed process: pull the failure logs, read the diff of the commit that triggered it, and check whether the same commit has both a passing and a failing run on record, which is the standard signal for flakiness rather than a real problem.

From there it takes exactly one of three actions. If the failure is flaky, it retriggers the workflow and stops, since neither a fix nor an escalation would be the right response to noise. If it's a dependency version conflict where the failing tool already named the exact correct fix, it opens a pull request applying that fix directly. If it genuinely needs judgment, like an API signature that no longer matches what the code expects, it escalates with the root cause, a stated confidence level, and the specific decision a human needs to make, never just a wall of raw logs.

Separately, it checks whether the same commit made the project's documentation stale, and opens a documentation fix if it can identify the exact outdated text with confidence.

It watches multiple repositories at once, including repositories on different GitHub accounts, and it isn't locked to one model provider. Gemini, Claude through Bedrock or the Anthropic API directly, OpenAI, and Groq are all supported through one configuration setting.

How I built it

Sentinel is built on the Strands Agents SDK, with a small set of purpose built tools rather than one monolithic function: fetching failed runs and logs, reading commit diffs, checking flakiness against a repo's own run history, opening pull requests, escalating with a structured summary, and checking documentation drift. The agent reasons over which tools to call and in what order, rather than following a hardcoded pipeline, which is what let it correctly handle two very different real failures without needing separate code paths for each.

For automatic operation, I wrote a GitHub Actions workflow that any watched repository can add. It listens for that repository's own workflow runs, and when one fails, checks out Sentinel's code fresh and runs it, with the exact failing run already identified through GitHub's own event data. This runs entirely inside GitHub's infrastructure, not on my machine.

What I learned

The most useful lesson was how easy it is for an agent to sound confident while being wrong, and how much that depends on what the surrounding tools actually return rather than the model itself. Early on, a tool would fail to open a pull request and the agent would still narrate a success, because nothing forced it to check the tool's own return value before reporting an outcome. The fix wasn't a smarter model, it was making failure impossible to miss: tools that return explicit error information instead of failing silently, and a system prompt that requires checking for that before claiming anything succeeded.

I also learned that a flakiness check is only as good as what it's scoped to. My first version compared a commit's pass and fail counts across every workflow in a repository at once, which meant an unrelated workflow succeeding got counted as evidence the actual broken build was fine. A repeatable, 100% reproducible dependency conflict was being misread as random flakiness because two different workflows' results were blended together. Scoping the check to one workflow at a time fixed it completely, and it was a good reminder that a metric is only meaningful within the boundary it was actually measured over.

Challenges I ran into

The most persistent challenge was Amazon Bedrock account verification, which blocked every model invocation for days despite correct IAM permissions, confirmed model access, and a fully funded account. That turned into a genuine strength rather than just a setback, since it forced Sentinel to be model agnostic from early on rather than as an afterthought, and Gemini ended up being the provider I proved everything against first.

The GitHub Actions side had its own real issues, each only visible once I actually triggered a run rather than assumed the code was correct. The default token GitHub provides to a workflow can't open pull requests unless that's explicitly enabled, which took a real failed run and a 403 error to discover. A stale branch left over from an earlier test silently blocked a later pull request from being recreated correctly, since the original code assumed branch creation failures were always harmless. Both are now handled properly, but only because I was willing to actually run the automation against real repositories and read what came back, rather than trust that the code looked right.

Accomplishments I'm proud of

The moment I'm most proud of is a specific one: a real dependency conflict in one of my own apps, a real GitHub Actions failure, an automatic trigger firing with zero manual step, and a pull request that I reviewed and merged because it was genuinely correct. Not a staged demo, an actual fix to an actual problem, found and resolved by the agent on its own.

What's next

The clearest next step is expanding the flakiness signal beyond a single commit's pass and fail history into a pattern across a repository's recent history more broadly, and extending documentation drift detection from identifier overlap into something closer to real static analysis. I'd also like to support a shared dashboard across an entire team's repositories, not just a local view of one person's configured set.

Built With

Share this project:

Updates

Submission history