The problem
Five pull requests merged yesterday. This morning's nightly evaluation is 17 points worse. Every test is green.
Monitoring tells you that a metric moved. Nothing does the investigation, because investigating means running experiments and touching code. So an engineer repeats the same ritual: find the last good run, check out commits one by one, retrain, compare, read diffs, write a fix, prove it, write it up.
What it does
Culprit watches recorded nightly evaluations. When a metric crosses a regression threshold it opens one investigation for that last-good to first-bad window, and a Strands agent:
- Finds the drop — reads the recorded metric history and isolates the window.
- Proves the culprit — checks each candidate commit out in an isolated git worktree and actually runs the training and evaluation. Diffs suggest hypotheses; measurements decide.
- Explains it — reads the culprit's diff plus the current code and states the mechanism.
- Fixes and verifies — edits code on a fix branch, re-runs the experiment to show the metric recovers, adds a regression test that would have caught the bug, runs the project's suite.
- Asks you once — pauses on a Strands interrupt. The only question it asks is "open this pull request?"
- Delivers — opens the PR, notifies the team channel, and emits a structured incident report.
A run only counts as a success when the culprit is bracketed by experiments the agent actually ran, the metric recovers on the fix branch when re-measured independently, and the project's tests pass. A fabricated report cannot pass.
How we built it
One Strands agent with twelve instance-bound tools covering git inspection, experiments, editing, testing, approval and delivery. Three Strands HookProviders enforce an experiment budget, limit repeated tool calls, record traces, and interrupt before consequential actions. FileSessionManager preserves the session across a pause and a later resume from a different process — CLI, web dashboard, or AgentCore. Structured output produces the incident report. FastAPI plus server-sent events power the dashboard; a Typer CLI provides the same workflow. Amazon Bedrock is the default model provider, with Anthropic and OpenAI alternatives and a deterministic offline provider for credential-free testing. An Amazon Bedrock AgentCore Runtime entrypoint is included.
What the demo proves
The video records the local application using the explicitly labelled scripted integration-test provider. That provider knows the churn repair; the git worktrees, ML training, measurements, tests, approval and report are all real. These results are not evidence of real-model generalization.
In the recorded investigation, nightly F1 dropped from 0.8301 to 0.6624. Five experiments identified commit eb367c9. A separate evaluator independently re-measured the fix branch: fast-config F1 recovered from 0.7027 to 0.8111, matching its fast-config good baseline of 0.8111. It reran the project tests successfully and confirmed the new regression-test file. Fast-config and nightly values are reported separately on purpose.
The approval was explicit in the dashboard. The recorded PR and notification use local adapters; a live GitHub PR for the generated ML repository is not claimed.
Challenges we ran into
Making the result impossible to fake was the hard part. Anything that reads a diff and names a suspect will sound convincing. Culprit's design forces every claim through an experiment: commits are checked out in detached worktrees, training runs per (sha, config) with caching, retries and process-group timeouts, and the verdict comes from measurements rather than from the model's prose.
The second hard part was the pause. Approval has to survive the process exiting, so sessions are persisted and any of the three front ends can resume the same run.
Accomplishments we're proud of
- An unseen second scenario (
fraud-risk) with a different domain, files, model, data, mechanism and fix. A source-level test forbids the offline policy from referencing it, so the generalization check stays honest. - An independent evaluator (
culprit evaluate) that re-measures the fix branch and re-runs the tests instead of trusting the agent's own report. - Release verification: 67 tests passing, Ruff clean, on Python 3.10, 3.11 and 3.12.
What we learned
Monitoring and CI answer "did something break". Neither answers "which change broke it", because that answer requires running experiments. Once an agent can run experiments, the interesting engineering shifts from prompting to evidence discipline — and to deciding exactly where the human boundary sits.
Limitations, stated up front
- Real Bedrock / Anthropic / OpenAI investigation has not been run; no authorized provider credentials were available.
- The AgentCore entrypoint is included and verified locally with the offline provider; cloud deployment and cloud invocation are not verified.
- The scheduled workflow is an opt-in example for a configured ML repository, not an active job in the Culprit source repository.
- Sessions are filesystem-based. The dashboard has token/demo-only controls, not a full user-account system.
What's next
Real-model generalization runs on the unseen fraud-risk scenario through Bedrock, and a verified AgentCore Runtime deployment.
Try it in three commands
No AWS account, no API key — CULPRIT_MODEL_PROVIDER=scripted drives the real agent loop offline. Clone the repo, run pytest -q, generate the churn demo, and start culprit serve. The run pauses before the PR; approve in the dashboard to complete it. culprit evaluate --scenario churn reproduces the independent offline scoring.
Built With
- amazon-bedrock
- fastapi
- git
- numpy
- pandas
- pydantic
- python
- scikit-learn
- strands-agents-sdk
- typer
Log in or sign up for Devpost to join the conversation.