-
-
Inbox — pending escalations, each specific to the actual bug found, confirmed by the gold-patch differential before it ever reaches you.
-
Architecture — one hook gates both the interruption budget and the git block.
-
The Watch — 16 issues triaged, 3 escalated. Quiet handling is the default; escalation cards are the rare exception.
-
Receipts — the actual generated script, real stdout, and the fail-then-pass gold-patch differential behind every "Reproduced" verdict.
-
Evidence -19% real reproduction rate vs. 88% a naive exit-code check would have claimed, plus the full failure taxonomy across 16 instances.
Inspiration
In March 2024, a backdoor came within days of reaching nearly every Linux server on earth. The entry point wasn't a zero-day — it was a burned-out maintainer. Lasse Collin, the solo unpaid maintainer of xz-utils, had said publicly that he was struggling to keep up. A patient attacker spent two years building trust, then shipped the backdoor. It was caught by accident, by one engineer investigating a 0.5-second SSH delay.
44% of maintainers who quit cite burnout. Most of that burnout isn't code review — it's triage: reading a bug report, trying to tell if it's real, and deciding whether it's worth interrupting your day. I wanted to know if an agent could do the "is this real" part honestly, and refuse to guess when it can't.
What it does
Understudy watches a repo. For each inbound bug report, it reads the issue text, writes a Python script meant to reproduce the bug, and runs it in a locked-down container. If the script itself is broken, not the bug, the script, it reads the error back and retries, up to three times. It reports one of four verdicts and can act on that without asking: post a comment, request the one missing detail. But interrupting a person costs from a daily budget enforced in code, not requested politely.
How I built it
Strands Agents SDK for the agent loop (Agent, @tool, HookProvider /
BeforeToolCallEvent), Pydantic-frozen schemas as the contract between stages, and Docker
for isolated execution. One hook does two jobs: it gates the interruption budget against a
durable SQLite ledger, and it blocks the agent from ever running git — SWE-bench images
clone the full repo and reset to base_commit, so the fix commit is still reachable
through history unless you block it explicitly.
Scoring is the part I care most about getting right: a script exiting non-zero doesn't
mean the bug is real (an ImportError exits non-zero too). The only check this project
trusts is the gold-patch differential — run the script at base_commit (must fail), run
the identical script again with the dataset's real fix applied (must pass). Only
fail-then-pass counts as reproduced, and the agent never sees the gold patch; only an
offline scorer applies it, after the verdict is already committed.
Challenges I ran into
AWS Bedrock access on this account returns ValidationException: Operation not allowed,
confirmed to persist across a full Free-to-Paid plan upgrade. Deploying through Bedrock
AgentCore hit the same wall from a different angle — a working bedrock-agentcore:* IAM
policy attached directly to the user still gets AccessDeniedException on every call. Two
Bedrock-family services blocked the same way reads as an account-wide restriction, not
something fixable from inside the account. The fallback was local inference
(qwen2.5-coder:7b via Ollama) instead of a paid API, and I measured the honest cost of
that instead of assuming it away.
The bigger challenge was almost shipping a wrong result. My first scoring pass checked whether the generated script exited non-zero on the buggy commit and called that "reproduced" — 14 of 16 instances, 88%. Running the same 16 instances through the real gold-patch differential instead: 3 of 16, 19%. Eleven of those "reproductions" were scripts broken for reasons that had nothing to do with the actual bug, and stayed broken after the real fix was applied. That gap is the actual headline finding of this project, not a footnote.
Accomplishments that I'm proud of
Publishing the wrong number next to the right one instead of hiding it. Also: a real, disclosed fork of tqdm (31k stars) carrying 3 real, unmodified upstream issues, with 3 real verdict comments posted by the agent — proof the full loop works outside a curated benchmark, on issues with no gold patch to check against. 74 tests pass, including real Docker integration and real Ollama calls where it matters, nothing mocked out.
What I learned
That the naive version of "did it work" is dangerously easy to ship by accident, and that disclosing a bad number is worth more to a reader's trust than a good one that turns out to be wrong. Also a very specific bug: one test instance ran on Python 3.5.6, and the model kept generating f-strings on every repair attempt because nothing told it the interpreter was that old — a small reminder that "the model is smart" doesn't substitute for telling it what it actually needs to know.
What's next for Understudy
Scale the eval past 16 instances, get a scheduled AgentCore + EventBridge deployment running instead of the current manually-triggered fork demo once AWS access clears up, and find out whether a bigger local model changes the reproduction rate without changing the interruption budget.
Log in or sign up for Devpost to join the conversation.