Inspiration

In March 2024, a backdoor came within days of reaching nearly every Linux server on earth. The entry point wasn't a zero-day — it was a burned-out maintainer. Lasse Collin, the solo unpaid maintainer of xz-utils, had said publicly that he was struggling to keep up. A patient attacker spent two years building trust, then shipped the backdoor. It was caught by accident, by one engineer investigating a 0.5-second SSH delay.

44% of maintainers who quit cite burnout. Most of that burnout isn't code review — it's triage: reading a bug report, trying to tell if it's real, and deciding whether it's worth interrupting your day. I wanted to know if an agent could do the "is this real" part honestly, and refuse to guess when it can't.

What it does

Understudy watches a repo. For each inbound bug report, it reads the issue text, writes a Python script meant to reproduce the bug, and runs it in a locked-down container. If the script itself is broken, not the bug, the script, it reads the error back and retries, up to three times. It reports one of four verdicts and can act on that without asking: post a comment, request the one missing detail. But interrupting a person costs from a daily budget enforced in code, not requested politely.

How I built it

Strands Agents SDK for the agent loop (Agent, @tool, HookProvider / BeforeToolCallEvent), Pydantic-frozen schemas as the contract between stages, and Docker for isolated execution. One hook does two jobs: it gates the interruption budget against a durable SQLite ledger, and it blocks the agent from ever running git — SWE-bench images clone the full repo and reset to base_commit, so the fix commit is still reachable through history unless you block it explicitly.

Scoring is the part I care most about getting right: a script exiting non-zero doesn't mean the bug is real (an ImportError exits non-zero too). The only check this project trusts is the gold-patch differential — run the script at base_commit (must fail), run the identical script again with the dataset's real fix applied (must pass). Only fail-then-pass counts as reproduced, and the agent never sees the gold patch; only an offline scorer applies it, after the verdict is already committed.

Challenges I ran into

AWS Bedrock access on this account returns ValidationException: Operation not allowed, confirmed to persist across a full Free-to-Paid plan upgrade. Deploying through Bedrock AgentCore hit the same wall from a different angle — a working bedrock-agentcore:* IAM policy attached directly to the user still gets AccessDeniedException on every call. Two Bedrock-family services blocked the same way reads as an account-wide restriction, not something fixable from inside the account. The fallback was local inference (qwen2.5-coder:7b via Ollama) instead of a paid API, and I measured the honest cost of that instead of assuming it away.

The bigger challenge was almost shipping a wrong result. My first scoring pass checked whether the generated script exited non-zero on the buggy commit and called that "reproduced" — 14 of 16 instances, 88%. Running the same 16 instances through the real gold-patch differential instead: 3 of 16, 19%. Eleven of those "reproductions" were scripts broken for reasons that had nothing to do with the actual bug, and stayed broken after the real fix was applied. That gap is the actual headline finding of this project, not a footnote.

Accomplishments that I'm proud of

Publishing the wrong number next to the right one instead of hiding it. Also: a real, disclosed fork of tqdm (31k stars) carrying 3 real, unmodified upstream issues, with 3 real verdict comments posted by the agent — proof the full loop works outside a curated benchmark, on issues with no gold patch to check against. 74 tests pass, including real Docker integration and real Ollama calls where it matters, nothing mocked out.

What I learned

That the naive version of "did it work" is dangerously easy to ship by accident, and that disclosing a bad number is worth more to a reader's trust than a good one that turns out to be wrong. Also a very specific bug: one test instance ran on Python 3.5.6, and the model kept generating f-strings on every repair attempt because nothing told it the interpreter was that old — a small reminder that "the model is smart" doesn't substitute for telling it what it actually needs to know.

What's next for Understudy

Scale the eval past 16 instances, get a scheduled AgentCore + EventBridge deployment running instead of the current manually-triggered fork demo once AWS access clears up, and find out whether a bigger local model changes the reproduction rate without changing the interruption budget.

Built With

Share this project:

Updates

Submission history