Inspiration

The idea started from a bad week, not a brainstorm. One of us maintains a side-project repo that quietly picked up real users, and with them, real bug reports — and one report about data loss sat unlabeled for eleven days under forty other issues because nobody read past the first three that week. It wasn't a hard bug. It wasn't even a surprising one. It was just buried, because triage is the chore everyone means to get to and nobody schedules time for, and the one week we didn't get to it was the week it mattered.

That's the moment this project actually comes from: not "issue triage is tedious," but "the cost of tedious triage is invisible until it isn't." The fix couldn't be another dashboard to check — anything that requires remembering to open it has the same failure mode as the problem it's solving. It had to be something that reads every issue the moment it lands, applies the same judgment every single time, and only taps you on the shoulder when a human actually has to decide something. That's the whole shape of IssueOps: an operator, not an app.

What it does

Point IssueOps at a repo and the triage desk runs itself:

  • Reads and judges every incoming issue — severity (P1/P2/P3) and type — the way a senior engineer would, not a keyword filter.
  • Proposes a fix and an owner, searching the actual codebase and reading ownership data to point at the right file and the right person.
  • Labels, assigns, and comments — real, live actions on a real tracker, not a preview pane.
  • Remembers everything. Every decision is written to long-term memory and recalled on the next similar issue, so IssueOps says "we've seen this before" instead of re-litigating a bug for the third time.
  • Drafts release notes on request, pulled straight from what actually shipped.
  • Runs itself. Start it once and it watches the queue continuously, triaging whatever shows up — nobody picks the issue, nobody types a per-ticket command. It surfaces only when a decision genuinely needs a human.
  • Isn't locked to one tracker. Every agent talks to a tracker through one shared interface instead of hardcoded API calls, so the same orchestrator, memory, and fix-suggestion logic already run against GitHub, GitLab, and Jira — one config value picks which. A people-directory resolves the same person to their GitHub handle, GitLab username, or Jira account automatically. Bitbucket and other trackers slot into the same interface next.

By the numbers

The case for automating this is arithmetic, not opinion:

manual_cost_per_week  = issues_per_week × minutes_per_triage
                       = 50 × 4 min
                       = 200 min  (≈ 3.3 hours, every week, forever)

issueops_cost_per_week = time_to_run `cortex watch <repo>` once
                       = ~5 seconds

That gap doesn't shrink as the team grows — it grows with it. And from this build's own live run: 6 issues triaged, 1 recurring pattern caught by memory instead of re-diagnosed, 1 draft release compiled, 0 issues silently dropped, 1 command typed by a human for the whole autonomous run.

The severity call itself is a function, not a guess: severity(issue) = f(blast_radius, workaround_exists, data_at_risk) → {P1, P2, P3} — applied identically on issue 1 and issue 1,000, which is exactly the property a rotating cast of human triagers can't guarantee.

Who it's for

Maintainers and small engineering teams drowning in an issue queue that grows faster than anyone has time to read it — open source maintainers carrying a project solo, small product teams with no dedicated triager, support queues that need the same rubric applied every single time.

Why it matters

Triage is the highest-frequency, lowest-glory job on any engineering team, and it's exactly where consistency compounds: the same severity call every time, nothing silently dropped, and an institutional memory that survives turnover instead of walking out the door with whoever used to do it by hand. This is the class of work the hackathon brief describes almost exactly — repetitive, judgment-heavy, and better run quietly in the background than opened as one more app to manage.

How we built it

  • Strands Agents SDK, built around its native "agents-as-tools" multi-agent pattern: a top-level orchestrator routes to a triage agent, which calls a fix-suggester agent as a tool of its own, alongside a separate release-notes agent — real delegation between agents, not one prompt doing five jobs.
  • Amazon Bedrock AgentCore, used across all three pillars: Runtime (the orchestrator deployed behind a real, invokable serverless endpoint, built and pushed via CodeBuild), Memory (a semantic long-term store queried before every triage call and written to after), and Observability (automatic tracing on every Runtime invocation, no instrumentation code written by hand).
  • A provider-agnostic tracker interface — IssueTracker — so no agent or tool ever hardcodes GitHub. GitHub, GitLab, and Jira each implement the same contract against their real APIs, selected by one environment variable, with a cross-provider people directory for assignment.
  • Zero mocks. Every demo action — every label, comment, assignment, and release note — is a real API call against a live repo.

Challenges we ran into

Getting comfortable with Amazon Bedrock AgentCore for the first time was its own learning curve — Runtime, Memory, and Observability are three genuinely different mental models bolted under one name, and knowing which one actually solves which problem took longer than writing the code that uses them.

The harder problems were about the idea, not the infrastructure. The biggest one: how much should this thing actually decide on its own? An agent that asks for confirmation before every label isn't autonomous, it's a form with extra steps — but an agent that silently reassigns ownership or closes issues on a hunch isn't trustworthy either. We kept coming back to the same test: would a careful senior engineer make this exact call without asking anyone? Severity, type, a fix suggestion — yes. Who owns it, if the signal's weak — flag it, don't guess. That line is the whole product, and it moved more than once before it felt right.

The second one was memory. It's easy to build an agent that logs every decision it makes — that's just a database. It's much harder to build one where recalling a past decision actually changes what it says next, in a way a human would find useful instead of noisy. Deciding what's even worth remembering — a severity call, sure, but not every passing detail of every issue — meant treating memory as a design surface, not a feature we bolted on because the platform offered it.

The third was the tracker abstraction. GitHub, GitLab, and Jira don't just have different APIs — they don't agree on what a "label" is, what counts as an owner, or whether "release" even means the same thing. Writing one interface honest enough to cover all three without quietly assuming GitHub's shape everywhere else took more redesign than code.

Accomplishments that we're proud of

  • A posted GitHub comment that reads, unprompted, "This is a recurring issue, as seen in Issue #1" — Amazon Bedrock AgentCore Memory changing what the agent says, live, not a scripted line.
  • A background operator that actually runs unattended: start it once, and a brand-new issue gets triaged before anyone touches the keyboard again.
  • A real Amazon Bedrock AgentCore Runtime deployment, with Observability traces flowing, reached by working through the exact permission and configuration issues a first production deployment actually hits.
  • One agent codebase, three trackers, zero duplicated logic — the provider-agnostic layer is real architecture, not a slide.

What we learned

The provider-agnostic tracker layer came together after GitHub was already working end to end, and the agents themselves barely had to change — almost all of the work was drawing the interface boundary correctly and pulling GitHub-specific logic behind it. That's the sign of a well-shaped abstraction: the second and third integrations are cheap once the first one is honest about what's actually tracker-specific.

What's next for IssueOps

  • Bring Bitbucket and other trackers onto the same IssueTracker interface.
  • Move the background watcher from a polling loop to a webhook or an EventBridge Scheduler rule invoking the AgentCore Runtime endpoint directly, so "runs in the background" needs no terminal window at all.
  • Auto-fix PRs for the issue patterns memory has seen often enough to be confident about.

Built with

strands-agents · amazon-bedrock-agentcore · amazon-bedrock (Nova Lite) · python · pygithub · python-gitlab · docker / AWS CodeBuild · amazon-ecr · amazon-cloudwatch


Built With

  • amazon-ecr
  • bedrock
  • python
  • strands
Share this project:

Updates

Submission history