Agent Triage
Tagline: From Slack decision to approved GitHub work—and a code-grounded plan.
Track: Professional Agents
Public repository: https://github.com/ArjunKaliyath/agent-triage
Inspiration
Engineering work rarely begins as a well-formed ticket. It usually emerges across Slack: a customer request, a correction from an engineer, a missing requirement, and finally a decision to act. Someone must reconstruct that conversation, determine whether related work already exists, create or update a GitHub issue, and inspect the code before implementation can begin.
We built Agent Triage to automate that repetitive handoff without giving an AI model unrestricted access to GitHub. Our guiding principle was simple: agents should provide judgment, while humans and deterministic software retain authority.
What it does
Agent Triage turns Slack conversations into approved, implementation-ready GitHub work.
A user invokes the Review work shortcut on a top-level Slack message. Agent Triage retrieves a bounded window of up to 20 top-level messages while excluding replies, bot output, and messages outside the configured channel.
A Strands orchestrator delegates the conversation to TriageAgent, which can:
- Separate multiple topics in the same conversation
- Identify new work
- Match follow-ups to existing GitHub issues
- Propose issue creation or updates
- Request clarification or recommend no action
- Cite the exact Slack messages supporting each proposal
Nothing is written automatically. The user reviews each structured proposal in Slack and explicitly approves the exact GitHub action.
A deterministic executor then verifies the requester, payload hash, repository boundary, duplicate candidates, and current issue state before performing the approved create or update.
After a successful write, GitAgent automatically inspects a bounded set of repository files at a pinned commit. It privately returns relevant paths, observed facts, suggested implementation steps, and unresolved questions in Slack App Home.
How we built it
Agent Triage is a Python 3.12 application built with the Strands Agents SDK and Amazon Bedrock. The current model is openai.gpt-oss-120b-1:0 in us-east-1.
The main components are:
- Slack Bolt and Socket Mode for the native message shortcut, modals, approvals, and private App Home results
- A Strands agents-as-tools orchestrator with specialized TriageAgent and GitAgent roles
- GitHub's REST API for bounded issue discovery, deterministic issue writes, and commit-pinned file reads
- SQLite for local workflow state
- Fernet encryption for persisted proposal and workflow payloads
- A read-only local MCP interface for bounded snapshots, issues, WorkLinks, and repository files
We deliberately separated framework-independent contracts, storage, model orchestration, and deterministic execution. The agents never receive a general-purpose GitHub write tool. Instead, approved actions pass through application code with explicit safety checks and idempotency protections.
The repository also includes synthetic Slack data and a small test repository fixture so the workflow can be demonstrated without exposing real company conversations or source code.
Challenges we ran into
One major challenge was model reliability. Smaller models were inexpensive and fast, but did not consistently distinguish new work from updates or follow the structured proposal contract. We evaluated several Bedrock models before selecting one that passed our five-case synthetic triage evaluation.
Another challenge was overlapping Slack context. A bounded window can contain several unrelated topics, including work that has already been processed. Focusing only on the selected message would miss relevant follow-ups, while repeatedly processing the entire window would create duplicate proposals. We solved this with success-only WorkLinks that associate Slack timestamps with their resulting GitHub issues.
Safe execution was also more complex than simply giving the model a GitHub tool. We had to handle duplicate candidates, stale issue state, repeated approvals, uncertain API results, fabricated evidence, and proposals created before repository state changed.
Finally, GitAgent initially produced plans that were too verbose or overly speculative when repository evidence was limited. We introduced strict file, byte, path, output, and formatting bounds so its advice remains concise and grounded.
Accomplishments that we're proud of
We completed and validated the entire local workflow:
- Slack shortcut to bounded context retrieval
- Multi-topic structured TriageAgent proposals
- Existing-issue recognition and duplicate prevention
- Independent human approval for each action
- Deterministic GitHub issue creation and updates
- Automatic, read-only GitAgent investigation
- Private implementation plans delivered through Slack App Home
We are particularly proud that the model cannot directly write to GitHub. Exact approval hashing, repository restrictions, operation markers, conflict detection, replay protection, and persisted execution state keep consequential actions deterministic and auditable.
The project has 37 offline application tests plus synthetic repository fixture tests. Live testing verified Slack context retrieval, structured proposals, approved GitHub creates and updates, and automatic GitAgent delivery. Our measured five-case TriageAgent evaluation passed all cases without repository-boundary or fabricated-evidence violations.
What we learned
The most important lesson was that effective agent systems do not need unlimited autonomy. Separating reasoning from authority made Agent Triage safer and also easier to test.
We also learned that bounded context is not enough by itself. An agent needs durable knowledge of which evidence has already produced work. WorkLinks became essential for distinguishing a new topic from an already-processed conversation inside overlapping Slack windows.
Structured output still requires deterministic validation. Schemas catch malformed data, but application rules must also verify evidence timestamps, repository boundaries, issue targets, duplicate candidates, and approval identity.
Finally, model evaluation needs realistic cases. A successful smoke test only proves that a model can be invoked. Our create, update, clarify, completed-work, and prompt-injection cases revealed behavioral differences that simple connectivity testing would have missed.
What's next for Agent Triage
The next milestone is integrating the official GitHub MCP transport while preserving deterministic, approval-gated writes.
We also plan to deploy Agent Triage through AgentCore and harden hosted persistence, encryption-key management, retention, deletion, observability, and multi-workspace isolation.
Another important extension is bringing Agent Triage into calls and meetings. It could analyze approved transcripts, identify spoken decisions, and propose evidence-backed GitHub issue creation or updates through the same human-controlled workflow.
Longer term, we would like to support additional collaboration and work-management systems, configurable organization policies, richer evaluation datasets, and carefully bounded implementation or pull-request planning—while preserving the project's central principle: agents reason, humans approve, and deterministic software executes.
Built With
- amazon-web-services
- bedrock
- github
- mcp
- python
- slack
- strands
Log in or sign up for Devpost to join the conversation.