Inspiration

When production breaks, the alarm tells you what is broken — "5xx is up" — but not why or what to do. That context lives somewhere else: a Slack thread where someone already reported the same symptom, the Jira ticket that changed authentication handling, the Confluence runbook for rolling it back. On-call engineers spend the first, most expensive minutes of an incident jumping between those tools.

We already build Gatepath, which connects Slack, Jira, Confluence, Google Drive and SharePoint behind a single permission-aware search exposed as an MCP server. We wanted to see what happens when an incident agent can ask Gatepath the same questions an experienced engineer would — and answer by voice, in a few sentences, with the evidence attached.

What it does

Incident Pathfinder is an incident investigation agent delivered as a simulated Alexa+ experience. You ask in plain language and keep the conversation going:

  1. "What happened in production?" — it checks the active incident and recent deployments, then searches company knowledge through Gatepath. Answer: the 5xx error rate jumped after the latest auth-service deployment, and Slack reports point at the new token validation middleware.
  2. "What changed?" — it reuses the incident from the session and links the deployment to the Jira ticket and the Slack reports.
  3. "What should I do?" — it retrieves the rollback runbook from Confluence and summarizes the steps.

Every answer shows Evidence from Gatepath (source, title, snippet, link) and the tools that were called. The agent is strictly read-only: no rollbacks, deploys or ticket updates. When asked to act, it says: "No production change will be executed without explicit confirmation."

How we built it

Simulated Alexa+ Web UI → API Gateway → AWS Lambda (Incident Pathfinder Agent) → Amazon Bedrock (Converse API, tool use) chooses the tools → get_active_incidents / get_recent_changes (demo data) → search_gatepath / get_runbook → MCP client → HTTP / MCP → Gatepath MCP Server → Slack · Jira · Confluence · Google Drive · SharePoint → DynamoDB (session state) · CloudWatch Logs

  • Amazon Bedrock (Amazon Nova 2 Lite via the Converse API) runs the agent loop and decides which of four read-only tools each question needs.
  • Gatepath integration is pure MCP. Incident Pathfinder is an MCP client: each turn it runs initialize → tools/list → tools/call against the existing Gatepath MCP server over Streamable HTTP, using Gatepath's own search and getDocument tools. We did not reimplement search, connectors or permissions, and we did not modify Gatepath.
  • AWS Lambda + API Gateway serve both the UI and the chat API. DynamoDB keeps the session (current incident, recent findings), so follow-up questions don't start from zero. CloudWatch Logs records every Bedrock round and MCP call as structured JSON. The Gatepath token is stored in SSM Parameter Store.
  • Everything is deployed with Terraform. The backend is plain Python with no third-party dependencies.

Challenges we ran into

  • Search queries that actually match. Gatepath requires every query term to appear in a document, so pasting the user's question returns nothing. We taught the agent to search with 2–3 distinctive keywords (service and component names) and added a fallback that broadens a zero-hit query.
  • Non-deterministic evidence. Because the model picks the search keywords each time, the same question can show two or three sources: an extra keyword (for example "middleware") can exclude the Confluence runbook under Gatepath's all-terms matching. The runbook is still retrieved when you ask "What should I do?". We kept the model in control of the search rather than hard-coding queries.
  • Getting a small model to follow the investigation. Nova 2 Lite sometimes skipped tools or repeated its previous answer verbatim. We added a lightweight plan per question (which tools are required, what the answer should focus on) and forced a tool call on the first round.
  • Not overclaiming. Correlation is not causation. The prompt and post-processing keep the language at "most likely related" and make the read-only policy explicit whenever actions come up.
  • Making evidence visible. Judges can't open our private Slack or Jira, so the UI shows a cleaned snippet (Slack markup stripped) under each source, not just a link.
  • Model access. Anthropic models on Bedrock were not yet enabled in our sandbox account, so we built on Amazon Nova; the model is a single configuration variable.

Accomplishments that we're proud of

  • A real, end-to-end chain on every question: Bedrock → MCP over HTTP → Gatepath → Slack, Jira and Confluence.
  • A three-turn conversation that holds context and returns evidence-backed answers in about 5–8 seconds.
  • A read-only design enforced by the tool set and the IAM role, not just by the prompt.

What we learned

  • MCP makes reuse cheap: an existing MCP server became an agent's knowledge layer with a small HTTP client and no changes on the server side.
  • In incident response, operational signals and organizational knowledge are only useful together. The alarm says what; Slack, Jira and the runbook say why and what next.
  • Small models need structure — an explicit plan and focus per question — more than longer prompts.

What's next for Incident Pathfinder

  • Real CloudWatch alarms and GitHub deployment history instead of demo incident data
  • Native Alexa+ integration (Agent Skill or a self-hosted MCP server) and voice interaction
  • Per-user authorization: each user's own Gatepath token, so answers include only what they may see
  • Human-approved actions such as rollbacks and Jira/Slack updates

Hackathon scope: incident and deployment data are demo data; Gatepath search is real.

Built With

Share this project:

Updates

Submission history