The Problem

It's 2am. An alarm fires. Your on-call engineer wakes up, opens PagerDuty, then AWS, then Datadog, then Slack, then a Confluence runbook from 2022. Forty minutes later, they finally understand what broke. Another hour to write the post-mortem.

This happens to every engineering team. And it costs more than most people realize.

The average production outage costs $5,600 per minute for enterprise companies. The average MTTR — mean time to resolve — is 287 minutes. Do the math: that's over $1.6 million per incident for a mid-size company.

The tools that exist today were not built for how AWS teams actually work. PagerDuty was built in 2009 — before Lambda, before ECS, before Aurora DSQL. It pages you. That's it. You still have to figure out the rest yourself.

We built IncidentMesh because the incident response workflow is broken, and fixing it requires starting from scratch with AWS-native architecture.


What We Built

IncidentMesh is a multi-region incident response platform that automates the entire lifecycle: detection, diagnosis, coordination, resolution, and post-mortem — all from a single platform built on top of Aurora DSQL.

Here is what happens when a production outage occurs:

1. Automatic Detection A CloudWatch alarm fires across five services — EC2, RDS, Lambda, ECS, and API Gateway simultaneously. SNS delivers the notification to IncidentMesh in under one second. An incident is created automatically with severity, affected services, and affected regions already populated. No human involvement required.

2. AI Root Cause Analysis The moment the incident is created, IncidentMesh sends the full context — alarm data, timeline, affected resources — to Claude Sonnet 4.5 running on AWS Bedrock. Within 20 seconds, the engineer has a root cause, a confidence score, specific remediation steps, and the exact AWS resources involved. What used to take 40 minutes now takes 20 seconds.

3. Slack-First Coordination The team gets an alert in Slack with two buttons: Acknowledge and Resolve. Clicking Acknowledge updates the incident status in Aurora DSQL instantly — the dashboard reflects the change in real time via Server-Sent Events. Engineers can run the entire incident from Slack using slash commands: /im-list, /im-status, /im-analyze, /im-resolve. They never need to leave the tool they live in.

4. Auto Post-Mortem When the incident is resolved, one click generates a complete post-mortem — root cause, impact assessment, resolution steps, lessons learned, and action items with owners and priorities — written by Claude Sonnet 4.5 in under 30 seconds.

5. Similar Incident Search Every incident is embedded using Amazon Titan Embeddings V2 and stored in Aurora DSQL. When a new incident arrives, IncidentMesh finds the five most semantically similar past incidents and surfaces their resolutions. The system gets smarter with every incident.


Why Aurora DSQL

This is not a project that uses Aurora DSQL as an afterthought. It is the architectural foundation.

Traditional incident response tools face a fundamental problem: when an outage affects multiple AWS regions simultaneously, the incident coordination database becomes a bottleneck. If your database is in us-east-1 and your engineers are updating incidents from eu-west-1, you get write conflicts, stale reads, and coordination failures — exactly when you need reliability the most.

Aurora DSQL solves this with active-active multi-region writes. Every engineer, regardless of their region, reads from and writes to the same globally consistent database. Optimistic Concurrency Control prevents conflicts when two engineers update the same incident simultaneously — the version field increments on every write, and conflicting updates are automatically detected and retried.

We surface this directly in the product. The incident detail page shows an OCC version counter. The sidebar shows live read and write latency to each DSQL region, updated every 10 seconds. The regional status grid shows the state of each affected region in real time.

Aurora DSQL is not infrastructure we deployed. It is the product feature that makes multi-region incident coordination possible.


Technical Architecture

Database: Aurora DSQL — 19 tables, active-active across us-east-1 and eu-west-1, OCC on all mutable tables, ASYNC indexes throughout

AI: AWS Bedrock — Claude Sonnet 4.5 for root cause analysis and post-mortem generation, Amazon Titan Embeddings V2 for semantic similarity search

Detection: CloudWatch composite alarm spanning EC2, RDS, Lambda, ECS, and API Gateway → SNS topic → HTTPS webhook → incident creation in under 5 seconds

Security: KMS-encrypted secrets (Slack tokens, API keys), IAM role assumption with external ID (confused deputy protection), per-organization data isolation

Frontend: Next.js 16 on Vercel, React 19, Server-Sent Events for real-time updates, D3.js for blast radius visualization

Integrations: Slack (OAuth, interactive buttons, 4 slash commands), PagerDuty (webhook import + manual import), AWS Cost Explorer (financial impact tracking), AWS EC2/ECS/Lambda/RDS (runbook execution)

Testing: 113 tests passing across 17 test files

Deployment: Vercel production deployment, environment variables managed via Vercel CLI


AWS Services Used

  • Aurora DSQL — core database, multi-region active-active
  • AWS Bedrock — Claude Sonnet 4.5, Titan Embeddings V2
  • Amazon CloudWatch — composite alarms across 5 services
  • Amazon SNS — alarm notification delivery
  • AWS KMS — secret encryption
  • AWS IAM — role assumption for customer AWS account access
  • AWS Cost Explorer — incident financial impact calculation
  • AWS EC2 / ECS / Lambda / RDS — runbook execution targets

The Business Case

IncidentMesh targets mid-market SaaS companies with 50 to 200 engineers. This segment spends heavily on AWS (typically $10K-$100K/month) and feels the pain of slow incident response acutely — they are too large to tolerate manual processes but too small to build their own tooling.

Pricing:

  • Starter: Free (up to 5 engineers)
  • Team: $499/month (up to 50 engineers)
  • Enterprise: $2,499/month (unlimited)

The competitive math is simple. PagerDuty charges $21 per user per month. A team of 20 engineers pays $5,040 per month — $60,480 per year. IncidentMesh Team plan costs $499 per month — $5,988 per year. That is $54,492 in annual savings, with better AWS integration and AI capabilities that PagerDuty does not offer at any price tier.

Go-to-market:

  1. Product Hunt launch targeting AWS-heavy YC startups
  2. Content marketing: "How we cut MTTR from 40 minutes to 20 seconds"
  3. Direct outreach to companies with AWS spend above $10K/month
  4. PagerDuty migration tool (already built — one-click import)

Path to revenue: First paying customer within 30 days of launch. $10K MRR within 90 days targeting 20 Team plan customers.


What Makes This Different

Every incident management tool on the market treats Slack as a notification destination. IncidentMesh treats Slack as the primary interface — engineers acknowledge, resolve, analyze, and query incidents without ever opening a browser.

Every incident management tool on the market uses a single-region database. IncidentMesh uses Aurora DSQL — the only serverless distributed SQL database that supports active-active multi-region writes. When an outage hits multiple regions, the coordination tool does not become the bottleneck.

Every incident management tool generates alerts and lets humans figure out the rest. IncidentMesh runs AI analysis automatically the moment an incident is created — by the time the engineer opens the dashboard, the diagnosis is already waiting.


What We Learned

Building on Aurora DSQL taught us something we did not expect: the hardest part of distributed systems is not the writes — it is making the consistency guarantees visible and understandable to the people using the system. We spent significant time surfacing OCC version numbers, regional latency, and conflict detection in the UI because an invisible database feature provides no competitive advantage.

The AI integration taught us the opposite lesson: the best AI features are invisible. When Claude analyzes an incident automatically the moment it is created, engineers do not think "the AI did this" — they think "IncidentMesh just told me what broke." That framing is the product.


Built With

Share this project:

Updates

Submission history