Inspiration
Being on-call at 3 AM is stressful. When an alert fires, engineers usually have to wake up, find their laptop, log in to a VPN, and dig through dashboards just to understand what is failing. We wanted to see if we could completely remove the laptop from the first step of incident response. The inspiration was to build a system that calls you, briefs you on the live state of your infrastructure, and lets you fix the problem using only your voice.
What it does
Incident Commander is a voice-driven Site Reliability Engineering assistant. When a CloudWatch alarm triggers, the system pulls live data from AWS APIs. It checks your ECS service health, recent logs, and load balancer targets. Then, it places an automated phone call to the on-call engineer using the CALL-E API.
The phone call reads out a summary of the incident based on the live AWS snapshot. The engineer can then reply with a spoken command to acknowledge the alert, escalate it to a secondary engineer, or take direct action like scaling up the service or rolling back to a previous deployment. If the engineer provides the required confirmation phrase, the app safely executes the change directly in AWS.
How we built it
We built the application using TypeScript. The backend listens for webhooks from CloudWatch SNS or can poll for alarms directly. When an incident occurs, we use the AWS SDK to gather real-time context from ECS, CloudWatch, and ELB.
We used the CALL-E Developer API to handle the voice interactions. We structured the application to use the CALL-E TypeScript SDK to create outbound calls and wait for the results. Once the call finishes and the engineer's spoken intent is parsed, our backend evaluates the response and maps it to specific, allowlisted AWS commands like UpdateService.
Challenges we ran into
We built the application using TypeScript. The backend listens for webhooks from CloudWatch SNS or can poll for alarms directly. When an incident occurs, we use the AWS SDK to gather real-time context from ECS, CloudWatch, and ELB.
We used the CALL-E Developer API to handle the voice interactions. We structured the application to use the CALL-E TypeScript SDK to create outbound calls and wait for the results. Once the call finishes and the engineer's spoken intent is parsed, our backend evaluates the response and maps it to specific, allowlisted AWS commands like UpdateService.
Accomplishments that we're proud of
We are really proud that we did not use any mock data or fake transcripts. Our project connects to real AWS accounts and runs real commands. Proving that you can safely scale an ECS cluster just by speaking into your phone while lying in bed was a massive success for us.
What we learned
We learned a lot about designing voice interfaces for technical users. When dealing with incidents, engineers want brevity and accuracy. We had to iterate on how the briefing was structured to make sure the most urgent information was presented first. We also gained deep experience with the CALL-E API and how to effectively link asynchronous phone calls back to long-running automated tasks.
What's next for Incident Commander
In the future, we want to expand the types of infrastructure we support beyond AWS ECS, integrating with Kubernetes and Terraform. We also plan to add a feature where the system can proactively suggest the most likely fix based on historical incident data, making the 3 AM phone call even shorter and more effective.
Built With
- amazon-cloudwatch
- amazon-ecs
- amazon-sns
- amazon-web-services
- api
- aws-elastic-load-balancing
- aws-sdk
- call-e-api
- ngrok
- node.js
- typescript
- webhooks
Log in or sign up for Devpost to join the conversation.