Amaze on Work

Autonomous Incident-to-Fix Engineering Agent built with Strands Agents SDK & Amazon Bedrock AgentCore
Track: Professional Agents | AWS Agents for Humans Hackathon


Inspiration

Anyone who has carried an on-call pager knows the 2:00 AM drill. You get jolted awake by a PagerDuty siren or an urgent Slack ping in #prod-incidents. An unhandled exception or race condition is blowing up payment transactions, and while you're still rubbing sleep out of your eyes, every minute of outage burns real money:

$$\text{Downtime Cost} \geq \$500{,}000 \times \Delta t_{\text{hours}}$$

What follows is hours of forensic engineering toil:

  • Digging through thousands of noisy log lines to isolate where the null pointer or schema mismatch originated.
  • Mentally piecing together caller graphs across microservice repositories to guess what else might break if you touch a function.
  • Typing out a rushed, caffeine-fueled hotfix under intense scrutiny, praying the patch doesn't trigger secondary regressions in production.
  • Waiting for CI builds, updating Jira tickets by hand, and jumping between fragmented Slack threads.

Most of this initial triage isn't creative work—it's deterministic forensic investigation. We realized that if an agent had direct access to our codebase's structural graph, an isolated sandbox to test diffs, and clean tool integrations, it could diagnose the bug, verify a fix, and open a ready-to-merge Pull Request before an engineer even finished booting up their laptop.

That's why we built Amaze on Work.


What It Does

Amaze on Work is an autonomous DevOps and software engineering incident-resolution agent. It plugs directly into your existing ops stack (Slack, GitHub, Jira), picks up production alerts the second they fire, finds the root cause using code property graphs, validates candidate fixes in isolated Docker sandboxes, and handles the handoff safely.

Here is how the pipeline executes when an alert lands:

  1. Alert Ingestion via MCP: Webhooks capture raw error payloads from Slack channels, GitHub issues, or Jira Service Desk, assigning an internal tracking ID (INC-XXXX).
  2. Telemetry & Symptom Parsing: An Incident Parser extracts stack traces, error types, involved file paths, and environment context without hallucinating details.
  3. GraphRAG Code Traversal: Rather than feeding blind chunks of code into an LLM, we query a Neo4j Code Property Graph to trace callers, callees, and AST dependencies, determining the exact blast radius of suspect functions.
  4. Adversarial Fix Synthesis: A Critic agent and a Fix Writer agent work together to brainstorm minimal-scope patches. Splitting these roles prevents confirmation bias — the model doesn't rubber-stamp its own mistakes.
  5. Containerized Sandbox Verification: Fixes are hot-patched into disposable Docker environments (amaze-python:base and amaze-node:base). We run baseline tests and post-patch tests side-by-side to guarantee zero regressions. If tests fail, an automated Strands retry edge loops back to the Fix Writer.
  6. Risk-Aware Delivery (Human-in-the-Loop): A principled scoring engine evaluates change risk (0–100) to govern autonomous deployment.

The Risk Score & Safety Guardrails

We didn't want the agent guessing whether a fix was "safe enough." Every candidate patch gets scored 0–100 based on five weighted parameters:

  • Blast Radius (30 pts): How many other functions/services depend on the code being modified.
  • Test Coverage (20 pts): How well the affected function is covered by unit tests.
  • Change Size (15 pts): Minimal unified diffs score lower risk; larger rewrites score higher risk.
  • Environment (10 pts): Touching live production vs. staging/dev.
  • Code Complexity (10 pts): Cyclomatic complexity and branching logic introduced.

Based on that composite score:

Risk Score Tier Action Taken Governance
0 – 25 🟢 Low Risk Auto-opens GitHub Pull Request with full reproduction steps and test deltas Autonomous
25 – 50 🟡 Medium Risk Opens PR and notifies lead reviewer on Slack with interactive "Approve & Merge" buttons Human-in-the-Loop (HITL)
50 – 100 🔴 High Risk No code applied. Generates comprehensive forensic diagnosis and Slack alert Human Required

The goal isn't reckless autonomy — it's knowing when autonomy is appropriate.


How We Built It

  • Strands Agents SDK (strands-agents): The core orchestrator. We used the official GraphBuilder pattern to construct a 6-node DAG pipeline (parse → analyze → review → fix → validate → score) with a conditional cyclic retry edge (validate → fix) for self-healing regressions.
  • Amazon Bedrock AgentCore: Deployment runtime adapter conforming to the Bedrock AgentCore specification (agentcore.json, @app.entrypoint, and MCP server gateways).
  • Amazon Bedrock & High-Throughput LLMs: Amazon Nova Pro and Anthropic Claude on Bedrock as the cognitive reasoning engine, with instant failover support for high-throughput evaluation.
  • Neo4j Code Property Graph (with NetworkX fallback): We parse repository ASTs into a queryable graph of functions, calls, and imports so agents reason about architecture rather than raw text.
  • Isolated Docker Sandboxes: Disposable ephemeral containers with resource caps and disabled network access to evaluate test deltas before touching Git branches.
  • Model Context Protocol (MCP): Standards-compliant gateways connecting GitHub (issues & PRs), Slack (Block Kit interactive cards), and Jira (ticket transitions).

What We Learned

  • Dumping whole files into the prompt doesn't work. It caused hallucinated imports and missed edge cases. Grounding the model in the specific call graph around the bug fixed most of that.
  • Generating a plausible fix is easy — validating one is hard. The sandbox test loop, not the LLM's first guess, is what actually makes a fix trustworthy.
  • Engineers won't trust a black box. Showing the blast radius, test results, and risk score alongside every PR turned skepticism into actual adoption in our testing.
  • Split the writer from the reviewer. One model doing both jobs leads to confirmation bias — it approves its own sloppy work far too often.

Challenges We Overcame

  1. Autonomy vs. Safety: We didn't want an agent silently pushing unverified code to production. Building the 3-tier risk-scoring engine solved this by creating a reliable Human-in-the-Loop guardrail.
  2. Ephemeral Sandbox Management: Sandboxed tests can hang or consume excess CPU. We implemented hard 30-second timeouts, strict container memory quotas (512MB), and isolated volumes.
  3. Multi-Agent Coordination: Moving from sequential procedural loops to a Strands GraphBuilder DAG required careful state management across nodes using Strands invocation_state.

What's Next

  • Auto-generating permanent regression tests for each bug resolved so issues never recur.
  • Incremental Code Property Graph updates via GitHub webhooks on push events.
  • Canary rollouts integrated directly with Amazon CloudWatch alarms, triggering automated rollback if error rates spike.
  • Cross-repository graph traversal to detect breaking API changes across distributed microservices.

Built With

Share this project:

Updates

Submission history