🛡️ AuditGuard (CloudBill Sheriff)

> **Autonomous FinOps & Cloud Waste Patrol Agent powered by AWS Strands Agents SDK**

---

### 💡 Inspiration
Cloud bill shock is a universal nightmare for modern engineering and finance teams:
1. **Silent Leakage**: Developers spin up high-performance GPU instances (e.g., `p3.2xlarge` or `g5.4xlarge`) for model fine-tuning or experiments on Friday afternoon and forget about them. Two weeks later, thousands of dollars have quietly evaporated.
2. **Alert Fatigue**: Existing cost management solutions send weekly 50-page PDF digests or passive email notifications that engineers habitually archive without reading.
3. **The Fear of Breaking Production**: Developers actively resist blind cron jobs and automated termination scripts because an aggressive cleanup script might accidentally kill an active production dependency.

Organizations leave idle resources running out of fear, burning an estimated **$40+ billion globally each year**.

We were inspired by the **AWS "Agents for Humans"** mission: **AI agents shouldn't just spam human inboxes with alerts, nor should they act as reckless black boxes.** Instead, an agent should continuously handle tedious background patrol tasks while strictly

honoring human agency. That is how AuditGuard (CloudBill Sheriff) was born.

---

### ⚡ What It Does
**AuditGuard** is an autonomous background AI agent built natively with the **AWS Strands Agents SDK**. It continuously patrols cloud infrastructure, discovers resource leakage, quantifies financial impact, and enforces a safety-guaranteed **Human-in-the-Loop

(HITL)** decision gate to recover thousands of dollars in wasted cloud spend with one click:

* **Autonomous Background Patrol**: Eliminates manual dashboard monitoring. The agent autonomously loops in the background using scheduled Strands reasoning workflows.
* **Quantified FinOps ROI**: Instantly translates raw technical metrics (low CPU utilization, 0 IOPS, unattached states) into concrete financial figures ($/month leakage and annualized budget recovery).
* **Safety-First Human-in-the-Loop Gate**: Strictly implements *"Agents propose, humans approve"*. Zero destructive operations happen without explicit cryptographic operator authorization.
* **100% Rollback Safety Net**: Automatically enforces pre-flight volume snapshots before terminating or stopping any resource, ensuring zero-risk rollbacks.

---

### 📐 FinOps Mathematical Model & Scoring

To balance waste elimination against operational risk, AuditGuard computes the **Monthly Financial Burn Rate ($B_{\text{monthly}}$)** and a multi-factor **Remediation Priority Index ($\Phi_{\text{risk}}$)** using formal FinOps telemetry:

#### 1. Monthly Financial Burn Rate
$$\text{Burn}_{\text{monthly}} = \sum_{i \in \mathcal{R}_{\text{idle}}} \Big( \mathcal{C}_{\text{compute}}(i) \cdot 730 + \mathcal{C}_{\text{storage}}(i) \cdot \mathcal{S}_i + \mathcal{C}_{\text{network}}(i) \Big)$$

Where:
* $\mathcal{C}_{\text{compute}}(i)$ is the hourly on-demand rate of instance $i$
* $\mathcal{S}_i$ is the provisioned volume capacity (GB)
* $730$ represents standard operational hours per month

#### 2. Remediation Priority & Blast Radius Index
$$\Phi_{\text{risk}}(i) = w_1 \cdot \left( \frac{\Delta t_{\text{idle}}}{168} \right) + w_2 \cdot \left( \frac{\text{Cost}(i)}{\text{Cost}_{\text{threshold}}} \right) + w_3 \cdot \big( 1 - \overline{\text{IOPS}}_i \big) - \mathcal{P}_{\text{tag}}$$

Where $w_1, w_2, w_3$ are weighted coefficients ($\sum w_k = 1.0$), and $\mathcal{P}_{\text{tag}}$ represents protected environment tags (`env:production` or `criticality:high`), ensuring production infrastructure is never penalized or mistakenly flagged.

---

### 🛠️ How We Built It

AuditGuard is engineered around three tightly integrated pillars:


┌─────────────────────────────────────────────┐
│Autonomous Patrol Engine (Strands Agents SDK)│
│                                             │
│                                             │
│  ┌───────────────────────────────────────┐  │
│  │                                       │  │
│  │     🕒 Scheduled Background Patrol    │  │
│  │                                       │  │
│  └───────────────────┬───────────────────┘  │
│                      │                      │
│                      ▼                      │
│  ┌───────────────────────────────────────┐  │
│  │                                       │  │
│  │     🛠️ Tool: scan_cloud_resources()   │  │
│  │                                       │  │
│  └───────────────────┬───────────────────┘  │
│                      │                      │
│                      ▼                      │
│  ┌───────────────────────────────────────┐  │
│  │                                       │  │
│  │     🛠️ Tool: analyze_finops_waste()   │  │
│  │                                       │  │
│  └───────────────────┬───────────────────┘  │
│                      │                      │
└──────────────────────┼──────────────────────┘
               Surfaces│Finding
 ┌─────────────────────┼─────────────────────┐
 │      Human-in-the-Loop Decision Gate      │
 │                     ▼                     │
 │ ┌───────────────────────────────────────┐ │
 │ │                                       │ │
 │ │       🔔 AuditGuard Finding Card      │ │
 │ │                                       │ │
 │ │     (Risk Score + Monthly Savings)    │ │
 │ │                                       │ │
 │ └───────────────────┬───────────────────┘ │
 │                     │                     │
 │                     ▼                     │
 │ ┌───────────────────────────────────────┐ │
 │ │                                       │ │
 │ │  👤 Human Operator: One-Click Approve │ │
 │ │                                       │ │
 │ └───────────────────┬───────────────────┘ │
 │                     │                     │
 └─────────────────────┼─────────────────────┘
                       │
 ┌─────────────────────┼─────────────────────┐
 │      Remediation & Safety Guarantee       │
 │                     ▼                     │
 │ ┌───────────────────────────────────────┐ │
 │ │                                       │ │
 │ │     🛠️ Tool: execute_remediation()    │ │
 │ │                                       │ │
 │ └───────────────────┬───────────────────┘ │
 │                     │                     │
 │                     ▼                     │
 │ ┌───────────────────────────────────────┐ │
 │ │                                       │ │
 │ │ 📸 Enforce Pre-flight Safety Snapshot │ │
 │ │                                       │ │
 │ └───────────────────┬───────────────────┘ │
 │                     │                     │
 │                     ▼                     │
 │ ┌───────────────────────────────────────┐ │
 │ │                                       │ │
 │ │        ☁️ Remediate Cloud Asset       │ │
 │ │                                       │ │
 │ │  (Stop GPU / Release EIP / Block S3)  │ │
 │ │                                       │ │
 │ └───────────────────┬───────────────────┘ │
 │                     │                     │
 │                     ▼                     │
 │ ┌───────────────────────────────────────┐ │
 │ │                                       │ │
 │ │      📝 Cryptographic Audit Trail     │ │
 │ │                                       │ │
 │ └───────────────────────────────────────┘ │
 │                                           │
 └───────────────────────────────────────────┘


1. **AWS Strands Agents SDK Core**:
   * `@tool scan_cloud_resources`: Collects multi-service inventory and CloudWatch telemetry across EC2, EBS, Elastic IP, and S3.
   * `@tool analyze_finops_waste`: Multi-factor reasoning engine that calculates idle duration, risk blast radius, and remediation roadmaps.
   * `@tool execute_remediation`: Safe execution tool protected by operator authorization tokens and automated snapshot hooks.
2. **Dual-Engine Architecture**:
   * **Mode A (Production AWS)**: Directly interfaces with Amazon Bedrock (**Claude 3.5 Sonnet**) and native AWS SDK (`boto3`).
   * **Mode B (Instant Sandbox / Evaluator Mode)**: Ships with a zero-setup, deterministic cloud state emulator (`mock_cloud.json`), allowing hackathon judges to verify end-to-end functionality in under 30 seconds without needing AWS credentials.
3. **Interactive FinOps Governance Interface**:
   * Built with FastAPI and a modern dark-themed web terminal featuring real-time Strands Agent reasoning traces, one-click authorization, and state rollbacks.

---

### 🧗 Challenges We Ran Into
* **Balancing Full Autonomy with Zero-Risk Safety**: Giving an AI agent permissions to execute cloud modifications (`ec2:StopInstances`, `ec2:ReleaseAddress`) is dangerous. We solved this by designing a **non-bypassable two-phase commit protocol**: the agent can

only generate cryptographically verified Action Proposals; the remediation tool strictly refuses execution unless presented with a valid operator approval signature. * Deterministic Tool Schemas in Strands Agents SDK: Ensuring the agent consistently produces well-structured JSON payloads with numeric risk ratings and exact resource IDs required strict Pydantic schemas and defensive validation in the tool layer. * Zero-Friction Evaluator Experience: We wanted hackathon judges to experience the live Strands reasoning loop immediately. Implementing the deterministic fallback engine ensured that anyone can run python3 server.py and see the complete agent in action within seconds.

---

### 🏆 Accomplishments That We're Proud Of
* **Immediate Financial Impact**: In our benchmark simulation, AuditGuard recovered **$2,417+ per month** in waste ($29,000+ annualized budget returned to R&D) across forgotten GPU instances and unattached disks.
* **True "Agents for Humans" Implementation**: Built an agent that augments human oversight rather than replacing it or creating more busywork.
* **100% Rollback Guarantee**: Zero production outages—every disruptive remediation action automatically triggers an immutable EBS volume snapshot before state change.

---

### 🎓 What We Learned
* **The Power of Strands Tool Decorators**: Building modular tools with the AWS Strands Agents SDK drastically simplifies LLM orchestrations compared to raw prompt engineering.
* **Human-Agent Collaboration Patterns**: The best agents don't hide their reasoning; exposing real-time thought traces, calculated risk percentages, and specific commands builds trust with DevOps and FinOps engineers.

---

### 🚀 What's Next for AuditGuard
* **Slack / Teams Conversational Approval Bot**: Enable engineers to approve remediation cards directly within Slack with interactive buttons (`/auditguard approve <token>`).
* **Predictive Right-Sizing**: Expand beyond idle resource shutdown to proactive instance rightsizing recommendations based on 30-day CloudWatch percentile trends.
* **Multi-Account AWS Organizations Support**: Extend the Strands Agent across multi-account AWS Landing Zones with centralized FinOps budget enforcement.

---

### 🔗 Project Links & Open Source
* **GitHub Repository**: https://github.com/masato25/auditguard (MIT Licensed)
* **AWS Builder Center Article**: [Building AuditGuard with AWS Strands Agents SDK](https://builder.aws.com/content/3JJ9XpQMh1lcwzpcUbdaOvcguqe/building-auditguard-the-autonomous-finops-waste-patrol-with-aws-strands-agents-sdk)

Built With

Share this project:

Updates

Submission history