Inspiration

FinOps tools and AWS cost reports are good at finding waste idle EC2, orphaned EBS, unused EIPs, and similar patterns. What usually does not get closed is the recovery loop: investigate root cause, assemble proof, decide what action is safe, get an authorized human to sign off, and record what was recovered.

We built Recoup for the Agents for Humans hackathon as a human-in-the-loop spend recovery operator: real read-only AWS scanners, evidence-backed recovery cases, deterministic safety and policy gates, and claim-bound approval before anything moves to Recovered in the ledger.

AWS provides the FinOps intelligence; Recoup closes the recovery loop.

What it does

Live demo (J-FULL - no AWS credentials required):

  1. Open Demo UI → consent → Demo Scan (9 parallel waste scanners).
  2. Review findings on Opportunities (~$87/mo detectable waste in the connected demo account).
  3. Start Recovery on findings (try three different services, e.g. EC2, EBS, RDS).
  4. On each opportunity detail page: Approve, Investigate further, or Decline - human-in-the-loop binds claim hash, amount, and state version (tampering → 409).
  5. Open Recovery Ledger - Remaining → Pending Approval → Recovered reflects your decisions.

On Approve, Recoup records recovery, updates the ledger, and sends an SNS recovery report. The public demo emphasizes governed closure and auditability (estimated monthly savings from scan), not unattended remediation on every resource type.

How we built it

  • Frontend: Next.js operator UI (scan, opportunities, claim-bound HITL on opportunity detail, Recovery Ledger).
  • Backend: FastAPI- scan orchestration, recovery assessment pipeline on promote, approvals, demo sessions (X-Demo-Session), quality scorecard.
  • AWS: STS AssumeRole, 9 read-only waste scanners (EC2, EBS, EIP, RDS, S3, Lambda, load balancers, CloudWatch Logs, Cost Explorer), DynamoDB, SNS, App Runner hosting.
  • Agents: Strands Agents on Amazon Bedrock - full 11-node graph for SLA credit recovery (golden replay in CI); live promote path uses a deterministic recovery assessment pipeline through the graph policy gate for reliable judge latency (RECOVERY_LLM_ON_PROMOTE=false in production).
  • Trust: Cedar-style default-deny policy (in-repo rules + deterministic evaluation on App Runner), evidence sanitizer, autonomy classes, zero unsafe external actions ship gate.

Two paths, one product: (1) Demo path - scan → promote → policy → HITL → ledger; (2) Depth path — optional POST /api/opportunities/{id}/run with use_strands for the canonical SLA scenario and CI golden replay.

Challenges we ran into

  • Latency vs. depth: Judges need a fast, reliable UI path; we default promote to the deterministic pipeline while keeping the full Strands graph provable in CI and optional via API.
  • Financial trust: Approvals must bind exact claim content, we enforce claim_hash, amount, and state_version, with Playwright adversarial tests for tampering.
  • Safe demo at scale: Per-guest demo sessions, masked account IDs, read-only scanners in the public flow, and production blocks on dangerous test reset endpoints.
  • Honest “recovery” language: V1 records estimated savings from scan and operator-approved closure; we separate that from claiming automated fix on every resource type.

Accomplishments we're proud of

  • End-to-end J-FULL journey on production App Runner, plus 127 Playwright tests (including full discovery → triage → ledger).
  • 419 backend tests, six-gate quality scorecard (golden SLA replay through the Strands graph, financial correctness, evidence recall, trace completeness, zero unsafe actions).
  • Real AWS scanner breadth, claim-bound HITL, Recovery Ledger, and SNS notification on approve.
  • Three AWS Builder articles documenting architecture, trust, and proof (links below).

What we learned

Autonomous cloud spend recovery is not “let the agent delete resources.” It is investigate → prove → policy → approve → record → verify. Strands and Bedrock shine on packaging and reasoning; deterministic gates must own money, policy, and execution boundaries. Testing the full operator journey mattered as much as the agent graph.

What's next for Recoup

  • Deeper realized vs. estimated savings in the ledger after approved remediation paths.
  • Broader remediation behind the same Cedar and HITL model for additional services.
  • Customer account onboarding (ExternalId cross-account role) beyond the hosted demo session model.
  • Optional AgentCore deployment path aligned with our CDK/IAM tool registry.

Try it

URL
Live demo https://7j3qm5yyga.us-east-1.awsapprunner.com/scan
GitHub https://github.com/swa01wk/recoup
Judge walkthrough https://github.com/swa01wk/recoup/blob/main/docs/judge-demo.md

Builder community (Agents for Humans):

  1. Building Recoup - Strands graph & safe recovery workflow
  2. Trust - Cedar, evidence redaction, HITL
  3. Proof - 419+127 tests & scorecard

Built With

  • agents
  • bedrock
  • cdk
  • cloudwatch
  • dynamodb
  • ec2
  • fastapi
  • lambda
  • next.js
  • python
  • rds
  • s3
  • sns
  • strands
  • typescript
Share this project:

Updates

Submission history