Inspiration
Cloud teams usually have two unconnected dashboards: one for security posture, one for cost. A misconfigured S3 bucket and an idle $560/month EC2 instance get treated as completely separate problems, even though both are really the same failure - nobody's watching the infrastructure closely enough, often enough. For Hilti ImagineHack 2026 (Track 2), we wanted to build something that watches both at once, explains why something's wrong in plain English, and critically - only acts on its own for things that are safe to undo. Anything irreversible waits for a human.
What it does
Feronia scans real AWS infrastructure with two parallel AI agents SecOps (security) and GreenOps (cost + carbon) and turns their findings into a concrete action plan. A deterministic validation layer (the Gatekeeper) checks every recommendation before it's allowed to exist, deduplicates and sorts everything, then pauses for human approval before executing anything destructive (terminating an instance, deleting a volume). Everything else resizing, encryption, security group fixes runs automatically, because it's reversible.
How we built it
The pipeline is a LangGraph state machine: ingest → build resource graph → route → SecOps/GreenOps (parallel) → Gatekeeper → Synthesizer → human approval gate → execute → report. We deliberately split judgment from verification: agents propose findings using real LLM calls, but a 6-check, zero-LLM Gatekeeper decides whether each one is actually trustworthy before it reaches the action plan.
SecOps grounds itself in named, real standards instead of inventing its own rules: CIS AWS Foundations Benchmark v3.0, CVSS v3.1 for severity, MITRE ATT&CK technique mapping, and an IAM privilege-escalation check built on permission-combination research published by the team behind CloudGoat - the same vulnerable-AWS tool the security industry uses to teach these exact attack paths.
GreenOps follows the Cloud Carbon Footprint open-source methodology rather than a made-up formula:
$$ \text{Average Watts} = W_{min} + u \cdot (W_{max} - W_{min}), \qquad CO_2e = \frac{\text{Watts} \times h}{1000} \times \text{PUE} \times I_{grid} $$
using CCF's published AWS coefficients and the real grid carbon intensity for our deployment region, scaled by each instance's actual measured CPU utilization rather than a flat assumption.
Once the demo was working against mock data, we wired it to real AWS - two separately scoped IAM identities (a read-only scanner, a write-limited executor restricted to our specific demo resources), so the whole pipeline genuinely scans and remediates a real, intentionally-vulnerable Terraform environment, not a JSON fixture.
Challenges we ran into
- LLM non-determinism, even at temperature=0. The same scan, run twice, could recommend resizing an instance one time and terminating it the next. We fixed the underlying ambiguity (overlapping "zombie" vs. "right-size" thresholds) rather than papering over symptoms.
- A LangGraph reducer bug that inflated our numbers. Our parallel-agent fan-out needed an
Annotated[list, operator.add]field - but Gatekeeper and Synthesizer were accidentally writing to that same field, so every downstream stage kept appending instead of replacing, and a real run of 6 findings reported as 18. - A real action/resource-type mismatch. GreenOps once recommended
terminate_instancefor an EBS volume - harmless in our mock JSON, but the kind of bug that matters once you're pointed at real AWS. We extended the Gatekeeper with a 6th check validating that every action is even possible for the resource type it targets. - Regional cost accuracy. Our cost model was quietly using US pricing for resources deployed in
ap-southeast-1, understating real savings by 20–30%. - A surreal macOS permissions incident mid-build, where a burst of concurrent file operations left our virtual environment unable to even resolve its own working directory - a reminder that not every bug lives in your code.
- Merging two parallel work streams without losing either. While we built the agent logic, a teammate independently restructured the whole repo into a monorepo. Reconciling weeks of bug fixes against a renamed directory tree, file by file, was its own exercise in patience.
What we learned
The most reliable way to make an agentic system trustworthy isn't a smarter prompt - it's a deterministic layer that never trusts the LLM's output until it's checked. Every real bug we found and fixed lived in the gap between "the agent said something plausible" and "the agent said something true," and the Gatekeeper is what closed that gap. We also learned that borrowing real, named industry standards (CIS, CVSS, MITRE, CCF) instead of inventing heuristics isn't just more credible to outsiders - it gives you an actual target to verify against when something looks wrong, which is exactly what let us catch and fix our own mistakes along the way.
Log in or sign up for Devpost to join the conversation.