Terraform Guardian
I've spent a fair amount of time deploying machine learning pipelines with GitHub Actions and AWS, and the ML code was rarely the hard part. The infrastructure was. Three separate times, I hit the same wall: a breaking change involving GitHub OIDC and AWS, mid-deployment. Each time I'd dig back through the config trying to figure out what moved — the IAM role, the trust policy, the permissions, the resources being provisioned, whether GitHub Actions was even assuming the right role. None of those questions are hard on their own. What makes them hard is how connected they are. Change one thing and the effect shows up somewhere else entirely. And when you're trying to ship an ML pipeline, the last place you want to spend your time is debugging infrastructure instead of the model.
That's what got me asking a different question: what if the pipeline reviewed changes before they hit production — not just checking that Terraform was syntactically valid, but understanding what a change actually meant? A new storage bucket and a widened IAM trust policy are not the same kind of risk, and a small cost increase isn't automatically a problem while a large unexpected one should raise a flag. I wanted a system that could tell those apart. That's Terraform Guardian.
What I built
Not another Terraform wrapper, not another deployment dashboard. Guardian sits between a terraform plan and the apply step. A developer opens a pull request with Terraform changes, the pipeline generates a plan, Guardian reviews it, and safe changes flow through automatically while risky ones get held for a human. It's built to pay attention to the parts of infrastructure that are easy to skim past: IAM roles and permissions, GitHub OIDC trust policies, new resource types, and cost.
The architecture is event-driven and split across two repos. The watched repository holds the Terraform code being protected and runs the GitHub Actions workflow. The Guardian repository holds the Lambda agent and the AWS infrastructure that reviews the plans it receives.
On a pull request, the workflow assumes a scoped plan-submitter role over OIDC, runs terraform plan -out=tfplan.binary, converts it to JSON with terraform show -json, uploads it to S3, and drops a message on SQS with the repo, PR, commit SHA, and plan location. That plan-submitter role can submit plans and nothing else — it's deliberately not the same role that applies infrastructure, because a workflow that reads a plan shouldn't carry the same permissions as one that changes production.
SQS triggers a containerized Lambda running the Guardian agent, built on Strands. The agent pulls the plan from S3 and works with three tools: terraform_plan_parse extracts resource changes and config diffs from the JSON; terraform_cost_estimate estimates the monthly cost impact, using Infracost when available and a heuristic fallback when it isn't; guardian_decision runs the plan against a deterministic escalation matrix and returns auto_apply or hold. If the agent can't assess a change reliably, it holds — safety by default, not by exception.
A held plan triggers an SNS notification and, when the token is configured, a comment on the GitHub PR. The decision also lands in DynamoDB, so there's a history and a place for humans to override it later.
A concrete case
Say a developer widens the GitHub OIDC trust policy on a deployment role — previously scoped to one repo and branch, now looser. Terraform applies cleanly. AWS reports no error. The deployment succeeds by every technical measure, and the security boundary still changed. Guardian catches that and returns:
Decision: hold
Risk: OIDC trust policy widened
Action: Human review required
Compare that to bumping a Lambda's memory allocation with cost still under threshold — that one goes straight through as auto_apply. Not every change deserves the same friction, and building a system that can tell the difference was the actual point.
What I learned building it
Infrastructure safety isn't just "did the deployment succeed." It can succeed technically and still be worse off — the OIDC example above is exactly that. I also learned that infrastructure config behaves like a system of relationships rather than a pile of independent files: an IAM role depends on its trust policy, the trust policy decides who can assume it, and GitHub Actions depends on all of that lining up correctly to authenticate at all. When something breaks, the cause is rarely visible in the deployment error itself. A role can exist and still fail to be assumed. A trust policy can be valid JSON and still not permit the identity you meant it to. A permission can be missing even when authentication succeeds.
That pushed a design principle I hadn't started with: a useful infrastructure tool has to explain why something is risky, not just report that it changed.
Designing the decision engine was its own problem. Hold everything and the tool becomes something people route around. Approve too much and it isn't doing its job. I landed on an escalation matrix with an OR-gate structure — different risk signals can each independently trigger a hold, so an OIDC widening escalates a plan on its own even if the cost impact is zero. That killed the idea I started with, that infrastructure risk could collapse into one number.
The rest of the build was less conceptual and more plumbing: getting the Lambda container image working meant fighting through ECR not finding the image, then Lambda rejecting the manifest and layer media types. None of that was about the agent logic, but it still had to work before any of the agent logic mattered. For demos, I wrote demo.py to run the decision engine locally against captured plan fixtures — python3 demo.py tests/fixtures/plan_safe.json for the clean path, python3 demo.py tests/fixtures/plan_oidc_trap.json for the held one — so I could show the reasoning without depending on a live AWS deployment during a recording.
Where it's going
Right now Guardian stops at detection: it flags the widened trust policy and holds the plan. The next version should keep going. If an OIDC trust policy gets widened, Guardian should explain what changed, identify the security boundary that was intended, and recommend the specific fix — Detect → Explain → Recommend → Human Approval → Re-plan. A developer reviews the proposed fix, approves it, and Guardian generates the corresponding Terraform change, which then goes through another plan and the same risk checks before anything applies. I want that loop human-approved end to end — the agent proposes, it never applies on its own.
Longer term, I want Guardian to learn from its own decision history: past holds, overrides, and recurring deployment failures, so it gets sharper at recognizing the configuration patterns that actually cause problems instead of just the ones I thought to encode by hand.
It started with the same OIDC break happening to me three times. I built this so the fourth time would go differently.
Built With
- amazon-web-services
- lambda
- python
- strands
- terraform
Log in or sign up for Devpost to join the conversation.