💡 Inspiration

I spent the last stretch of my Placement on greenfield projects that came out of the Cloud migration, and one failure mode never stopped stinging: a Terraform PR would come back fully green from terraform plan, get reviewed, get merged — and then apply would die halfway through because the deploying service account was missing a permission, or a quota was already exhausted. The plan had no idea. Nobody did, until it was already expensive.

The IAM side was worse. We'd moved from Google Groups to Entra time-bound roles, and the policy checks never made it clear what could and couldn't be onboarded as a role. Devs were guessing, security was gatekeeping, and everyone was frustrated. PreFlight is the tool I wished we'd had: something that reads the change with the environment's ground truth and says plainly what will break and why.

🔍 What it does

PreFlight is a pre-merge check for Terraform on GCP. It takes three inputs — the plan JSON, the Terraform source, and a small environment-context file (the deploying service account's roles, live quota usage, and org policies) — and catches what terraform plan can't see:

  • Predicted apply failures — permission gaps that will 403 at apply time, with the exact missing permission named
  • Quota walls — creates that would exceed a live quota limit
  • IAM / org-policy violations — explained in plain English, including permanent grants of privileged roles that policy says must be time-bound

The differentiator is the remediation loop: PreFlight proposes a fix as a diff, applies it only to an in-memory copy, and re-analyses the patched config — a fix is never marked resolved just because it looks plausible. It's also precise rather than noisy: the sample's allowed member domain is deliberately not flagged.

🛠️ How I built it

I built it with Codex from a written build brief, starting from a JSON Schema contract and a synthetic "hero scenario" before any engine code existed. At runtime, GPT-5.6 Terra is the reasoning engine: an ANALYSE pass produces findings as strict Pydantic structured output, a CRITIQUE pass challenges every finding against its evidence and drops anything unsupported, and a REMEDIATE pass re-analyses the patched config to verify fixes.

The division of labour that made it work: Terra owns the diagnosis, evidence, explanations, and fixes; Python owns stable finding IDs, the severity taxonomy, schema validation, and patch safety. The patcher rejects unknown files, path traversal, and ambiguous hunks, and never touches the real source. It's deployed on DigitalOcean App Platform with prompt caching and a content-hash response cache to keep model costs down.

Codex carried far more of this than I expected — it turned the schema and sample scenario into the CLI, the loaders, and the test suite in a fraction of the time it would have taken me, and even the demo video was cut with ffmpeg commands Codex generated.

⚔️ Challenges I ran into

The one that taught me the most: Terra kept diagnosing the right three issues, but between runs it would reorder them, change their IDs, or pick a different severity for the IAM finding. My first instinct was to prompt harder. The actual fix was to stop fighting the model for determinism — let it own the diagnosis, and let Python canonicalize IDs, ordering, counts, and the severity taxonomy after validation.

Deployment had its own drama: the London region's containers couldn't resolve DNS for pypi.org or api.openai.com at all, which I only worked out after staring at "Temporary failure in name resolution" for far too long. Moving the app to Frankfurt fixed it in one deploy. Synchronous requests also timed out during model runs, so analysis moved to background jobs the UI polls.

🏆 Accomplishments I'm proud of

A 15-test deterministic suite that passes against the model's real behaviour, a verified remediation run in production, and a live public demo — built solo in under a week. The moment the CRITIQUE pass deleted one of its own unsupported findings for the first time was genuinely satisfying.

📚 What I learned

Don't fight a probabilistic model for determinism — give it the reasoning and give the code the contract. Structured outputs plus a self-critique pass plus deterministic canonicalization turned "impressive but flaky" into something I could write regression tests against. This was also my first time deploying on DigitalOcean App Platform and my first serious use of OpenAI's structured-output and prompt-caching APIs.

🚀 What's next for PreFlight

A GitHub Action that posts the verified report as a required pull-request check, and populating the environment context live from read-only GCP IAM, quota, and org-policy APIs instead of a hand-written file. After that: broader Terraform resource coverage, auth and audit logs for team use, and other clouds once the GCP workflow is operationally complete.

Built With

Share this project:

Updates

Submission history