Seven Stage Infrastructure as Code Security Cheatsheet for DevOps

September 09, 2026•11 min read

Isometric infrastructure security title card

Infrastructure as code security is the practice of applying shift-left controls, static scanning, and policy-as-code enforcement to infrastructure definitions before they ever provision a resource. The single highest-leverage action any team can take is gating pull requests and CI pipelines on new critical findings using a scanner paired with a policy engine such as Open Policy Agent. Anchor this work to the OWASP Infrastructure as Code Security Cheatsheet, not ad hoc rules built in isolation.


TL;DR:

  • Gating on new critical findings is essential to prevent bottlenecks, but overblocking the entire backlog can halt deployment pipelines altogether.

  • State files must be stored in encrypted remote backends with restricted access, as they can expose secrets and internal resource details if mishandled.

  • Incorporating IaC scans at multiple pipeline checkpoints, including IDE, PR, CI gate, and drift detection, significantly reduces infrastructure misconfigurations and security incidents.

  • Medium-term improvements include establishing a shared policy library, enforcing plan reviews before production, and adopting workload identity federation to eliminate long-lived credentials.

  • Success is measured by reduced time-to-remediate, lower drift rates, and higher gate pass rates, with focus on gating strategy more critical than the choice of scanning tools.


Projectjth

Bring Clarity to IaC Security

Projectjth provides technical advisory services and risk assessments for companies managing security and operational challenges.

Explore Projectjth

Table of Contents

What Infrastructure as Code Security Actually Protects

Infrastructure as code security means treating Terraform, CloudFormation, Pulumi, or Bicep files with the same scrutiny you’d apply to application code, because a single misconfigured module can replicate across every environment it touches. This is the shift-left argument in concrete form: catch a public S3 bucket definition in a pull request, and you prevent it from existing anywhere. Catch it after deployment, and you’re chasing it across dev, staging, and production simultaneously.

The blast radius problem is what separates IaC risk from ordinary application vulnerabilities. One flawed module referenced by twenty services doesn’t create twenty separate bugs. It creates one bug with twenty consequences, often provisioned in minutes by an automated pipeline nobody double checked.

IaC security complements, rather than replaces, runtime tools. A Cloud Security Posture Management (CSPM) platform and dynamic application security testing (DAST) both operate after resources exist. IaC scanning operates before they do. Microsoft’s DevSecOps guidance describes correlating IaC scanning with Defender for Cloud so pre-deploy findings and runtime posture inform each other rather than living in separate dashboards.

Where IaC scanning fits relative to other layers:

  • Static IaC analysis catches misconfigurations before provisioning.

  • CSPM and CNAPP tools catch drift and exposure after resources exist.

  • DAST catches application-layer flaws once services are running.

Common Risks and Failure Modes in IaC

Most IaC incidents trace back to a small set of repeat offenders. Knowing them by name makes code review faster and audits more targeted.

  1. Hardcoded secrets. API keys and credentials committed directly into .tf or .yaml files persist in git history long after anyone remembers deleting them.

  2. Exposed state files. Terraform state can contain plaintext secrets, resource IDs, and internal network mappings, and must live in an encrypted remote backend with tightly restricted read and write access.

  3. Over-privileged IAM. Wildcard permissions and broad trust policies get written once “to make it work” and rarely get scoped down afterward.

  4. Publicly exposed resources. Storage buckets, databases, and security groups open to 0.0.0.0/0 are still one of the most common findings in any IaC scan.

  5. Configuration drift. Manual console changes bypass version control entirely, so the deployed environment no longer matches what your repository claims is true.

  6. Supply-chain exposure. Compromised scanner artifacts or mutable container tags let an attacker inject malicious logic into the tooling meant to catch problems.

  7. Template rendering pitfalls. Helm charts and Terraform modules that interpolate values dynamically can silently produce insecure defaults that no static read of the source file would reveal.

Pro Tip: Run a search across your entire repository history for patterns like AKIA or -----BEGIN PRIVATE KEY-----, not just the current branch. Secrets committed and later deleted are still sitting in git history unless the history itself was rewritten.

State files deserve special attention because they’re often treated as build artifacts rather than production assets, even though they can expose everything a scanner would flag in the source code plus the actual values behind them.

Best Practices by Lifecycle Phase

Security controls belong at every stage of the IaC lifecycle, not bolted on as a single pre-deploy gate. Here’s how to distribute them.

Authoring. Developers should get feedback before they ever open a pull request. IDE scanning plugins, a library of vetted secure modules, and a strict no-secrets-in-code policy catch the cheapest fixes at the cheapest point in the process. CODEOWNERS rules on infrastructure directories add a mandatory second set of eyes.

PR and CI. This is where policy becomes enforceable. Gate merges on new critical findings rather than the entire historical backlog, or you’ll create a pipeline nobody can merge into. Pair static scanning with secret detection and branch protection rules that block direct pushes to main.

Plan and review. Never apply blind. Reviewing the terraform plan output before it runs shows exactly what will change, and production applies should require a named human approval step, not just an automated pass.

Apply and runtime. Use OIDC or workload identity federation instead of long-lived service account keys, so a compromised runner exposes a short-lived credential rather than a permanent one. Restrict who can trigger an apply, and run scheduled plan checks against live state to catch drift before it compounds.

State and secrets. Store state in an encrypted remote backend, isolate workspaces per environment, and integrate a secrets manager rather than environment variables scattered across CI configuration.

  • Use per-environment, per-account credentials so a single compromised pipeline can’t reach every environment at once.

  • Lock state during applies to prevent concurrent, conflicting writes.

  • Rotate credentials on a schedule rather than only after an incident.

Where to Run IaC Scans and Which Tools Belong There

Scanning at a single point in the pipeline is the most common gap platform teams miss. Coverage needs to span six distinct checkpoints, each catching what the others can’t.

  • IDE and pre-commit hooks give developers instant feedback before code ever reaches a shared branch, using scanners like Checkov, KICS, or Trivy running locally.

  • PR-level scanning posts inline comments on the exact line that triggered a finding, which shortens the loop between detection and fix.

  • CI gates run the comprehensive rule set and block merges on new critical issues, ideally paired with a policy engine such as OPA or HashiCorp Sentinel for organization-specific rules.

  • Pre-deploy plan checks review the actual terraform plan diff against policy before anything touches real infrastructure.

  • Post-deploy drift detection runs scheduled plan checks against live state and feeds a CSPM or CNAPP platform for correlation with runtime posture.

Tool maturity tends to progress in stages: starter teams run scanning only in CI, intermediate teams add IDE feedback and compliance mappings, and advanced teams enforce policy-as-code with runtime drift correlation built in.

Supply-chain hygiene matters as much as the scanning logic itself. Pin scanner versions, verify checksums, and avoid mutable tags, since a compromised scanning tool defeats the entire point of running one.

Mapping the Seven-Stage IaC Security Pipeline

A staged pipeline model catches more than any single scan point, and teams that adopt this structure report fewer production incidents tied to infrastructure misconfiguration.

  1. Pre-commit. Local scanning and secret detection stop the cheapest mistakes before they leave a developer’s machine.

  2. Branch protection. Required reviews and status checks prevent unreviewed infrastructure changes from landing on main.

  3. PR scanning. Inline findings give the author immediate, contextual feedback tied to specific lines.

  4. CI gate. The comprehensive scan and policy evaluation run here, blocking merges on new critical issues while tracked backlog items move to remediation sprints instead of stalling the pipeline.

  5. Plan review. A human reviews the actual diff against live infrastructure before approving an apply.

  6. Post-deploy drift detection. Scheduled plan checks compare live state against source of truth on a recurring basis.

  7. Application validation. Runtime and DAST checks catch what static analysis structurally cannot see.

Gating strategy matters more than tool choice at this stage. Blocking on the full historical backlog produces an unmergeable pipeline; blocking on new critical findings while working the backlog on a schedule keeps momentum without ignoring risk.

A Practitioner’s Checklist From the Field

Engagements that focus on IaC security tend to converge on the same fixes: standardized secure module defaults, a gating policy that blocks new critical findings without freezing the backlog, a centralized policy library instead of per-team rules, and a measurable drop in drift incidents within the first few review cycles.

Quick wins worth doing this month:

  • Add IDE or pre-commit scanning so developers see findings before they open a PR.

  • Move to OIDC or workload identity federation to eliminate long-lived deploy credentials.

  • Migrate state files into an encrypted remote backend with restricted access.

Medium-term work centers on building a shared policy-as-code library and requiring plan review before production applies. Longer-term maturity means correlating IaC findings with a CNAPP platform for full posture visibility. Track progress with concrete numbers: time-to-remediate on new findings, drift rate per environment, and the ratio of gated PRs to total infrastructure merges.

Why the Standard Advice Undersells the Hard Part

Most guidance on this topic treats IaC security as a scanner selection problem: pick a tool, wire it into CI, done. That’s the easy 20 percent. The hard part is gating strategy, and almost nobody talks about it honestly.

Why the Standard Advice Undersells the Hard Part — overview diagram

Teams that block every existing finding on day one build a pipeline nobody can merge into, and within weeks someone quietly disables the gate entirely. The teams that succeed block only new critical findings while working the existing backlog through tracked remediation sprints. That distinction isn’t a technicality. It’s the difference between a control that survives contact with a real engineering org and one that gets bypassed by month two.

The other underrated point: state file exposure gets far less attention than public S3 buckets, despite often being worse, since a state file can contain the actual secret values a scanner would only flag as a pattern elsewhere. Fix your gating strategy and your state backend before you shop for another scanner. Everything else is secondary.

— Jesse

How Projectjth Approaches IaC Security Implementation

Buying a scanner is the easy part. Getting gating rules, policy-as-code, and secure module defaults actually adopted across a real engineering team, without breaking their delivery pace, is where most implementations stall. Technical advisory work focuses on closing that gap: pipeline hardening, policy-as-code rollout, and Linux infrastructure remediation tailored to team workflows rather than a generic playbook.

Projectjth

An engagement typically starts with a risk assessment of your current IaC pipeline, mapping where scanning, secret detection, and state management gaps actually sit against the seven-stage model. From there, the deliverables are concrete: a gating policy tuned to your backlog size, a centralized module library, and a remediation plan your platform team can maintain after the engagement ends. Success gets measured the same way described above: time-to-remediate, drift rate, and gate pass rates, not vague assurances. If your infrastructure automation has grown faster than your security controls have kept up, start a conversation with Projectjth about a pipeline security assessment.

Sources

FAQ

Is CISA Still a Thing in IaC Security Guidance?

Yes. The Cybersecurity and Infrastructure Security Agency remains an active source of guidance on securing cloud and automated deployment pipelines, and its recommendations often align with the shift-left principles found in the OWASP IaC cheatsheet.

Is IaC Part of CI/CD?

Yes, IaC definitions are typically stored in version control and deployed through the same CI/CD pipeline as application code, which is why scanning and policy gates belong directly inside that pipeline rather than as a separate process.

What Is IaC Security Scanning?

IaC security scanning is static analysis of infrastructure definition files (Terraform, CloudFormation, Bicep) that flags misconfigurations, exposed secrets, and policy violations before resources are provisioned.

What Is the Best IaC Tool?

There’s no single best tool. The strongest setups pair a static scanner (Checkov, KICS, or Trivy) with a policy engine like Open Policy Agent, layered across IDE, PR, and CI checkpoints rather than relying on one tool at one stage.

How Does IaC Security Differ From Runtime Cloud Security?

IaC security catches misconfigurations before deployment, while runtime tools like CSPM and CNAPP platforms detect drift and exposure in resources that already exist, and the two layers work best when their findings are correlated together.

Back to Blog

Put It Into Practice

Find Your Next Clear Step.

Need help with a technical challenge or a leadership event? Tell Jesse what you are working through and start a practical conversation.

Start a Conversation

PROJECT-JTH LLC · Hoover, Alabama
Calmer Leadership. Better Systems. Practical Tools.

Home   /   Contact