Infrastructure Drift: The State File Isn't an Audit
Share this post

Infrastructure Drift: The State File Is Not an Audit Record

A platform team at a mid-sized fintech company — a hypothetical composite built from patterns common enough to be unremarkable, not a real client engagement — had a green Terraform pipeline for eleven straight months. Every pull request against the infrastructure repository triggered a plan, every plan was reviewed, every apply succeeded, and the dashboard the VP of Engineering checked before board meetings showed exactly what an executive wants to see: zero pending changes, zero drift warnings, a clean state.

During those eleven months, an on-call engineer opened port 5432 to 0.0.0.0/0 on a database security group at two in the morning, chasing a connectivity failure during an incident nobody wanted to extend by waiting for a pull request to merge. The fix worked. The incident closed. The engineer meant to open a follow-up ticket to revert the change and file the correct rule through Terraform. The ticket never got filed, because the incident retro focused on the application bug that caused the outage, not on the infrastructure workaround used to diagnose it. The security group stayed open to the internet for the next nine months, invisible to every plan the team ran, because nobody touched that particular Terraform resource again, and a security group rule that Terraform did not need to reconcile was a security group rule Terraform never looked at.

This is not a story about a careless engineer. It is a story about what a CI/CD pipeline actually verifies when it validates an infrastructure-as-code plan, and what it does not.

What "the pipeline is green" actually means

Most engineering organizations that have adopted Terraform, Pulumi, CloudFormation, or OpenTofu have also adopted an unstated assumption alongside the tooling: if the pipeline shows no planned changes, production infrastructure matches the code. That assumption is wrong often enough, and wrong in high-consequence enough ways, to be worth taking apart mechanically rather than treating as a minor edge case.

Start with what actually happens when a CI/CD pipeline runs a Terraform plan. Terraform maintains a state file — a JSON record mapping every resource in your configuration to a real object in your cloud provider, along with the last known values of that object's attributes. By default, running terraform plan (and terraform apply) performs a refresh: Terraform queries the provider's API for each resource already tracked in state and updates its in-memory record of that resource's current attributes before computing what changes, if any, are needed to make reality match the configuration. HashiCorp's own documentation describes this refresh as "reconciling the resources tracked by the state file with the real world," and by default, terraform plan "compares your state file to real infrastructure" every time it runs.

That is a real, functioning check, and it is not nothing. For resources Terraform already tracks, a plan run genuinely does query the live cloud API and will surface a difference between what is declared and what is actually running, at the moment the plan executes. The gap is not that this mechanism is fake. The gap is in when it runs, what it covers, and what a passing result actually certifies — three separate limitations that compound into the eleven-month blind spot in the fintech example above.

The gap is temporal: plans run on code push, not on a schedule

The first and most consequential limitation is timing. In the overwhelming majority of organizations, a Terraform, Pulumi, or CloudFormation plan runs because someone changed the code — opened a pull request, pushed a commit, or triggered a deploy. Nothing runs a plan simply because time has passed. If nobody touches the module that manages a particular security group, load balancer, or IAM role for nine months, no plan checks that resource for nine months, refresh mechanism or not. The refresh only fires when a plan fires, and a plan only fires on a code event.

This is not a hypothetical gap invented for this article. It is the specific problem HashiCorp built a standalone product feature to solve. When HCP Terraform's Drift Detection reached general availability, HashiCorp's own announcement described it as continuously monitoring infrastructure so that "when infrastructure changes happen outside of the Terraform workflow," teams find out through a scheduled check rather than by accident — a capability layered explicitly on top of ordinary plan-on-push behavior, because ordinary plan-on-push behavior does not provide it. The same announcement is candid about the consequence of not having it: without a scheduled check, "applications can suddenly crash, deployments can unexpectedly fail, thousands of dollars in monthly costs can be wasted due to unused resources, and systems or unknown resources can be left open to public access" before anyone notices, because nothing was watching between deploys. That is a vendor's own framing of the problem this article is about, and it is consistent with how the mechanism actually works: routine CI/CD plan checks are event-triggered, not continuous, and the interval between events is exactly the window in which unreviewed change accumulates.

Pulumi's own materials describe the identical structural gap from the other side of the tooling landscape. Pulumi distinguishes between three states worth naming precisely, because the vocabulary clarifies where the failure actually sits: the desired state is what your code says should exist, the current state is what your state file (or Pulumi's managed state backend) last recorded, and the actual state is what is really running in the cloud account right now. Drift, in Pulumi's own definition, is what opens up "when the cloud resources running in your account no longer match the desired state described in your infrastructure as code" — and critically, the current recorded state can itself be stale relative to the actual running state for as long as nobody runs a refresh, which is exactly the interval between code pushes.

Pulumi's recommended technical response is telling: pulumi refresh --preview-only, run on a schedule independent of any code change, specifically to query live provider state and compare it against the recorded state without waiting for someone to touch the code. The --expect-no-changes flag exists so this can be wired into CI as a pass/fail gate on its own, decoupled entirely from a deploy pipeline. Pulumi built this as a distinct, schedulable check for the same reason HashiCorp built Drift Detection as a distinct product feature: the ordinary deploy pipeline does not provide it.

The gap is scoped: unmanaged resources are invisible by construction

The second limitation is more structural than the first, and less fixable by adding a schedule. HashiCorp's own tutorial on managing resource drift states the boundary plainly: "Terraform cannot detect drift of resources and their associated attributes that are not managed using Terraform." A resource created directly in the AWS console, provisioned by a script that never wrote to Terraform state, or spun up by a different team's tooling entirely, does not exist as far as any Terraform plan is concerned. It cannot drift from a declared state, because no state declares it. It is not flagged, not warned about, not visible in any dashboard built on top of Terraform state — it simply is not part of the universe the tool reasons about.

This matters more than it sounds like it should, because the resources most likely to be created this way are exactly the resources created under the conditions least conducive to careful review: incident response, a one-off proof of concept that quietly became load-bearing, a security exception granted verbally and implemented by hand rather than routed through the normal change process. The database security group in the fintech example above was, in that specific case, already Terraform-managed — the rule that got added by hand was a manual mutation of a managed resource, which is the more common and in some ways more dangerous pattern, because it looks like drift a plan should catch and frequently does not, for the third reason below. But plenty of organizations also carry a population of entirely unmanaged resources: a load balancer someone stood up during a launch crunch and never got around to importing, a database instance created directly because Terraform's provider didn't yet support a needed feature, an S3 bucket created by a data science team outside the platform team's Terraform root modules entirely. None of it shows up in a drift report, because a drift report is fundamentally a comparison against a list, and these resources were never added to the list.

AWS CloudFormation's own drift detection documentation illustrates a narrower version of the same scoping problem. Drift detection in CloudFormation works only on resources "that support drift detection" — the service does not claim universal coverage across every resource type CloudFormation can provision, and it is the engineering team's responsibility to know, resource type by resource type, which parts of a stack are actually being checked and which are silently excluded. A stack can report a clean IN_SYNC status while containing resource types the detection mechanism does not evaluate at all, and nothing in the console distinguishes "verified and matching" from "not evaluated" at a glance unless you read the per-resource detail.

The gap is granular: ignored attributes are sanctioned blind spots

The third limitation is the one platform teams build deliberately, for good reasons, and then frequently lose track of. Terraform's lifecycle block includes an ignore_changes argument specifically so that attributes legitimately managed outside Terraform — the desired instance count of an autoscaling group that an external autoscaler adjusts continuously, a tag applied by a separate tagging-compliance tool, an image digest updated by a deployment pipeline that runs more often than the infrastructure code changes — do not generate noisy, false-positive plan diffs every time someone runs terraform plan. This is a correct and necessary feature. Without it, Terraform would constantly propose reverting changes that other, equally legitimate systems are making on purpose.

The failure mode is scope creep in how that exclusion gets written. ignore_changes = [tags] on an autoscaling group is precise: it excludes exactly the attribute an external system legitimately owns. ignore_changes = all on the same resource — a pattern that shows up more often than platform teams like to admit, usually added under time pressure to silence a noisy plan diff without taking the time to identify which specific attribute is actually the source of the noise — excludes everything, including the security-relevant attributes nobody meant to exempt from review. From that point forward, a manual change to that resource's security group associations, its IAM instance profile, or its public IP assignment produces the same silence as the legitimate autoscaler-managed attribute the exclusion was originally written to accommodate. The plan is green not because nothing changed, but because the plan was explicitly told to stop looking.

None of these three gaps — temporal, scoped, or granular — requires a tooling failure or a bug. Terraform, Pulumi, and CloudFormation are all doing exactly what they are documented to do. The failure is organizational: treating a mechanism that checks a specific, event-triggered, partially-scoped comparison as if it were a continuous, comprehensive audit of what is actually running in production. It is neither continuous nor comprehensive by default, and knowing precisely where the coverage ends is the starting point for closing the gap deliberately rather than discovering it during an incident.

A taxonomy of what actually drifts, and why it happens

Drift is not one phenomenon with one cause. Distinguishing the sources matters because each one implies a different fix, and a program built to catch only one kind will have blind spots against the others.

Source of drift Typical trigger Why it happens Typical remediation approach
Incident-driven manual change An outage or security event where the fastest fix is a console edit, not a pull request Waiting for CI/CD approval during an active incident carries its own cost; the operational instinct to fix first and document later is often locally rational Post-incident reconciliation step, treated as a required part of incident closure, not an optional follow-up
Provider-side automatic change Cloud provider deprecates a default, rotates a managed certificate, auto-patches an underlying AMI, or changes a default setting on a managed service The provider owns part of the resource's lifecycle and changes it independent of your configuration Scheduled drift detection with provider-change awareness; explicit ignore_changes for genuinely provider-owned attributes
Cross-team tooling collision A security team's compliance scanner, a cost-optimization bot, or another team's automation modifies a resource Terraform also manages Multiple systems have write access to the same resource with no single source of truth agreed between them Resource ownership boundaries enforced through tagging, IAM permission boundaries, or splitting overlapping automation's write scope
Deliberate but undocumented exception An engineer or team makes a change they believe is justified — a temporary firewall opening, a broadened IAM policy for a debugging session — and never files the corresponding code change The workaround solves the immediate problem; updating the source of truth is a second, separately motivated step that competes with the next task Time-boxed exceptions with an expiry, tracked explicitly rather than left to memory
Console convenience during low-stakes changes An engineer finds it faster to toggle a setting in the console than to write, review, and merge a Terraform change for something that feels trivial The perceived cost of the IaC workflow exceeds the perceived risk of a small manual edit, in the moment Lowering the friction of small, well-scoped IaC changes so the console is never the faster path
Stale or partial imports A resource was migrated into Terraform management, but the import captured only some attributes, or a related sub-resource (an inline policy, a bucket policy document, a security group rule) was never imported at all Import is a manual, resource-by-resource operation with real gaps in provider coverage, not an automatic, guaranteed-complete operation Post-import verification comparing every attribute the provider exposes against what state actually captured

Two patterns in this table deserve emphasis because they are the ones most engineering leaders underweight. First, provider-side drift is not a hypothetical category invented for symmetry with the others — cloud providers routinely change attributes of resources you did not touch, through automatic minor-version upgrades on managed database engines, automatic certificate rotation, or changes to default security settings that ship as part of a platform update. This kind of drift is not caused by anyone on your team doing anything wrong, and no amount of process discipline aimed at engineers prevents it; only detection catches it. Second, deliberate-but-undocumented exceptions are the category most directly tied to security exposure, because they are, by definition, changes someone believed were justified at the time and never revisited — which is precisely the profile of the eleven-month open security group in the opening example, and the profile that shows up repeatedly in the case studies below.

A risk map: not all drift carries the same consequence

Treating every instance of drift as equally urgent is a fast way to either alarm-fatigue a team into ignoring drift reports entirely, or to burn remediation effort on low-consequence noise while a genuinely dangerous change sits unaddressed. A risk map organized around consequence, not just detection difficulty, is more useful for prioritization than a flat list of "things that changed."

Drift category Representative example Primary consequence Typical detection lag if unmanaged Reversibility
Network exposure Security group or firewall rule opened beyond its declared scope Direct security exposure — a reachable attack surface that did not exist in the threat model Weeks to months, often until a penetration test or audit Usually instantly reversible once found, but the exposure window has already occurred
Identity and access widening IAM policy, role trust relationship, or service account permission broadened beyond declared scope Privilege escalation path, lateral movement risk if the credential is later compromised Months, often surfacing only during a security review or compliance audit Reversible, but requires confirming nothing legitimate now depends on the broadened scope
Encryption or logging configuration Encryption-at-rest disabled during a migration, or access logging turned off to reduce noise or cost and never re-enabled Compliance violation, forensic blind spot if an incident later occurs in the unlogged window Months to indefinitely, since the absence of logs is not something anyone notices until logs are needed Reversible going forward; unlogged historical activity cannot be recovered retroactively
Capacity and topology drift Autoscaling limits, instance types, or subnet routing manually adjusted and never reconciled with declared values Reliability risk that surfaces specifically when declared capacity is relied upon — most commonly during a failover or disaster recovery attempt Often zero lag in normal operation, but full-severity lag until the exact moment declared capacity is tested under load Reversible in normal operation; the danger is that it is discovered under the worst possible conditions
Cost-relevant drift Orphaned resources, oversized instances left running after a manual resize, unused reserved capacity Budget waste, generally without direct security or reliability consequence Weeks to months, typically caught by a billing review rather than an infrastructure review Reversible with no operational risk once identified
Cosmetic or provider-managed drift Tag values changed by a compliance tool, minor version bumps applied automatically by the provider Little to no direct consequence; mainly a source of plan noise if not explicitly ignored Immediate, at the next plan run Not a genuine problem to reverse; the fix is usually to formally exempt the attribute

The ordering in this table is deliberate. Network exposure and identity widening sit at the top because they combine the two properties that make a risk genuinely dangerous rather than merely untidy: a long plausible detection lag, and a consequence that does not require any other failure to occur first — an open security group is a live risk the moment it exists, independent of whether anything else goes wrong. Capacity and topology drift sits in the middle for a different reason: in isolation it is often harmless, sometimes even a reasonable operational adjustment, but its risk is conditional and concentrated at exactly the moment an organization can least afford a surprise — a regional failover, a disaster recovery drill, or a sudden traffic spike. Cost drift and cosmetic drift sit at the bottom not because they are unimportant in aggregate, but because neither carries the acute, single-incident consequence that justifies the same urgency as the categories above them.

Case study one: the open port discovered by a penetration test, not a pipeline

This is a hypothetical, illustrative scenario constructed for this article. It does not describe a real QAtronic client or a documented public incident.

Initial situation. A payments infrastructure team at a fintech company manages its VPC, security groups, and RDS instances through Terraform, with a mature CI/CD pipeline that runs a plan on every pull request and blocks merges on unreviewed changes. The team considers its infrastructure well-governed, and by the measure of pipeline health — merge frequency, review coverage, time-to-deploy — it is.

The hidden assumption. The team assumes that because their Terraform pipeline is clean, their security group configuration accurately reflects what is declared in code. This assumption has never been tested against reality directly; it has only been tested against the pipeline's own output, which is a different thing.

The technical and organizational cause. During a database performance incident eight months earlier, an on-call engineer temporarily widened the RDS instance's security group ingress rule from the application subnet's CIDR range to 0.0.0.0/0, intending to rule out a network path issue as the cause of elevated query latency. The network path was not the cause. The engineer reverted every other diagnostic change made during the incident, but the security group rule, added directly through the console rather than through the Terraform module that normally manages that resource, was not tracked in any checklist and was missed during cleanup. Because no pull request touched that specific security group resource in the following eight months — the module was stable, and nobody had reason to modify it — no plan ever refreshed its state and compared it against the code's declared ingress rule.

The consequence. An external penetration test commissioned ahead of a Series C due-diligence process discovered the open ingress rule directly, through an unauthenticated network scan, well before any internal process did. The database itself was not compromised — authentication still required valid credentials — but the finding became a material item in the due-diligence report, requiring the company to explain to a private equity investor's technical diligence team why a database explicitly declared, in code, as accessible only from the application subnet was in fact reachable from anywhere on the internet, and why internal controls had not caught it in eight months.

The decision that needed to be made. The immediate fix — reverting the security group rule to match the Terraform declaration — took minutes. The harder decision was structural: whether to treat this as an isolated engineer error, or as evidence that the organization had no mechanism for detecting the difference between "the pipeline is green" and "production matches the declaration" for any resource the pipeline was not actively asked to check.

The better approach. The team implemented a scheduled, standalone drift detection job — independent of any pull request — running against every managed resource on a recurring basis rather than only when code changed, with network-exposure-relevant resources (security groups, network ACLs, public IP assignments) flagged for same-day review rather than folded into a general weekly report. They also instituted a rule specific to incident response: any manual infrastructure change made during an active incident is logged in the incident ticket by resource ID, and incident closure is blocked until each logged manual change is either reverted or formally reflected back into the IaC source within one business day — treating the reconciliation step as part of resolving the incident, not as optional cleanup that competes with the next sprint's priorities.

The detection landscape: what each major tool actually gives you

Understanding what your specific toolchain checks, and what it does not, is a prerequisite for deciding where to invest additional detection effort. The following comparison focuses on mechanism, not marketing, and distinguishes native platform capability from third-party tooling built on top of it.

Tool / mechanism What it compares Trigger model Remediation Notable limitation
Terraform / OpenTofu plan (default) State file against live provider API, for resources already in state On-demand, typically wired to code push in CI/CD None automatic; a human applies the plan No coverage for unmanaged resources; attributes under ignore_changes are excluded entirely
Terraform -refresh-only plan Same comparison, but explicitly separated from a change proposal so drift can be reviewed without risking an unintended apply On-demand or scheduled, if a team builds the schedule itself None automatic; produces a state-update-only plan for review Still scoped to tracked resources only; does not run on its own without a trigger someone builds
HCP Terraform Drift Detection State against live infrastructure, across all workspaces with the feature enabled Scheduled, continuous, independent of code push (the specific gap this feature was built to close) Notifies via email, Slack, or webhook; remediation still requires a human-reviewed plan Commercial feature of HCP Terraform Business tier and above; still limited to Terraform-managed resources
Pulumi refresh --preview-only / --expect-no-changes Recorded state against live provider state On-demand or scheduled via Pulumi Deployments Two paths: remediate (reapply code to revert drift) or adopt (update code to match reality) Same tracked-resource scoping limitation as Terraform; adoption path requires a human decision each time
AWS CloudFormation drift detection Template-declared configuration against live resource properties, per stack On-demand via console, CLI, or API; can be scheduled through a custom Lambda/EventBridge pipeline, but this is not built in None automatic; a human must decide to update the stack or the resource Only covers resource types that support drift detection; single concurrent detection operation per stack; no built-in schedule
AWS Config Continuous recording of resource configuration changes against defined Config Rules, independent of any IaC tool Continuous, event-driven, does not require a code push of any kind Can trigger automated remediation actions through Config Rules with associated remediation configurations Evaluates configuration compliance against rules you define; does not natively know what your Terraform or CloudFormation source of truth declares unless you build that mapping yourself
Third-party drift scanners (commercial: Spacelift, env0, Firefly, and similar; formerly open-source driftctl) Live cloud inventory against IaC state across multiple tools and providers in one view Typically scheduled and continuous, as the core product feature Varies by product; several offer guided or automated remediation workflows Requires broad read (and sometimes write) access across cloud accounts; adds a third-party dependency into the security perimeter

Two rows in this table are worth a closer look because they represent genuinely different architectural approaches, not just different vendors selling the same thing.

AWS Config does not compare against your Terraform code at all. It continuously records configuration changes to supported resource types and evaluates them against rules you define — "does this security group have any rule allowing unrestricted ingress," "is this S3 bucket encrypted," "is this IAM policy free of wildcard actions" — independent of whether the resource in question is managed by Terraform, CloudFormation, or a person clicking through the console. This makes it a genuinely different and complementary layer: it will catch the open security group in the case study above through its own continuous evaluation, regardless of whether Terraform ever touches that resource again, because it is not waiting for a code push at all. The tradeoff is that Config Rules tell you a resource violates a policy; they do not, on their own, tell you that a resource no longer matches what your Terraform module declares, unless you build that specific correlation yourself.

driftctl, an open-source tool that once served as a popular independent option for scanning cloud accounts against Terraform state across providers, is now formally in what its own maintainers describe as "maintenance mode," with the project's README stating plainly that the team "cannot promise to review contributions" going forward. This matters less as commentary on any one tool and more as a reminder that the drift detection tooling landscape has consolidated toward paid platform features (HCP Terraform, Pulumi Deployments) and commercial third-party products, rather than toward mature, actively maintained open-source alternatives — a genuine constraint worth factoring into a build-versus-buy decision, not a reason to avoid the category.

Chart 1: illustrative drift accumulation without a reconciliation cadence

Weeks since last full reconciliation Illustrative percentage of managed resources drifted
0 0%
4 3%
8 7%
12 11%
20 18%
30 24%
40 29%
52 33%

What this illustrates: the curve is drawn to rise quickly at first and then flatten, on the reasoning that most drift accumulates through everyday incident response and console convenience — a roughly steady rate of new manual changes per week — while the rate of new drift slows slightly over time as teams that already drifted on a given resource have less remaining untouched configuration left to drift on. The specific numbers are illustrative, not measured; the mechanism behind the shape — that unmanaged intervals accumulate discrepancies roughly continuously, and nothing removes a discrepancy from the pool until someone deliberately reconciles it — is the same mechanism documented directly in HashiCorp's own description of why HCP Terraform built a scheduled, continuous drift detection feature rather than relying on plan-on-push alone.

Case study two: the IAM policy nobody meant to keep

This is a hypothetical, illustrative scenario constructed for this article. It does not describe a real QAtronic client or a documented public incident.

Initial situation. A B2B SaaS company runs a suite of internal microservices on AWS, with each service's IAM execution role defined in Terraform and scoped, in principle, to the specific S3 buckets and DynamoDB tables that service needs. The platform team enforces least-privilege IAM as a stated architectural principle and reviews new role definitions during code review.

The hidden assumption. The team assumes that because every IAM role originates in a reviewed Terraform pull request, the roles running in production still match what was reviewed. Nothing in their process distinguishes "this policy was reviewed when it was written" from "this policy, as currently attached in AWS, still matches what was reviewed."

The technical and organizational cause. A background job responsible for exporting customer data began failing intermittently with an access-denied error against a newly added S3 bucket that had not existed when the job's IAM role was originally scoped. Under pressure to restore the export feature — a customer-visible feature with contractual SLA implications — an engineer attached an AWS-managed AmazonS3FullAccess policy directly to the role through the console as an immediate fix, intending it as temporary, with a plan to replace it with a correctly scoped inline policy once the fix shipped through the normal Terraform workflow. The export job started working. The temporary managed policy attachment, because it was added outside Terraform, was not represented anywhere in the module that defined the role, and because no further pull request touched that specific role for the following five months, no plan run ever surfaced the discrepancy.

The consequence. During preparation for a customer-requested security questionnaire, the security team ran an IAM policy analysis across the AWS account using the account's own IAM Access Analyzer output rather than relying on the Terraform codebase as the source of truth, specifically because a prior finding had already taught them not to. The analysis surfaced the export job's role as having broad, unscoped S3 access across every bucket in the account — including buckets belonging to entirely unrelated services and, notably, a bucket storing infrastructure backups with more sensitive access implications than the original export bucket the fix was meant to address. The Terraform code, read on its own, showed a correctly scoped policy; it had simply stopped being the operative policy the moment the console attachment was added.

The decision that needed to be made. Whether to trust the Terraform repository as ground truth for IAM configuration going forward, or to accept that any policy attached outside a Terraform apply — a mechanically trivial thing to do in AWS, requiring no special privilege beyond IAM write access most senior engineers already hold — could silently override what the code declared, indefinitely, with no signal anywhere in the normal development workflow.

The better approach. The team adopted two changes rather than one. First, they scoped the affected role correctly and removed the managed policy attachment, restoring the codebase and live account to agreement. Second, and more durably, they configured a scheduled AWS Config rule specifically targeting IAM role policy attachments, comparing the live attached-policy set for every Terraform-managed role against an expected baseline recorded at apply time, and alerting when the two diverged — a check that runs independent of any pull request and does not depend on anyone remembering to touch that specific Terraform resource again. They also formalized a rule the export-job incident had violated informally: any IAM policy change made outside the normal Terraform workflow, for any reason including incident response, must be time-boxed with an explicit expiry recorded at the moment it is made, and a role with an unresolved temporary attachment past its expiry is automatically flagged for review rather than left to persist on the assumption that someone remembers.

Why detection is necessary but not sufficient

A detection mechanism that fires correctly and lands in a channel nobody is accountable for reading is, in practical effect, indistinguishable from having no detection mechanism at all. Most organizations that struggle with drift are not struggling because no tool told them about it; they are struggling because the tool told someone, and that someone did not have a clear mandate, the authority, or the time to act on it before the next thing demanded their attention.

This is the more durable failure mode, and it is organizational rather than technical. A Slack alert from a scheduled drift job routes, by default, to whichever channel the team that configured the integration happened to be watching at setup time — often the platform team's own operations channel, which is also where every deploy notification, every CI failure, and every dependency-update bot posts. A security-relevant drift alert competing for attention in that channel against dozens of routine notifications per day has a predictable fate: it gets acknowledged with an emoji reaction, or it does not get acknowledged at all, and either way it does not reliably result in a decision about whether the drift is acceptable, needs immediate reversion, or needs to be adopted into the codebase as the new correct state.

Responsibility Platform / infrastructure team Security engineering On-call / SRE Service-owning engineering team Engineering leadership
Configure and maintain drift detection tooling Owns Consults — — Approves budget/tooling choice
Triage a new drift finding for severity Consults Owns for security-relevant categories (network, identity, encryption) Consults if operational Owns for service-specific attributes —
Decide: revert to declared state, or adopt drift into code Consults Approves for security-relevant categories — Owns, with sign-off from security when relevant Escalation path if teams disagree
Execute the reversion or adoption change Platform team executes for shared infrastructure — — Service team executes for service-owned resources —
Set and enforce time-boxed exceptions for legitimate manual changes Owns the mechanism Sets policy for what qualifies as time-boxable Requests exceptions during incidents Requests exceptions for service-specific needs Reviews aggregate exception volume quarterly
Own the consequence if undetected drift causes an incident Shared accountability with the service-owning team Shared accountability for security-relevant categories — Shared accountability with platform team Ultimate accountability

The table above is deliberately built around a specific failure this article has already illustrated twice: drift ownership diffused across "everyone" in principle tends, in practice, to belong to no one. Assigning the platform team ownership of the detection mechanism itself, while assigning triage-and-decision authority to whichever team actually owns the drifted resource, avoids two opposite failure patterns: a platform team drowning in triage decisions for resources it does not have the context to evaluate, and a service team with no visibility into drift because the detection tooling belongs entirely to a different group that never routes findings back to them.

Case study three: the disaster recovery test that revealed the declaration was fiction

This is a hypothetical, illustrative scenario constructed for this article. It does not describe a real QAtronic client or a documented public incident.

Initial situation. An e-commerce platform runs its production workload in a single AWS region, with a documented disaster recovery plan calling for a full infrastructure rebuild in a secondary region using the same Terraform codebase, targeted at a four-hour recovery time objective ahead of the company's peak seasonal sales period. The DR plan had been reviewed and approved by leadership the previous year and was treated, reasonably, as a solved problem — the Terraform code that built production could, by design, build an equivalent environment anywhere.

The hidden assumption. The team assumed that because the Terraform code, applied fresh, would produce infrastructure matching its own declarations, applying that same code in a new region would produce infrastructure matching production — treating "matches the code" and "matches what's actually running" as the same claim, when eighteen months of incremental manual adjustments to the live production environment had made them meaningfully different.

The technical and organizational cause. Over the eighteen months since the DR plan was last tested, production had accumulated several categories of drift that the case studies above have already described individually: an autoscaling group's minimum and desired capacity had been manually raised during two separate high-traffic events and never reverted in code, because reverting felt like inviting the same emergency scaling scramble next time; a load balancer's target group had been repointed to a replacement service during a migration, with the change made directly against the load balancer rather than through the Terraform module managing it, because the migration's own tooling handled traffic cutover directly; and a database parameter group's connection limit had been raised by a vendor support engineer during a troubleshooting session for an unrelated issue, through the console, with no corresponding Terraform change ever filed. None of these changes were reckless in isolation. Each one solved a real, immediate problem, and each one left the Terraform codebase describing an environment that no longer matched the one actually serving customer traffic.

The consequence. When the DR team executed a planned test — applying the same Terraform code to build a full environment in the secondary region, ahead of the seasonal peak specifically to validate readiness — the new environment came up successfully by every measure Terraform itself reported: the apply completed without error, every resource in the plan was created, and the pipeline reported success. It was also meaningfully undersized relative to what production actually needed, because the Terraform code's declared autoscaling minimums, connection limits, and capacity settings reflected the state from eighteen months earlier, not the manually adjusted values production had been running on since. The test environment, had it been called on to actually take over production traffic, would have hit connection limits and autoscaling ceilings within the first hour of realistic load — a failure the DR test caught only because someone thought to load-test the rebuilt environment against realistic traffic rather than treating a successful terraform apply as proof the environment was production-equivalent.

The decision that needed to be made. Whether "the Terraform code applies successfully" was an adequate definition of disaster recovery readiness, or whether readiness required an explicit, separate step verifying that the code's declarations still matched what production actually needed — a distinction the team had not previously drawn, because in the absence of drift the two claims are identical, and the team had no reason to believe drift had accumulated until the test specifically surfaced it.

The better approach. The team treated the DR test failure as a drift-detection failure first and a capacity-planning failure second, and restructured their DR validation process accordingly: before any future DR test, a full drift reconciliation pass runs against production, comparing every attribute the DR-critical modules declare against live values, with any discrepancy resolved — either reverted in production or adopted into the code — before the DR rebuild is attempted. They also added a standing rule that any manual production change made to a resource within a DR-critical module is flagged for reconciliation within one week rather than left for the next scheduled full audit, specifically because the connection-limit change made by a vendor support engineer during an unrelated troubleshooting session illustrated that drift does not only originate from the team's own engineers, and a detection cadence built around "someone on our team remembered to file a ticket" would have missed it regardless.

A maturity model for drift detection and remediation

Organizations rarely move from zero drift visibility to a fully governed reconciliation program in one step, and treating this as a binary — "we have drift detection" or "we don't" — obscures the more useful question of what capability actually needs to be built next. The five levels below are organized around what an organization can actually answer with confidence at each stage, not around which specific tool it has purchased.

Level Name What the organization can answer Typical detection mechanism Typical failure mode at this level
0 Unmonitored Nothing. Drift is discovered only through an incident, an audit, or a penetration test None; "the pipeline is green" is treated, incorrectly, as sufficient evidence Long, invisible exposure windows; the eleven-month open security group in this article's opening scenario is a Level 0 organization by definition
1 Event-triggered Whether declared configuration matched live configuration the last time a related code change was pushed Default terraform plan / pulumi preview / CloudFormation change sets, run only on code push Resources that go untouched for long periods accumulate undetected drift regardless of pipeline discipline
2 Scheduled Whether declared configuration currently matches live configuration for every managed resource, checked on a recurring cadence independent of code changes Scheduled -refresh-only plans, HCP Terraform Drift Detection, Pulumi Deployments scheduled refresh, or a custom scheduled job Findings exist but often lack a clear owner or a triage process; alerts get acknowledged and not acted on
3 Owned and triaged Who is accountable for deciding what happens to each finding, and within what time frame Scheduled detection plus an assigned RACI-style ownership process and a defined triage severity scale Reconciliation happens, but the organization has no way to know whether new drift is accumulating faster than it is being resolved
4 Governed with enforced boundaries Whether the rate of new drift is increasing, decreasing, or stable, and whether unmanaged (shadow) resources exist in the account at all Continuous, policy-as-code-gated detection covering both managed-resource drift and unmanaged-resource discovery, with time-boxed exceptions enforced automatically Requires sustained investment; the main risk is treating Level 4 as a project with an end date rather than an ongoing discipline

Most organizations that have never deliberately invested in this problem sit at Level 1, often while genuinely believing, on the strength of a consistently green pipeline, that they are closer to Level 3 or 4. The distance between believing your organization is well-governed on this dimension and actually being well-governed on it is, specifically, the distance this article has spent its middle sections describing: the gap between what a plan checks and what a plan is silently assumed to guarantee.

Chart 2: illustrative detection lag by maturity level

Maturity level Illustrative median detection lag (days)
0 — Unmonitored 240
1 — Event-triggered 95
2 — Scheduled 12
3 — Owned and triaged 6
4 — Governed with enforced boundaries 1

What this illustrates: the largest single improvement in this illustrative model comes from moving off Level 0 and Level 1 — that is, from having no detection or detection that depends entirely on unrelated code changes — to any form of scheduled, independent checking at Level 2. The gains from Level 2 to Level 4 are real but smaller in relative terms, because the harder remaining problem at that stage is not detection speed but triage speed and organizational authority to act, which is exactly why Level 3 in the maturity model above is defined by ownership rather than by tooling.

A practical drift triage framework

Detection produces findings. Findings require a repeatable way to decide what happens next, or they pile up as an ungoverned backlog that eventually gets ignored wholesale — the same fate that met the Slack channel described earlier in this article. The following framework is built specifically for this article, organized as a sequence of questions any team can walk through for a given finding, rather than a generic incident-response checklist repurposed for a different problem.

Step 1 — Classify the resource against the risk map. Does the drifted attribute fall into network exposure, identity and access, encryption/logging, capacity/topology, cost, or cosmetic, using the categories defined earlier in this article? This single classification step determines both urgency and who should be in the room for the next step.

Step 2 — Determine whether the live state or the declared state is actually correct. This is the question teams most often skip, defaulting to "revert to what's declared" as though the code is automatically right. Sometimes it is not: the manual change may have been a legitimate fix for a real problem the code never anticipated, in which case the correct action is updating the code to adopt the live state, not reverting production to a configuration that was already known to be wrong.

Step 3 — Check for an active, time-boxed exception. Was this change logged as a deliberate, temporary exception during an incident or a debugging session, with an expiry date, per the process described in the case studies above? If so, and the expiry has not passed, no action is required yet beyond confirming the exception is still tracked. If the expiry has passed, treat it as unresolved drift requiring an immediate decision, not a new finding starting from zero.

Step 4 — Assign a decision owner and a time frame based on the Step 1 classification. Network exposure and identity findings route to security engineering with same-day review. Capacity, topology, and cost findings route to the service-owning team with a standard weekly cadence. Cosmetic and provider-managed findings are either formally exempted via a precise ignore_changes entry or closed with no further action.

Step 5 — Execute the decision: revert, adopt, or exempt. Reverting means applying the declared code to bring live infrastructure back into agreement. Adopting means updating the code — through a normal, reviewed pull request, using terraform import or Pulumi's equivalent adoption workflow where the resource was previously unmanaged — to reflect the live state as the new correct declaration. Exempting means adding a precise, narrowly scoped ignore_changes entry (never a blanket all) with a comment explaining which external system legitimately owns the attribute.

Step 6 — Record the decision, not just the fix. A drift finding resolved without a record of why it was resolved that way is a finding the organization will rediscover from scratch the next time a similar case comes up, with no institutional memory of the previous reasoning. A short, dated note attached to the relevant Terraform module — even a code comment — closes this loop cheaply.

Step 7 — Feed recurring patterns back into prevention. If the same resource, the same team, or the same category of drift shows up in triage repeatedly, the fix is not another round of triage; it is addressing why the manual change keeps happening — often a genuine gap in how fast the IaC workflow can respond to a legitimate, time-sensitive need, which is a process problem worth fixing on its own terms rather than a compliance problem to keep re-litigating.

Testing the detection mechanism itself, not just the infrastructure

Every recommendation so far in this article assumes the drift detection job actually works — that it runs on schedule, that it actually queries live infrastructure rather than a cached copy, and that a genuine discrepancy actually produces an alert someone sees. That assumption deserves the same skepticism this article has applied to the plan-is-green assumption, because a broken or silently degraded detection job is functionally identical to having none, with the added danger that a team who believes they have Level 2 or Level 3 coverage will stop looking for other ways to catch the same problem.

A drift detection job can fail quietly in several specific, unglamorous ways. A scheduled job can lose its IAM permissions after an unrelated cleanup of over-broad roles, and continue "running" while every API call it makes returns an authorization error the job's own error handling swallows rather than escalates. A job can be pointed at the wrong workspace or the wrong state backend after a migration, and report a clean result because it is comparing the right code against the wrong, unrelated infrastructure. A webhook integration into a Slack channel can silently fail after a token rotation, so the job itself succeeds and produces findings that never reach anyone. None of these failures look like a failure from the outside; each one produces exactly the same visible signal as a healthy system with no drift to report — silence — which is the same signal a genuinely drift-free environment produces, and the two are indistinguishable without a deliberate check.

The practical fix, and the one most consistent with how QAtronic approaches verification generally, is to test the detection mechanism the same way you would test any other production safeguard: by deliberately triggering the condition it exists to catch, on a controlled resource, and confirming the expected signal actually arrives within the expected time.

A concrete version of this exercise: in a non-production account or a deliberately isolated test resource, make a known, intentional manual change — widen a security group rule by one CIDR block, or attach an additional managed policy to a role that exists specifically for this purpose — and record the exact time the change was made. Then measure, without prompting anyone, how long it takes for the scheduled detection job to flag it, whether the alert reaches the channel and the person the RACI table above assigns it to, and whether the content of the alert is specific enough for that person to act without first investigating what actually changed. If the alert never arrives, arrives late enough to exceed the time frame assigned in the triage framework's Step 4, or arrives with too little detail to act on, the finding is not "the security group was open" — it is that the detection program itself has a gap, discovered under controlled conditions rather than during a real incident.

This exercise is worth running on a recurring basis, not once, for the same reason any other safeguard needs periodic revalidation: the permissions, integrations, and configurations a detection job depends on change over time, often for reasons entirely unrelated to drift detection itself, and a program that passed this test a year ago carries no guarantee it still passes today. A reasonable cadence is quarterly for organizations at maturity Level 2 or above, run as a scheduled exercise with the same seriousness as a disaster recovery test — because, as the third case study in this article demonstrated, an untested recovery plan and an untested detection program fail for structurally identical reasons: both assume a mechanism works because nobody has watched it fail.

Where this investment does not make sense to push further

None of this justifies unlimited investment in drift tooling, and a few boundaries are worth stating directly rather than leaving implicit.

For a small team running a handful of services with a genuinely small infrastructure footprint, Level 4 governance is disproportionate. A scheduled -refresh-only plan run weekly, reviewed by whichever engineer is on platform duty that week, closes most of the meaningful gap at a fraction of the operational overhead a dedicated drift-governance program requires, and the RACI table above can collapse to a single line — "the founding engineering team reviews it" — without losing its function.

For resources genuinely and permanently owned by another automated system — an autoscaler, a service mesh's dynamically managed routing rules, a managed database's provider-controlled maintenance window — the correct response is a precise, well-documented ignore_changes entry, not a recurring triage cycle that will flag the same non-issue every week. Treating every detected difference as a finding requiring a decision, when the difference is a known and intended division of ownership, produces exactly the alert fatigue that causes teams to stop reading drift reports at all.

And for organizations already running a mature GitOps model on Kubernetes — where a controller like Argo CD or Flux continuously reconciles cluster state against a Git repository, correcting drift automatically and by design as a core part of how the system operates rather than as an add-on check — much of this article's argument does not apply in the same form, because continuous reconciliation is the default behavior, not a gap that needs to be closed. The distinction is worth naming precisely: Terraform, Pulumi, and CloudFormation are fundamentally imperative, plan-and-apply tools that act when invoked, while a GitOps controller is a continuously running reconciliation loop by architecture. Extending genuinely continuous reconciliation to cloud infrastructure outside Kubernetes is possible — it is what HCP Terraform's Drift Detection and Pulumi's scheduled refresh are approximating — but it requires deliberately building or buying that continuous layer, because the underlying tools were not designed to provide it on their own.

What executives should ask their infrastructure teams

A useful diagnostic for a leadership team that has not previously asked these questions directly: how many of the following can your platform or infrastructure lead answer with confidence, right now, without needing a week to investigate?

  • When was the last time every managed resource's live configuration was compared against its declared configuration, independent of whether any code was recently changed?
  • Which resources in production were created outside Terraform, Pulumi, or CloudFormation entirely, and how do we know the list is complete?
  • If a security group, IAM policy, or encryption setting drifted from its declaration today, who would find out, how quickly, and who has the authority to decide whether to revert it or adopt it?
  • Do any of our ignore_changes or equivalent exclusions apply to an entire resource rather than to specific, named attributes?
  • If we had to rebuild production from our current infrastructure-as-code in a new region tomorrow, would the result match what is actually running today, or what was running whenever these modules were last touched?

An organization that can answer all five with specifics is meaningfully further along than a green pipeline history alone would suggest. An organization that cannot is not necessarily poorly run — most organizations cannot, including many with genuinely disciplined engineering cultures — but it is carrying an unmeasured risk that the rest of this article has tried to make specific and addressable rather than abstract.

How QAtronic approaches this

Drift detection is fundamentally a verification problem, and verification is the discipline QAtronic works with engineering teams to build into existing delivery pipelines rather than bolt on as a separate audit. That means helping teams design the scheduled, independent reconciliation checks this article describes — distinct from event-triggered plan runs — mapping detected drift against a risk classification specific to the organization's own infrastructure rather than a generic severity scale, and testing disaster recovery and failover procedures against what production actually runs today, not against what the infrastructure code declared the last time someone happened to touch it. If your organization has a green pipeline and has never separately verified that green means what it appears to mean, that verification is a bounded, well-scoped place to start.

Conclusion: the distinction worth taking back to your team

A passing CI/CD check on an infrastructure-as-code pull request answers one specific question: does this proposed code change, applied to the resources it already tracks, produce the intended result. It does not answer, and was never built to answer, the broader question every executive assumes it answers by default: does production currently match what we've declared it should be. Those are different claims, and the gap between them is not a bug in Terraform, Pulumi, or CloudFormation — every limitation described in this article follows directly from how these tools are documented to work. The gap exists because organizations adopted a code-review discipline for infrastructure changes without adopting an equivalent, independent discipline for verifying that the declarations those reviews approved still describe reality.

The distinction worth bringing back to your own engineering organization is this: a state file is a cache of what was true the last time someone checked, not a continuously verified record of what is true now, and the interval between those two things is exactly as long as the interval since your last independent reconciliation — not since your last deploy. Ask your platform team when that last independent reconciliation actually happened, for every resource that matters, not just the ones a recent pull request happened to touch. If the honest answer is "we're not sure," that uncertainty is the finding, and it is a more useful place to start than any dashboard currently telling you everything is green.

Frequently Asked Questions

Does running terraform plan more often solve this problem?

Partially, and only for the temporal gap. Running plans on a schedule rather than only on code push closes the window where drift can sit undetected simply because nobody touched the relevant module. It does not address the scoping gap — a plan still cannot see resources that were never imported into state — or the granularity gap created by broad ignore_changes entries. A useful mental model: scheduling closes the "when" gap, but the "what" gap requires separate attention to import completeness and exclusion scope.

Is AWS Config a replacement for Terraform drift detection?

No, but it is a genuinely complementary layer, for a specific reason: AWS Config evaluates live resource configuration against rules you define, continuously and independent of any IaC tool, which means it catches policy violations even on resources Terraform never touches. It does not, on its own, know what your Terraform code declares a given resource should look like, so it cannot tell you "this no longer matches your IaC" without additional integration work mapping Config's findings back to your Terraform source. Organizations with mature drift governance tend to run both: IaC-native drift checks for declaration fidelity, and continuous policy evaluation for security-relevant configuration regardless of origin.

Should every detected drift be reverted automatically?

No. Automatic reversion is appropriate for a narrow set of high-confidence, low-ambiguity cases — a security group rule that unambiguously violates a network policy, for instance — but applying it broadly risks reverting a legitimate emergency fix before anyone has verified the underlying problem it addressed is actually resolved, potentially reintroducing an outage. The triage framework in this article treats reversion as one of three possible outcomes, alongside adoption and formal exemption, decided deliberately rather than applied uniformly.

How is this different from a configuration management tool like Ansible or Chef continuously enforcing state?

Configuration management tools built around continuous enforcement — repeatedly applying a desired state and correcting deviations as a core operating loop — solve a version of this problem by architecture, similar to how GitOps controllers solve it for Kubernetes. Most cloud infrastructure provisioning tools (Terraform, Pulumi, CloudFormation) are not built this way by default; they are invoked on demand rather than running as a continuous loop, which is precisely why scheduled drift detection has to be added deliberately rather than assumed.

Our compliance framework requires periodic configuration reviews. Doesn't that already cover this?

It covers part of it, on whatever cadence your compliance framework specifies — often quarterly or annually. The gap this article describes is what happens between those reviews, and the case studies above are specifically examples of drift that existed for months before a compliance-driven review, penetration test, or security questionnaire happened to surface it. A compliance review is a valid detection mechanism; it is a slow one relative to the exposure window a security-relevant drift category can create.

What is the single highest-leverage first step if we have never done any of this deliberately?

Run a full reconciliation pass across every managed resource once, independent of any pending code change, and separately inventory which resources in your cloud accounts exist outside your IaC state entirely. The first pass will surface a specific, non-hypothetical list for your own environment rather than an abstract concern, and in most organizations doing this for the first time, that list is what actually secures budget and attention for building the scheduled process described in the maturity model above.

Recent posts

October 2, 2026
FHIR Interoperability Testing: Certified, Not Connected
October 2, 2026
Definition of Done Erosion: Why Standards Quietly Slip
October 2, 2026
Kubernetes Admission Control Testing: A Field Guide