Skip to content

Day 9 of 100

The Terraform Drift

TerraformIaCCloudIntermediate15–20 min

🔥 Problem Statement

A routine `terraform plan` on the production workspace comes back with unexpected drift — a resource's configuration doesn't match what Terraform expects, because someone modified it directly in the cloud console. Digging into it, the change turns out to be small: a security group rule adjusted to unblock an urgent request last week, made outside any PR, and never reflected back into the Terraform configuration.

In the team channel, the read is: "It's only a small change. Let's ignore it and move on."

🏗️ Environment

  • Terraform: remote backend, production workspace
  • Cloud: production account, resources managed by Terraform
  • Change process: PR review for code changes; console access still available to some engineers

🤔 Your Challenge

Should this manual change be kept, reverted, imported into Terraform, or explicitly reflected in the configuration — and does 'ignore it' actually make the drift go away?

  • If the plan is ignored this time, what happens the next time someone runs terraform apply on this resource?
  • What's the difference between the manual change being wrong and the manual change simply not being represented in code?
  • Why does it matter whether infrastructure state lives in the console, in Terraform state, or in both, disagreeing?
  • What would prevent this same situation from happening again next week, for a different resource?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

"Ignore it" isn't actually an available option, even though it feels like the low-effort path. Terraform's plan will keep flagging this drift on every future run until either the configuration is updated to match reality, or the manual change is reverted to match the configuration — there's no third state where Terraform quietly stops noticing. Choosing to do nothing just defers the decision to whoever runs the next apply and gets surprised by a plan they didn't expect.

Step 2 — Identify the likely bottleneck

The actual risk isn't the security group rule itself — the scenario says it's small, and that might well be true. The risk is what "small, made outside review, not reflected in code" represents as a pattern: a manual change to production infrastructure with no PR, no review, and no record beyond whatever's in the cloud provider's own audit log. If that becomes normal practice under time pressure, Terraform state and actual infrastructure diverge a little more each time, and eventually a terraform apply reverts someone's forgotten manual fix without anyone realizing it was going to until production is already affected.

Step 3 — Investigation

  • Confirm exactly what changed — pull the specific diff between the Terraform-expected state and the actual resource configuration from the plan output.
  • Find out why the change was made — was it a legitimate, still-needed fix, or a temporary workaround that should have been reverted already?
  • Check whether it's still needed. If the urgent request it was solving is resolved, reverting to match the Terraform configuration is often the simplest correct answer.
  • If it's still needed, decide how it should be represented in code — either update the Terraform configuration to match the new desired state (making the manual change permanent and reviewed after the fact), or use terraform import/state manipulation carefully if the resource itself needs to stay exactly as it is.
  • Check for a pattern — is this the only manually-changed resource, or would a broader audit of the production workspace turn up more undocumented drift?

Pick one of the two real options — update the code to match the infrastructure, or update the infrastructure to match the code — rather than leaving the plan permanently dirty. If the change was legitimate and should stay, it belongs in a reviewed PR updating the Terraform configuration, even after the fact; that turns an undocumented manual change into a documented, reviewed one, which is most of the value of catching drift at all.

Beyond this one resource, this is worth a conversation about why direct console access to production was available in the moment someone needed an urgent fix. If manual changes are sometimes genuinely necessary — a real incident, a genuine emergency — the fix isn't punishing that decision after the fact, it's having a fast, reviewed path to update Terraform immediately afterward, so drift gets caught in minutes rather than discovered by accident during an unrelated plan a week later.

Step 5 — Engineering lesson

Choosing to manage infrastructure with Terraform is a claim that the code is the source of truth — and that claim only holds if every change, including the fast, well-intentioned emergency ones, eventually makes it back into the code. A small unreflected change ignored once sets the precedent that drift is acceptable, and drift acceptable once tends to compound quietly until a routine apply does something nobody expected.

🧠 Today’s Takeaway

If infrastructure is managed by code, the code must remain the source of truth.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.