Skip to content
Platform Engineering

Day 9: Manage Everything Like Code. Not Like Chaos.

Terraform gets infrastructure into Git. That's one system out of six that can still drift silently. Everything-as-code is what happens when the same discipline applies everywhere else too.

Thamunkpillai 5 min read

Five days ago this series made the case for treating infrastructure as code. Most teams that adopt that idea stop at exactly one system: servers and networks go into Terraform, and everything else — Kubernetes manifests applied by hand, IAM policies clicked together in a console, monitoring dashboards built once and never touched again — stays exactly as undeclared and driftable as it always was. The insight behind IaC wasn't "Terraform is good." It was "if it exists, it can drift; if it's code, it can be versioned, reviewed, and automated." That insight doesn't stop applying just because you've moved past the network layer.

Everything-as-code is that same argument, generalized: any system whose state matters should have its state declared in a file, in version control, changed only through review — not through whoever has console access and fifteen minutes.

The inventory, once you actually take it

SystemOld wayAs code
InfrastructureManual provisioningTerraform, Pulumi, CloudFormation
Kubernetes workloadskubectl edit in prodManifests via Helm/Kustomize, applied by a pipeline
Access & identityIAM changes in the consoleAccess-as-code (Terraform, AWS PAM modules)
Policy & securityA wiki page nobody readsOPA/Conftest rules, enforced in CI
ObservabilityDashboards built by hand in the UIDashboards-as-code (Grafana JSON, alerts as YAML)
ConfigurationSSH in and edit a fileAnsible/Chef/Puppet, applied from a repo
YAML
# Access as code — an IAM policy that's reviewed like any other change,
# not clicked together once and forgotten
resource "aws_iam_policy" "read_only_billing" {
  name   = "read-only-billing"
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect   = "Allow"
      Action   = ["ce:Get*", "ce:Describe*"]
      Resource = "*"
    }]
  })
}
rego
# Policy as code — a rule that runs in CI, not a Slack message asking
# people nicely not to do this
package terraform.security
 
deny[msg] {
  input.resource.aws_s3_bucket[name].acl == "public-read"
  msg := sprintf("bucket %v must not be publicly readable", [name])
}

Neither of those examples is exotic tooling. They're the same three properties IaC already gave you for servers — versioned, reviewed, reproducible — applied to systems that usually don't get them.

Note

The common failure pattern: a team is genuinely disciplined about Terraform, then someone grants themselves S3 access "just for now" directly in the AWS console during an incident, and it's still there eight months later because it was never in a file anyone would notice was wrong. The infrastructure was code. The access to it wasn't. The blast radius of the thing that wasn't code turned out to matter more.

Why "manage everything like code" beats "buy more tools"

The instinct when a system feels chaotic is to reach for a new tool — a better dashboard, a nicer console, a more polished UI for managing whatever's currently a mess. Everything-as-code takes the opposite bet: the fix isn't a better interface for making undeclared, unreviewed changes. It's removing the undeclared, unreviewed change as an option at all.

That trade shows up directly in what breaks and how it gets fixed:

  • Consistent and repeatable — the tenth environment matches the first one, because both came from the same file, not from someone's memory of what they did last time.
  • Versioned and reviewed — every change to policy, access, or config goes through the same pull request process as application code, with the same paper trail.
  • Auditable by default — "who changed this and why" is a git log, not a support ticket asking whoever's been here longest.
  • Reproducible from scratch — a fresh environment isn't a checklist and a prayer, it's apply and a wait.

Start with the system that's currently the scariest

Trying to convert every system to code simultaneously is how everything-as-code initiatives stall — it's a multi-quarter program with no early win to point at. The better entry point is picking whichever undeclared system currently causes the most incidents or the most "wait, who changed this?" conversations, and converting just that one first. For a lot of teams that's access and IAM, because it's high-blast-radius and usually the least governed system in the stack; for others it's Kubernetes RBAC, or the alerting rules nobody can find the source of.

Where this is heading

Once infrastructure, workloads, policy and access are all declared in Git, a natural question shows up: if Git is the source of truth for all of it, why is a human — or a CI pipeline that only knows what it sent — still the one pushing changes into the cluster? Day 29 of this series is that exact question, and the answer is GitOps: a controller inside the cluster whose only job is making live state match what Git says it should be, continuously, without anyone running apply by hand.

Terraform got infrastructure out of consoles and into code. Everything-as-code is what happens when you stop treating that as the finish line and start asking which other system in your stack is still one console click away from silent drift.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.