I'm currently building a shared Terraform module set for landing zones — one module, applied consistently across subscriptions instead of every environment getting hand-rolled. The moment it's shared code, it stops being "infrastructure I run once" and becomes something closer to a library: other people's environments depend on it, and a change I make on a Tuesday can break something three teams and two weeks away from me.
Imagine this: someone opens a PR against that module changing a variable's default value — a small, reasonable-looking tweak, maybe tightening a network CIDR range or changing a default SKU. terraform plan runs in CI, the diff looks sane, a reviewer skims it, it merges. Two weeks later, a subscription that never explicitly set that variable — because the old default was exactly what it needed — quietly starts provisioning something different. Nobody notices until an audit, or until something in that subscription breaks in a way that takes an afternoon to trace back to a module change from two weeks ago.
Why this is a real software engineering problem, not just an ops one
If that same kind of change happened in an application's shared library, most teams would expect a test suite to catch it before merge. Infrastructure code rarely gets that same discipline, and I think it's mostly habit rather than a good reason — terraform plan shows you a diff for this configuration, right now, but it doesn't tell you whether the module still behaves correctly for every other configuration depending on it. A green plan on the PR's own example doesn't mean the module's contract with everyone else who consumes it is still intact.
That's the part that makes this a testing-strategy problem, not just a "review changes carefully" problem. Careful review catches things a human can reason about by reading a diff. It doesn't catch a default that mattered to a consumer nobody thought to check.
The concept: treat the module's inputs and outputs as a contract
The useful shift is thinking of a Terraform module the way you'd think about a function signature: given these inputs, it should produce these outputs and this shape of resources, every time. Testing means verifying that contract holds — not just "does this specific plan look right," but "does the module still behave correctly across the range of inputs it's actually used with."
That naturally splits into a few layers, roughly cheapest-and-fastest to most-expensive-and-slowest:
- Static validation — does the code even parse correctly, follow conventions, avoid known-bad patterns? Catches typos and style issues before anything else runs.
- Plan-time / unit tests — given specific inputs, does the module produce the expected plan, without touching real infrastructure? Fast, free, and where most of the value is.
- Integration tests — apply the module against real cloud resources, assert on actual state, then tear it down. Slow, costs real money, but the only layer that catches things a plan can't predict — actual provider behavior, real API limits, real dependency ordering.
What I'd actually use to build this
terraform validateandtflintfor the static layer — fast, catches syntax and common misconfigurations before anything else runs.- Terraform's native
terraform testframework (.tftest.hclfiles) for plan-time assertions — it's built into Terraform itself now, no extra tooling required, and it can run entirely againstplanwithout provisioning anything. - Checkov or Conftest/OPA for policy-level checks — the same "policy as code" idea applied to the module itself, not just what it provisions. If the landing zone module should never produce a public storage endpoint, that's a rule the test suite should enforce, not something a reviewer has to remember to look for.
- Terratest (Go-based) for the integration layer, when a change is significant enough to justify actually standing up real resources and tearing them down again.
- GitHub Actions to run static and unit-layer tests on every PR, with the slower integration layer gated to run less often — nightly, or only on changes to specific high-risk parts of the module.
What I'd actually investigate first
- List every place this module is actually consumed, and what inputs each consumer passes — including which ones rely on defaults rather than setting values explicitly. This is the step that reveals how much implicit behavior a "small" change could actually touch.
- Identify which module behaviors are load-bearing — a default CIDR range that three teams rely on is a bigger deal to change than an internal tag format nobody reads.
- Write the plan-time tests for the contract that already exists, before changing anything — tests that describe current behavior are what let you change the module with confidence later, not tests written after the fact to match whatever the code happens to do now.
- Decide what's worth an integration test versus what plan-time assertions already cover — not everything needs a real
apply, and treating everything as equally risky just makes the suite too slow for anyone to want to run it. - Wire the fast layers into CI on every PR, and make the slow layer's absence a visible, deliberate trade-off — not something that silently never runs because nobody set it up.
What would I do?
- Write
.tftest.hclplan-time tests that assert on the shape of the plan for the module's most common input combinations — including the defaults, since that's exactly what broke in the scenario above. - Add a policy check (Checkov or Conftest) for the specific things that would be a real incident if they slipped through — public network exposure, missing required tags, an overly broad IAM assignment.
- Reserve Terratest-style real-apply integration tests for the parts of the module that are genuinely hard to predict from a plan — provider-specific quirks, resources with asynchronous provisioning behavior, anything that's bitten the team before.
- Run the fast layers (validate, lint, plan-time tests, policy checks) on every single PR, blocking merge on failure.
- Treat a changed default value as a deliberate, documented decision — not a drive-by tweak — precisely because it's invisible to every consumer who didn't explicitly override it.
What I would NOT do
The shortcuts that cause the actual incidents
Trusting a clean terraform plan on one example configuration as proof the module is safe for every consumer — a plan only tells you about the inputs you gave it. Skipping tests because "it's just infrastructure code" — the blast radius of a bad merge here is often larger than an application bug, not smaller. Writing only integration tests because they feel more "real" — they're slow enough that people stop running them locally, and a test suite nobody runs might as well not exist. And changing a shared default without checking who currently depends on the old one — that's precisely the failure mode from the opening scenario, and it's checkable in a few minutes if you actually look.
What a reasonable test pyramid looks like for a module like this
/\
/ \ Integration (Terratest)
/----\ slow, real cost, real infra — sparingly
/ \
/--------\ Plan-time tests (terraform test)
/ \ fast, free — most of the value lives here
/------------\
/--------------\ Static validation (validate, tflint, policy)
/________________\ fastest, runs on every keystrokeA concrete example
A plan-time test asserting the module still produces the expected resource shape for its default inputs — no real infrastructure touched:
# tests/defaults.tftest.hcl
run "default_network_cidr" {
command = plan
assert {
condition = azurerm_virtual_network.this.address_space[0] == "10.0.0.0/16"
error_message = "Default CIDR changed — check every consumer relying on this default."
}
}
run "no_public_storage_by_default" {
command = plan
assert {
condition = azurerm_storage_account.this.public_network_access_enabled == false
error_message = "Storage must not be publicly accessible by default."
}
}That second assertion is the kind of thing a reviewer might miss on a busy Friday and a test never will. It runs in seconds, costs nothing, and turns "someone has to remember to check this" into "CI checks this, every time, automatically."
The trade-offs, honestly
Writing plan-time tests takes real time up front, and it's easy to under-invest in this the first time you build a module — nothing forces you to until the first silent breakage happens. Integration tests catch real bugs plan-time tests structurally can't, but they cost actual cloud spend, run slowly enough that people avoid running them casually, and can leave orphaned resources behind if a test fails mid-teardown — which means the integration layer itself needs its own cleanup discipline. And policy checks are only as good as the rules someone thought to write; they won't catch a novel failure mode nobody's encoded yet. None of this makes the module bulletproof. It shifts the odds — from "a bad change is caught if someone happens to notice" to "a bad change is caught automatically, every time, for the things worth encoding as a rule."
Where this connects beyond Terraform
This is a platform engineering problem as much as a testing one — a landing zone module is infrastructure serving as a product, and a product needs a test suite the same way any other internal platform does, not because it's trendy but because "many teams depend on this and can't easily tell when it changed underneath them" is exactly the situation tests exist for. It connects to Day 41's policy-as-code argument directly — the module-level checks here are the same idea, just enforcing rules about the module's own output instead of what gets deployed through it. And it's relevant to SRE thinking too: a silent infrastructure regression is a reliability problem before it's ever an incident, and catching it at PR time is strictly cheaper than catching it in production. Worth noting for the AI-assisted future too — if an agent is ever the one proposing changes to a shared module, the test suite is exactly what lets you trust its output without reviewing every line by hand; the tests become the contract the agent has to satisfy, not just the humans.
My takeaway
- A clean
terraform planon one example proves the module works for that example — nothing about every other consumer. - Treat a module's inputs and defaults as a contract with whoever depends on them, and treat changing a default as changing that contract.
- Plan-time tests (
terraform test) are where most of the value is — fast, free, and they run on every PR without anyone having to remember to run them. - Reserve real-apply integration tests for what plan-time tests genuinely can't predict — don't make the whole suite slow trying to cover everything at that layer.
- The point isn't a bulletproof module. It's catching the specific, predictable ways this actually breaks, automatically, instead of hoping a reviewer notices.
One thing to actually try
Pick one Terraform module you already maintain and write a single .tftest.hcl file that asserts on its current default behavior — nothing fancy, just "given no overrides, this is what gets planned." Run it, watch it pass. Now you have a change detector for the exact kind of silent default-value drift that's easy to miss in a review and easy to catch in a test.





