Skip to content
Azure

Azure Landing Zone IAM: Why I Avoid the Owner Role (And What I Use Instead)

Owner is the fastest way to unblock someone in a new Azure subscription, and the fastest way to lose track of who can do what across your landing zones. Here's how I actually think about IAM while building landing zone modules in Terraform.

Thamunkpillai 9 min read

I'm currently building a Terraform module set meant to standardize landing zones across our Azure subscriptions — one module, applied consistently, instead of every subscription getting hand-rolled the day someone needs one. Most of that work is boring in a good way: management groups, policy assignments, networking baselines. The part I keep circling back to, rewriting, and second-guessing is access.

Imagine this: it's Thursday afternoon, a team lead pings you that their deployment pipeline is failing in a brand-new subscription because it "doesn't have permissions yet," and they need it fixed before a demo tomorrow. You could open the portal, assign Owner to the pipeline's service principal at the subscription scope, and be unblocked in ninety seconds. It always works. That's exactly the problem.

Why "just give it Owner" becomes a real problem later

Owner at subscription scope isn't just "full access to this subscription" — it's full access to every resource group under it, forever, until someone remembers to revoke it. Nobody remembers to revoke it. Six months later you've got a service principal that can delete anything in a subscription, was granted that access to fix a demo, and nobody currently on the team can tell you why.

Multiply that by every subscription in a landing zone model and you get what I think of as RBAC sprawl: dozens of direct role assignments, made under time pressure, at inconsistent scopes, with no shared pattern. When an audit asks "who can modify production networking," the honest answer is "we'd have to check," and that answer is the actual failure — not any single overprivileged assignment, but the fact that nobody can reason about access as a system anymore.

The concept that actually fixes this: scope discipline, not more tickets

Azure RBAC is hierarchical — management group, then subscription, then resource group, then resource — and a role assigned higher in that chain is inherited everywhere below it. That inheritance is exactly why Owner-at-subscription is so dangerous: it's not one grant, it's every grant, silently, at every layer underneath.

The fix isn't a stricter approval process for Owner requests. It's assigning the narrowest built-in or custom role, at the lowest scope that satisfies the actual need, through a group, not a person or a raw service principal. Almost nobody actually needs Owner. Most "I need Owner" requests are really "I need to manage resources in this resource group," which is Contributor scoped to a resource group, or narrower still.

What I'd actually use to build this

  • Terraform (azurerm provider) for every role assignment, defined as code and reviewed like any other change — no console click-ops that isn't reflected anywhere.
  • Microsoft Entra ID groups as the thing roles get assigned to, never individual users. People move between groups; role assignments don't need to change when they do.
  • Privileged Identity Management (PIM) for anything above read-level human access — engineers get eligible assignments they activate for a time-boxed window, not standing access.
  • Workload identity federation (OIDC) for CI/CD — GitHub Actions authenticates to Azure with a federated credential scoped to one subscription and one narrow role, no client secret sitting in a repo or a vault that can be copied out.
  • Azure Policy as the backstop — a policy that flags or denies direct role assignments made outside the pipeline, so scope discipline doesn't quietly erode the first time someone's in a hurry.

What I'd actually investigate first

Before writing a single Terraform resource, I'd want to know what the current state actually is — not what I assume it is:

  1. Pull every existing role assignment across every subscription. az role assignment list --all gives you the real picture, and it's usually worse than expected. This is the step people skip because it's tedious, and it's the step that tells you whether you have a design problem or a decade of accumulated shortcuts.
  2. Group assignments by "why does this exist." Anything nobody can explain is a candidate for removal, not for grandfathering in.
  3. Map real access needs by team and environment, not by role title — a platform engineer's dev-subscription needs look nothing like their production needs, and treating them the same is how Owner-everywhere happens in the first place.
  4. Design the management group hierarchy first, so scoping decisions have somewhere sane to land — an access policy without a hierarchy underneath it just becomes exceptions layered on exceptions.
  5. Write the custom roles that don't exist yet. Built-in roles are a reasonable starting point, but "deploy and manage this specific class of resource, nothing else" often has no built-in match, and forcing people into Contributor because the precise role doesn't exist is how scope creep sneaks back in.

What would I do?

  • Default every human access request to PIM-eligible, time-boxed, scoped to a resource group — not standing, not subscription-wide, unless there's a specific, written reason it needs to be.
  • Assign roles to Entra ID groups, and manage group membership as the access-control surface, not individual role assignments.
  • Give CI/CD pipelines federated workload identity scoped to exactly the subscription and resource type they deploy — a Terraform pipeline that manages networking doesn't need permission to touch Key Vault.
  • Write the landing zone's IAM baseline as a Terraform module, so every new subscription inherits the same access pattern by default instead of getting it improvised.
  • Add an Azure Policy that catches direct portal-based role assignments outside the pipeline, so the guardrail holds even under Thursday-afternoon pressure.

What I would NOT do

The shortcuts that cause the actual incidents

Assigning Owner "temporarily" to unblock something — temporary access has a way of becoming permanent the moment nobody's job is to remember to remove it. Granting roles to individual users instead of groups — you'll be cleaning up stale assignments for people who left the team two years ago. Giving a CI/CD service principal broader scope "in case we need it later" — that's not solving today's problem, it's pre-authorizing tomorrow's incident. And skipping the audit step because you're confident you already know what access looks like — you don't, not exactly, and "exactly" is what an audit actually asks for.

How the access flow actually looks end to end

Text
Engineer requests access
        |
        v
Entra ID group membership (PIM-eligible)
        |
        v
   Activate (time-boxed, justified)
        |
        v
Role assignment, scoped to resource group
        |
        v
   Access expires automatically
Text
CI/CD pipeline (GitHub Actions)
        |
        v
OIDC federated credential (no stored secret)
        |
        v
Role assignment, scoped to one subscription + role
        |
        v
   Terraform plan / apply

A concrete example

A group-based, resource-group-scoped assignment in Terraform looks like this — deliberately narrow, deliberately reviewable:

Terraform
resource "azurerm_role_assignment" "platform_contributor" {
  scope                = azurerm_resource_group.landing_zone.id
  role_definition_name = "Contributor"
  principal_id          = data.azuread_group.platform_engineers.object_id
}

Worth being honest about a real limitation here: Terraform's azurerm provider doesn't manage PIM eligible assignments natively — that still goes through the Entra ID Governance / Microsoft Graph API. In practice that means standing role assignments are code-managed and reviewable, but the time-boxed elevation layer on top needs a separate integration. I'd rather have that gap explicitly than pretend Terraform covers all of it.

The trade-offs, honestly

None of this is free. PIM activation adds friction to someone's day — a few seconds to a couple of minutes, depending on approval requirements — and the first reaction from an engineer used to standing access is usually mild annoyance. Custom roles are more precise than built-in ones, but somebody has to maintain them as Azure's own role definitions evolve, and that's ongoing work, not a one-time setup. And a fully centralized IAM module is more consistent but slower to change than letting individual teams manage their own subscription access — you're trading velocity for auditability, and which one you need more of depends on what stage the org is actually at, not on which one sounds better in a blog post.

Where this connects beyond Azure

This is really a platform engineering problem wearing an IAM costume. A landing zone is a platform, and "how do I get access to deploy here" is exactly the kind of request a golden path should answer through self-service — a PR against the access module, reviewed and merged — rather than a Slack message and a rushed portal click. It's also squarely an SRE and security concern: the blast radius of a compromised or misconfigured identity is the same conversation as the blast radius of an outage, just triggered by a different kind of failure. And it's worth thinking about now rather than later, because the next thing requesting access to your landing zone might not be a person or a CI pipeline — it might be an AI agent triaging infrastructure or proposing changes, and it will need exactly this same discipline: scoped, time-boxed, attributable identity, not a standing credential nobody's watching.

My takeaway

  • Owner is almost never the right answer — it's the fast answer, and those aren't the same thing.
  • Assign roles to groups, not people or raw service principals, or you'll spend years cleaning up assignments for people who've already moved on.
  • Time-boxed, activated access (PIM) beats standing access every time the actual need is occasional, which is most of the time.
  • Workload identity federation removes an entire category of secret-leak incidents for CI/CD — if you're still managing a client secret for a deploy pipeline, that's the next thing to fix.
  • An access model you can't audit in an afternoon isn't a model — it's just a history of individual decisions nobody remembers making.

One thing to actually try

Run az role assignment list --all --query "[?roleDefinitionName=='Owner']" -o table against one of your own subscriptions. Before you design anything, just look at how many Owner assignments already exist, and see how many of them you can actually explain. That list is usually the most convincing argument for scope discipline you'll find — more convincing than anything I could write here.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.