Skip to content
GCP

GCP IAM: Why I'd Never Download a Service Account Key Again

A downloaded JSON key with project-level Editor is the fastest way to get a CI/CD pipeline working on GCP — and the fastest way to hand out a credential that never expires and no one is watching. Here's how I'd actually set this up.

Thamunkpillai 10 min read

Imagine this: you inherit a GitHub Actions pipeline that deploys to GKE, and buried in the repo's secrets is GCP_SA_KEY — a base64-encoded JSON service account key, downloaded from the console at some point, pasted into a GitHub secret, and never touched since. You go looking for the service account it belongs to and find it has roles/editor at the project level. Nobody remembers exactly why. It's just always been there, and the pipeline works, so nobody's questioned it.

That key doesn't expire. It doesn't rotate unless someone remembers to rotate it. And roles/editor means whoever holds that key — or whoever finds it, if that secret ever leaks — can modify almost everything in the project: compute, storage, networking, IAM policies on most resources. Not because anyone made a deliberate decision to grant that much access, but because it was the fastest way to get a pipeline unblocked, and "fastest" quietly became "permanent."

Why this specific shortcut is worse than it looks

A downloaded service account key is a long-lived, static, bearer credential — whoever has the file has full access, indistinguishable from the legitimate pipeline, with no built-in expiry. Compare that to a human logging in: they authenticate every session, through an identity provider, with MFA, and their access can be revoked centrally in seconds. A JSON key sitting in a CI secret store has none of that. If it leaks — committed to a repo by accident, pulled from a compromised CI runner, copied out of a secrets manager by someone who shouldn't have access — it's valid until someone notices and manually revokes it, which in practice can be a very long time.

And roles/editor compounds the problem. It's one of GCP's three primitive roles (Owner, Editor, Viewer), and it predates IAM's more granular predefined and custom roles. It's broad by design — meant for a time when "can edit most things in this project" was an acceptable grant. It almost never matches what a specific pipeline or service actually needs to do.

The concept that actually fixes this: no keys, just federated trust

GCP's IAM resource hierarchy — Organization → Folder → Project → Resource — works the same way Azure's does: a role granted higher up is inherited everywhere below it, so scope matters as much as which role you pick. But the deeper fix here isn't just "pick a narrower role." It's removing the static key entirely.

Workload Identity Federation (WIF) lets an external identity — a GitHub Actions run, an AWS workload, anything that can produce a signed token — exchange that token for short-lived GCP credentials, without a service account key ever existing. GitHub's OIDC token proves "this is a specific workflow, in this specific repo, on this specific branch," GCP verifies that against a configured trust policy, and hands back credentials that expire in about an hour. There's no file to leak, because there's no file.

The equivalent for workloads running inside GKE is GKE Workload Identity, which lets a Kubernetes ServiceAccount impersonate a GCP service account natively — a pod authenticates as itself, tied to its Kubernetes identity, again with no key file mounted into the container.

What I'd actually use to build this

  • Terraform (google provider) to define the workload identity pool, provider, and IAM bindings as code — reviewable, versioned, no console click-ops.
  • google-github-actions/auth in the GitHub Actions workflow, configured against the WIF provider instead of a stored key.
  • GKE Workload Identity for anything running as a pod that needs to call GCP APIs — Cloud Storage, Pub/Sub, whatever the service actually touches.
  • Custom or predefined IAM roles, scoped to the specific APIs a pipeline or workload actually calls, instead of Editor.
  • An Org Policy constraint (constraints/iam.disableServiceAccountKeyCreation) to stop new keys from being created in the first place — the guardrail that holds even when someone's in a hurry and reaches for the old habit.

What I'd actually investigate first

  1. List every service account key that currently exists. gcloud iam service-accounts keys list --iam-account=<sa-email> per service account, across every project — this is tedious and exactly the step that reveals how much is actually out there versus what you assumed.
  2. Check what each key is actually used for, not what it was originally requested for. Pipelines get repurposed; the key's current blast radius is what matters, not its history.
  3. Map each use case to the narrowest role that covers it — a deploy pipeline pushing to GKE needs specific GKE and Artifact Registry permissions, not project-wide Editor.
  4. Set up the WIF pool and provider once, centrally, so every new pipeline plugs into the same trust configuration instead of each team inventing its own.
  5. Migrate one pipeline first, prove it end to end, then move the rest — this is exactly the kind of change where a broken deploy pipeline is loud and immediate, so you want to validate the pattern before applying it everywhere.

What would I do?

  • Set up Workload Identity Federation for every CI/CD pipeline that deploys to GCP, and treat any remaining service account key as a migration backlog item, not a permanent fixture.
  • Use GKE Workload Identity for every pod that calls a GCP API — no mounted key files, no GOOGLE_APPLICATION_CREDENTIALS pointing at a secret volume.
  • Scope every service account to the narrowest predefined or custom role that covers its actual calls — start narrow, widen only when something concretely breaks, not preemptively.
  • Turn on the org policy constraint disabling service account key creation, project by project, once the WIF migration is far enough along that it won't just get worked around.
  • Run IAM Recommender periodically — it flags roles granted that haven't actually been used, which is the fastest way to find scope you can safely remove.

What I would NOT do

The shortcuts that cause the actual incidents

Granting Editor or Owner because figuring out the precise predefined role feels like it'll slow the pipeline down today — it will slow down an incident review far more later. Downloading a "just for testing" service account key and letting it quietly become the permanent CI credential — temporary keys have the same habit as temporary access anywhere else: nobody's job is to remove them. Storing a key in a secrets manager and calling that "secure" — a secrets manager reduces exposure, it doesn't remove the fact that the credential is static and long-lived. And disabling key creation org-wide before any pipeline has actually been migrated to WIF — that just breaks deploys and teaches people to route around the platform team instead of trusting it.

How the trust flow actually looks

Text
GitHub Actions workflow run
        |
        v
GitHub OIDC token (short-lived, identifies repo + branch)
        |
        v
GCP Workload Identity Pool  --verifies trust policy-->  Provider
        |
        v
Short-lived GCP access token (~1 hour)
        |
        v
Scoped IAM role, specific APIs only
        |
        v
   Deploy runs, token expires

A concrete example

The Terraform side — a workload identity pool and provider trusting a specific GitHub repo, nothing broader:

Terraform
resource "google_iam_workload_identity_pool" "github" {
  workload_identity_pool_id = "github-actions-pool"
}
 
resource "google_iam_workload_identity_pool_provider" "github" {
  workload_identity_pool_id          = google_iam_workload_identity_pool.github.workload_identity_pool_id
  workload_identity_pool_provider_id = "github-provider"
  attribute_mapping = {
    "google.subject"       = "assertion.sub"
    "attribute.repository" = "assertion.repository"
  }
  attribute_condition = "assertion.repository == 'my-org/my-repo'"
  oidc {
    issuer_uri = "https://token.actions.githubusercontent.com"
  }
}

And the GitHub Actions side — no GCP_SA_KEY secret anywhere:

YAML
- uses: google-github-actions/auth@v2
  with:
    workload_identity_provider: projects/123456/locations/global/workloadIdentityPools/github-actions-pool/providers/github-provider
    service_account: deploy-pipeline@my-project.iam.gserviceaccount.com

The attribute_condition line is the part worth not skipping — without it, the pool trusts GitHub's OIDC issuer generally, which means any GitHub Actions workflow, in any repository, could potentially authenticate as that service account. Scoping it to a specific repository (and branch, if you want to go further) is what makes this actually narrow rather than just keyless.

The trade-offs, honestly

Setting up WIF has more moving parts up front than pasting a key into a secret — a pool, a provider, an attribute mapping, and it takes real testing to get the trust conditions right before the first deploy succeeds. For a single pipeline in a hurry, a key is genuinely faster today. The payoff is entirely about what happens later: no key to rotate, no key to leak, and a credential that's cryptographically tied to a specific workflow run instead of a file that works for anyone who has it. Custom IAM roles are more precise than primitive ones, but somebody has to keep them in sync as a service's actual API usage changes — a role that's too narrow breaks the pipeline the moment the service starts calling one more API, which is annoying but a far better failure mode than "too broad and nobody notices."

Where this connects beyond GCP

This is the same identity discipline the Azure landing zone piece got to from a different direction — static, long-lived, broadly-scoped credentials are the problem, regardless of which cloud you're standing in. It's a platform engineering concern because the fix belongs in a shared module and a golden path, not in every team's separately-invented pipeline. It's a security and SRE concern because a leaked static credential and a misconfigured deploy are two different root causes with the same blast-radius conversation. And it matters for AI agents specifically — an agent given standing, broadly-scoped access to provision or modify cloud infrastructure is exactly the downloaded-key problem again, just with a non-human actor behind it. The same answer applies: short-lived, narrowly-scoped, attributable to a specific identity, never a static credential sitting somewhere waiting to be found.

My takeaway

  • A downloaded service account key is a static, unexpiring bearer credential — the risk isn't hypothetical, it's just waiting for the file to end up somewhere it shouldn't.
  • Workload Identity Federation removes the key entirely instead of managing it more carefully — that's a better trade than any rotation policy.
  • Primitive roles (Owner, Editor, Viewer) are almost always broader than what's actually needed — treat a request for one as a sign the real requirement hasn't been mapped out yet.
  • attribute_condition is what makes WIF actually scoped — set it up without one and you've built a keyless system that's still too trusting.
  • Migrate one pipeline first. This is a change where "it broke the deploy" is loud and immediate, so prove the pattern before rolling it out everywhere.

One thing to actually try

Run gcloud iam service-accounts keys list --iam-account=<your-sa-email> against whatever service account your main CI/CD pipeline uses. If you see a USER_MANAGED key with a creation date from months ago, that's your answer to whether this is worth doing — and a real candidate for the first pipeline to migrate to Workload Identity Federation.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.