The questions that decide a Google Cloud platform are the same ones that decide an AWS platform: what failure it must survive, what it may cost, and how a change reaches it. What differs is the primitives, and two things Google Cloud makes notably easier — regional clusters and identity federation.

This describes the shape we build, and the two parts teams most often get wrong: environments that drift, and test data that is either useless or unlawful.

The cluster is regional, not zonal

A zonal GKE cluster has its control plane in one zone. When that zone has a bad day, you keep your workloads but lose the ability to change them, which during an incident is close to the same thing.

A regional cluster replicates the control plane across three zones in the region and spreads nodes across them. The cost difference is the control plane charge and some cross-zone traffic. The difference in what you can do at 3am is total.

Regional GKE, one project per environment
Regional GKE cluster across three zones, with a global front endAn anycast global external load balancer with Cloud CDN and Cloud Armor terminates TLS at the edge. Behind it a regional GKE cluster replicates its control plane across three zones and spreads private nodes across them. State is in Cloud SQL with regional high availability, with Memorystore for caching and an optional read replica.GOOGLE CLOUDone project per environmentVPCprivate nodes, Cloud NAT egressZONEaZONEbZONEcingressasyncCloud CDNanycast front endCloud ArmorWAF and DDoSGlobal external load balancerTLS at the edgeGKE control planeregional: replicated across all three zonesNodesno public IPNodesno public IPNodesno public IPCloud SQLregional HA, synchronousMemorystorecacheRead replicaoptional
Cluster
regional, private nodes
Egress
Cloud NAT, no public node addresses
Identity
Workload Identity Federation, no exported keys
Isolation
one project per environment, so IAM follows the boundary
This diagram as text
  • Cloud CDN — anycast front end
  • Cloud Armor — WAF and DDoS
  • Global external load balancer — TLS at the edge
  • Google Cloud one project per environment
    • VPC private nodes, Cloud NAT egress
      • GKE control plane — regional: replicated across all three zones
      • Zone a
        • Nodes — no public IP
      • Zone b
        • Nodes — no public IP
      • Zone c
        • Nodes — no public IP
      • Cloud SQL — regional HA, synchronous
      • Memorystore — cache
      • Read replica — optional

Relationships

  • Cloud CDN → Cloud Armor
  • Cloud Armor → Global external load balancer
  • Global external load balancer → GKE control plane — ingress
  • GKE control plane → Nodes (Zone a)
  • GKE control plane → Nodes (Zone b)
  • GKE control plane → Nodes (Zone c)
  • Cloud SQL → Read replica — async

One project per environment rather than one project with namespaces for everything. It makes the IAM boundary the project boundary, makes cost attribution automatic, and means a mistake in a lower environment cannot reach production IAM at all.

No service account keys, anywhere

An exported service account key is a password in a JSON file. It does not expire, it is frequently committed by accident, and once copied it is indistinguishable from legitimate use.

Workload Identity Federation removes them in both directions:

  • For CI: GitHub Actions presents its OIDC token to a workload identity pool, which exchanges it for a short-lived access token. The pool’s attribute condition pins it to the repository and the branch, so a fork cannot obtain production credentials.
  • For workloads in the cluster: a Kubernetes service account is bound to an IAM principal, so a pod calls Google APIs as itself. The node identity is not the workload identity, which is what keeps a compromised pod from inheriting everything the node can do.

The result is an estate where gcloud iam service-accounts keys list returns nothing, and that is a control you can demonstrate to an auditor in one command.

Ephemeral environments, per pull request

A permanent staging environment is a liability dressed as a safety net. It drifts from production with every manual fix, its data diverges, and the first time anyone notices is when something passes staging and fails in production.

We provision a complete environment per pull request, from the same Terraform modules production uses, with only the variables changed.

Pull request lifecycle
Per-pull-request environment lifecycle, provisioned and destroyed from the production modulesOpening a pull request runs terraform apply against the same modules production uses, with only the variables changed, and seeds the database from the latest obfuscated snapshot. The environment URL is posted back to the pull request. Migrations run forward, then the integration, end-to-end, accessibility and performance suites run. On merge or at a time-to-live, a reaper destroys it.GATESPull request openedterraform applyproduction modules, different variablesSeed the databaselatest obfuscated snapshotEnvironment URL posted back to the pull requestMigrations run forwardon real-shaped dataIntegration and end-to-endAccessibility and performance budgetsterraform destroyon merge, or when the TTL reaper runs
Built from
production modules, different variables
Lifetime
until merge, or a TTL that reaps it
Data
obfuscated snapshot, never raw production
This diagram as text
  • Pull request opened
  • terraform apply — production modules, different variables
  • Seed the database — latest obfuscated snapshot
  • Environment URL posted back to the pull request
  • Gates
    • Migrations run forward — on real-shaped data
    • Integration and end-to-end
    • Accessibility and performance budgets
  • terraform destroy — on merge, or when the TTL reaper runs

Relationships

  • Pull request opened → terraform apply
  • Pull request opened → Seed the database
  • terraform apply → Environment URL posted back to the pull request
  • Environment URL posted back to the pull request → Migrations run forward
  • Environment URL posted back to the pull request → Integration and end-to-end
  • Environment URL posted back to the pull request → Accessibility and performance budgets
  • Migrations run forward → terraform destroy
  • Integration and end-to-end → terraform destroy
  • Accessibility and performance budgets → terraform destroy

Two details decide whether this works in practice. The environment must carry a time-to-live and a reaper job, because pull requests get abandoned and an environment nobody destroys is a bill nobody expected. And the database migration must run forward on real-shaped data in that environment, because a migration that works on an empty schema and locks a large table for nine minutes is a production incident that was fully preventable.

Refreshing lower environments safely

Teams need production-shaped data to test against. Synthetic data misses the distributions and the edge cases that cause the failures worth catching. Copying production is how organisations end up with customer records in a developer laptop’s port-forward.

The order of operations is the whole control:

Scheduled refresh — obfuscation happens before export
Scheduled data refresh where obfuscation happens inside the production boundaryA point-in-time snapshot is restored into an isolated scratch instance with no application and no human access. An obfuscation job masks direct identifiers, re-keys consistently and shifts dates by a constant offset per subject. A verification scan confirms no personal-data patterns remain. Only then does anything leave the production boundary, into staging, ephemeral pull-request environments and the performance rig.PRODUCTION BOUNDARYnothing leaves un-obfuscatedLOWER ENVIRONMENTSrestoreexportSnapshotpoint in timeIsolated scratch instanceno application, no human accessObfuscation jobmask · re-key consistently · shift datesVerification scanno personal-data patterns remainStagingEphemeral PR environmentsPerformance rig
Cadence
scheduled, and before any major release
Referential integrity
preserved by consistent re-keying
Dates
shifted by a constant offset per subject, so intervals survive
Evidence
the verification result is retained per run
This diagram as text
  • Production boundary nothing leaves un-obfuscated
    • Snapshot — point in time
    • Isolated scratch instance — no application, no human access
    • Obfuscation job — mask · re-key consistently · shift dates
    • Verification scan — no personal-data patterns remain
  • Lower environments
    • Staging
    • Ephemeral PR environments
    • Performance rig

Relationships

  • Snapshot → Isolated scratch instance — restore
  • Isolated scratch instance → Obfuscation job
  • Obfuscation job → Verification scan
  • Verification scan → Staging — export
  • Verification scan → Ephemeral PR environments
  • Verification scan → Performance rig

What the obfuscation job actually does matters as much as that it runs:

  • Mask direct identifiers — names, emails, phone numbers, national identifiers — with generated values of the same shape, so format validation still exercises the same code paths.
  • Re-key consistently, so the same source value maps to the same replacement everywhere. Without this, joins break and the data stops being useful for exactly the tests you wanted it for.
  • Shift dates by a constant offset per subject rather than randomising them, which preserves intervals and sequences while destroying the real calendar.
  • Drop columns that cannot be safely masked — free-text notes, uploaded documents, anything where personal data may be embedded in prose. A column you cannot verify is a column that does not travel.
  • Verify before export, with a scan for the patterns that must not survive. The scan result is retained, which turns the whole process into evidence for ISO/IEC 27001 Annex A 8.33 and for GDPR Article 32.

Running this on a schedule, and always before a major release, is what makes the performance test and the migration rehearsal mean something.

Scaling and cost

Node auto-provisioning plays the role Karpenter plays on AWS: it creates node pools shaped to the pending workload instead of scaling one pool of one machine type. On Autopilot the question disappears, and you pay per pod request instead — which makes honest resource requests a direct cost control rather than an indirect one.

The levers that matter, in the order they usually pay:

  1. Requests derived from measurement, not from a template. Over-stated requests are the largest single source of waste on any Kubernetes bill, and on Autopilot you are billed for them directly.
  2. Spot or preemptible nodes for anything interruptible, with PodDisruptionBudgets and graceful termination actually implemented rather than assumed.
  3. Committed use discounts once a baseline is genuinely stable, and not before.
  4. Lower environments that switch off. Ephemeral environments reaped on merge, and a schedule that scales non-production to zero outside working hours.

The same acceptance criteria

We accept a Google Cloud platform against the Google Cloud Architecture Framework the same way we accept an AWS platform against Well-Architected: operational excellence, security, reliability, performance and cost optimisation, each with a question that has to have a demonstrable answer. The framework differs; the discipline does not.

How STP approaches this

We build the cluster regional from the start, remove service account keys as a precondition rather than a follow-up, make environments ephemeral so they cannot drift, and put the obfuscation inside the production boundary so the refresh is lawful as well as useful. Where a platform already exists, we start with an architecture review that looks for the exported keys, the long-lived staging environment and the untested migration path.

More on Kubernetes platform engineering and DevOps and platform automation, or start a conversation.