The questions that decide a Google Cloud platform are the same ones that decide an AWS platform: what failure it must survive, what it may cost, and how a change reaches it. What differs is the primitives, and two things Google Cloud makes notably easier — regional clusters and identity federation.
This describes the shape we build, and the two parts teams most often get wrong: environments that drift, and test data that is either useless or unlawful.
The cluster is regional, not zonal
A zonal GKE cluster has its control plane in one zone. When that zone has a bad day, you keep your workloads but lose the ability to change them, which during an incident is close to the same thing.
A regional cluster replicates the control plane across three zones in the region and spreads nodes across them. The cost difference is the control plane charge and some cross-zone traffic. The difference in what you can do at 3am is total.
- Cluster
- regional, private nodes
- Egress
- Cloud NAT, no public node addresses
- Identity
- Workload Identity Federation, no exported keys
- Isolation
- one project per environment, so IAM follows the boundary
This diagram as text
- Cloud CDN — anycast front end
- Cloud Armor — WAF and DDoS
- Global external load balancer — TLS at the edge
- Google Cloud one project per environment
- VPC private nodes, Cloud NAT egress
- GKE control plane — regional: replicated across all three zones
- Zone a
- Nodes — no public IP
- Zone b
- Nodes — no public IP
- Zone c
- Nodes — no public IP
- Cloud SQL — regional HA, synchronous
- Memorystore — cache
- Read replica — optional
- VPC private nodes, Cloud NAT egress
Relationships
- Cloud CDN → Cloud Armor
- Cloud Armor → Global external load balancer
- Global external load balancer → GKE control plane — ingress
- GKE control plane → Nodes (Zone a)
- GKE control plane → Nodes (Zone b)
- GKE control plane → Nodes (Zone c)
- Cloud SQL → Read replica — async
One project per environment rather than one project with namespaces for everything. It makes the IAM boundary the project boundary, makes cost attribution automatic, and means a mistake in a lower environment cannot reach production IAM at all.
No service account keys, anywhere
An exported service account key is a password in a JSON file. It does not expire, it is frequently committed by accident, and once copied it is indistinguishable from legitimate use.
Workload Identity Federation removes them in both directions:
- For CI: GitHub Actions presents its OIDC token to a workload identity pool, which exchanges it for a short-lived access token. The pool’s attribute condition pins it to the repository and the branch, so a fork cannot obtain production credentials.
- For workloads in the cluster: a Kubernetes service account is bound to an IAM principal, so a pod calls Google APIs as itself. The node identity is not the workload identity, which is what keeps a compromised pod from inheriting everything the node can do.
The result is an estate where gcloud iam service-accounts keys list returns nothing, and that is a
control you can demonstrate to an auditor in one command.
Ephemeral environments, per pull request
A permanent staging environment is a liability dressed as a safety net. It drifts from production with every manual fix, its data diverges, and the first time anyone notices is when something passes staging and fails in production.
We provision a complete environment per pull request, from the same Terraform modules production uses, with only the variables changed.
- Built from
- production modules, different variables
- Lifetime
- until merge, or a TTL that reaps it
- Data
- obfuscated snapshot, never raw production
This diagram as text
- Pull request opened
- terraform apply — production modules, different variables
- Seed the database — latest obfuscated snapshot
- Environment URL posted back to the pull request
- Gates
- Migrations run forward — on real-shaped data
- Integration and end-to-end
- Accessibility and performance budgets
- terraform destroy — on merge, or when the TTL reaper runs
Relationships
- Pull request opened → terraform apply
- Pull request opened → Seed the database
- terraform apply → Environment URL posted back to the pull request
- Environment URL posted back to the pull request → Migrations run forward
- Environment URL posted back to the pull request → Integration and end-to-end
- Environment URL posted back to the pull request → Accessibility and performance budgets
- Migrations run forward → terraform destroy
- Integration and end-to-end → terraform destroy
- Accessibility and performance budgets → terraform destroy
Two details decide whether this works in practice. The environment must carry a time-to-live and a reaper job, because pull requests get abandoned and an environment nobody destroys is a bill nobody expected. And the database migration must run forward on real-shaped data in that environment, because a migration that works on an empty schema and locks a large table for nine minutes is a production incident that was fully preventable.
Refreshing lower environments safely
Teams need production-shaped data to test against. Synthetic data misses the distributions and the edge cases that cause the failures worth catching. Copying production is how organisations end up with customer records in a developer laptop’s port-forward.
The order of operations is the whole control:
- Cadence
- scheduled, and before any major release
- Referential integrity
- preserved by consistent re-keying
- Dates
- shifted by a constant offset per subject, so intervals survive
- Evidence
- the verification result is retained per run
This diagram as text
- Production boundary nothing leaves un-obfuscated
- Snapshot — point in time
- Isolated scratch instance — no application, no human access
- Obfuscation job — mask · re-key consistently · shift dates
- Verification scan — no personal-data patterns remain
- Lower environments
- Staging
- Ephemeral PR environments
- Performance rig
Relationships
- Snapshot → Isolated scratch instance — restore
- Isolated scratch instance → Obfuscation job
- Obfuscation job → Verification scan
- Verification scan → Staging — export
- Verification scan → Ephemeral PR environments
- Verification scan → Performance rig
What the obfuscation job actually does matters as much as that it runs:
- Mask direct identifiers — names, emails, phone numbers, national identifiers — with generated values of the same shape, so format validation still exercises the same code paths.
- Re-key consistently, so the same source value maps to the same replacement everywhere. Without this, joins break and the data stops being useful for exactly the tests you wanted it for.
- Shift dates by a constant offset per subject rather than randomising them, which preserves intervals and sequences while destroying the real calendar.
- Drop columns that cannot be safely masked — free-text notes, uploaded documents, anything where personal data may be embedded in prose. A column you cannot verify is a column that does not travel.
- Verify before export, with a scan for the patterns that must not survive. The scan result is retained, which turns the whole process into evidence for ISO/IEC 27001 Annex A 8.33 and for GDPR Article 32.
Running this on a schedule, and always before a major release, is what makes the performance test and the migration rehearsal mean something.
Scaling and cost
Node auto-provisioning plays the role Karpenter plays on AWS: it creates node pools shaped to the pending workload instead of scaling one pool of one machine type. On Autopilot the question disappears, and you pay per pod request instead — which makes honest resource requests a direct cost control rather than an indirect one.
The levers that matter, in the order they usually pay:
- Requests derived from measurement, not from a template. Over-stated requests are the largest single source of waste on any Kubernetes bill, and on Autopilot you are billed for them directly.
- Spot or preemptible nodes for anything interruptible, with PodDisruptionBudgets and graceful termination actually implemented rather than assumed.
- Committed use discounts once a baseline is genuinely stable, and not before.
- Lower environments that switch off. Ephemeral environments reaped on merge, and a schedule that scales non-production to zero outside working hours.
The same acceptance criteria
We accept a Google Cloud platform against the Google Cloud Architecture Framework the same way we accept an AWS platform against Well-Architected: operational excellence, security, reliability, performance and cost optimisation, each with a question that has to have a demonstrable answer. The framework differs; the discipline does not.
How STP approaches this
We build the cluster regional from the start, remove service account keys as a precondition rather than a follow-up, make environments ephemeral so they cannot drift, and put the obfuscation inside the production boundary so the refresh is lawful as well as useful. Where a platform already exists, we start with an architecture review that looks for the exported keys, the long-lived staging environment and the untested migration path.
More on Kubernetes platform engineering and DevOps and platform automation, or start a conversation.

