Most AWS cost and reliability problems are not caused by the services chosen. They are caused by a platform that grew one decision at a time, where nobody wrote down what it was supposed to survive or what it was supposed to cost. The design below is the shape we build to, and the reasoning behind each choice, so the trade-offs are visible rather than inherited.

Start from the failure you must survive

Before any service is selected, two numbers get agreed in writing: how long the service may be unavailable, and how much data may be lost. Everything downstream is derived from them.

A single Availability Zone is a single building. Three zones is the default because it is the point at which a quorum-based datastore can lose one and keep writing. A second region is a different question entirely, and most estates do not need it.

Baseline production topology, one region
Baseline production topology on AWS across three availability zonesUsers resolve through Route 53 to CloudFront with AWS WAF at the edge. CloudFront reaches an Application Load Balancer in the public subnets of a VPC. The load balancer distributes to EKS node groups in private subnets in three availability zones. State is in RDS PostgreSQL with a synchronous standby in a second zone and an asynchronous read replica in a third.AWSone regionVPCpublic and private subnetsAVAILABILITY ZONEaAVAILABILITY ZONEbAVAILABILITY ZONEcresolvealiasoriginsyncasyncUsersHTTPSRoute 53health-checkedCloudFront + WAFTLS terminated at the edgeApplication Load Balancerpublic subnets, one per zonetarget group per serviceEKS nodesprivate subnetRDS primarywritesEKS nodesprivate subnetRDS standbyautomatic failoverEKS nodesprivate subnetRead replicaoptional
Zones
three, minimum
Public surface
CloudFront and the load balancer, nothing else
Node subnets
private, egress through NAT
Data tier
Multi-AZ, failover operated by AWS
This diagram as text
  • Users — HTTPS
  • Route 53 — health-checked
  • CloudFront + WAF — TLS terminated at the edge
  • AWS one region
    • VPC public and private subnets
      • Application Load Balancer — public subnets, one per zone — target group per service
      • Availability zone a
        • EKS nodes — private subnet
        • RDS primary — writes
      • Availability zone b
        • EKS nodes — private subnet
        • RDS standby — automatic failover
      • Availability zone c
        • EKS nodes — private subnet
        • Read replica — optional

Relationships

  • Users → Route 53 — resolve
  • Route 53 → CloudFront + WAF — alias
  • CloudFront + WAF → Application Load Balancer — origin
  • Application Load Balancer → EKS nodes (Availability zone a)
  • Application Load Balancer → EKS nodes (Availability zone b)
  • Application Load Balancer → EKS nodes (Availability zone c)
  • RDS primary ↔ RDS standby — sync
  • RDS standby → Read replica — async

Nodes sit in private subnets. The only things with a public address are CloudFront and the load balancer. This is not a formality: it is what makes the blast radius of a compromised container a lateral movement problem rather than an immediate internet-facing one.

EKS, ECS or Fargate

This is usually discussed as a question about team size and Kubernetes experience. It is not. It is a question about how much the business is willing to be locked in, decided years before the lock-in is felt.

ECS is proprietary. It is a good orchestrator, it is cheaper to operate than a cluster, and it exists only on AWS. Every ECS task definition, every service discovery integration and every surrounding AWS primitive is work that does not move. That is a perfectly reasonable trade when the business is certain it will stay on AWS for the long term, and the certainty has to be real rather than assumed, because the bill for changing your mind arrives as a rewrite.

EKS is Kubernetes. The same substrate runs on another cloud and on hardware in your own room. What EKS buys is that AWS operates the control plane, which is the heaviest and least interesting part of running Kubernetes, while the workloads themselves stay portable. For an organisation that holds to not locking itself into proprietary technology, that is the point: you are renting the operational burden, not the architecture.

ECS when staying is a decision. EKS when leaving has to remain possible.

The second half of that only holds if portability is maintained deliberately. An EKS cluster wired into a dozen AWS-only services through the application code is locked in exactly as firmly as ECS, and has paid the Kubernetes complexity for nothing. Keeping the option open means:

  • Infrastructure as code written so the platform can be provisioned elsewhere. The modules are parameterised by provider-specific detail rather than built around it, and somebody has actually stood the platform up somewhere else, at least once, to prove the claim.
  • Managed services reached through interfaces the application does not hard-code. A managed database is sensible. A managed database whose proprietary extensions are in the query layer is a migration.
  • Portability tested as part of disaster recovery, not as a theory. If the recovery plan says the platform can be rebuilt on another provider, that is a claim with a clock on it, and it is cheap to check once a year.

Done that way, the same manifests and the same pipelines that run on EKS run on GKE or on a cluster on your own hardware, and a provider decision stops being irreversible.

Choose it when What you are accepting
ECS on Fargate Staying on AWS is a settled business decision, and you want the least to operate Proprietary. Moving later is a rewrite, not a migration
ECS on EC2 Same, with predictable load where the instance discount matters Same lock-in, plus you own the instances
EKS with Karpenter Portability is a requirement, and you want AWS to run the control plane Kubernetes complexity, an upgrade cadence, and the discipline to stay portable
EKS Fargate profiles Isolation per pod, or bursty jobs you do not want to keep nodes for Higher unit cost, no DaemonSets, slower start

A single estate frequently uses two. A common and sensible pattern is EKS for the application platform with a Fargate profile carrying the system namespace, so cluster add-ons are not competing with application pods for nodes and a node problem cannot take out the controllers that would otherwise fix it.

Node provisioning: Karpenter, not fixed node groups

A managed node group scales a fixed instance type between a minimum and a maximum. When a pod cannot be scheduled, the group adds another of the same instance, whether or not that shape fits.

Karpenter watches for unschedulable pods and provisions an instance that actually fits what is pending, choosing across families, sizes and purchase options. The part that produces the saving is consolidation: as load falls, it repacks workloads onto fewer nodes and terminates what is left empty.

What we set, and why:

  • Multiple instance families in the requirements, not one. Constraining Karpenter to a single family removes most of its ability to find capacity and to find it cheaply. Constrain on what the workload truly needs — architecture, CPU-to-memory ratio, local storage — and let it choose.
  • Graviton where the image supports it. ARM64 instances carry a materially better price-performance ratio. The work is in the build pipeline, not the platform: a multi-architecture image, and a test run that actually exercises the ARM build before it is promoted.
  • Spot for anything that can be interrupted, on-demand for what cannot. Karpenter handles the two-minute interruption notice by cordoning and draining, which is only safe if the workload has a PodDisruptionBudget and terminates gracefully. Spot without those two things is how a cost optimisation becomes an incident.
  • Consolidation enabled, with a disruption budget. Otherwise repacking can move more pods at once than the application can absorb.

Scaling has three independent axes

Teams routinely conflate these, then wonder why a scaling event did nothing.

Three scaling mechanisms, three different problems
Three scaling mechanisms on AWS, the signal each reads and what it changesThe horizontal pod autoscaler reads requests per second or queue depth and changes the replica count. The vertical pod autoscaler reads observed usage and recommends resource requests. Karpenter reads pending pods and changes the number and shape of nodes.READSCHANGESRequests per second, concurrency, queuedepthObserved CPU and memory, over timePods that cannot be scheduledHorizontal pod autoscalerHPAVertical pod autoscalerVPA, recommendation modeKarpenternode provisioningReplica counthandles trafficResource requestssets what the scheduler reservesNode count and instance shapesupplies the room
Do not combine
HPA and VPA on one deployment and one metric will oscillate
Scale on
the thing that describes arriving work, not CPU
Queue consumers
queue depth, and nothing else
This diagram as text
  • Reads
    • Requests per second, concurrency, queue depth
    • Observed CPU and memory, over time
    • Pods that cannot be scheduled
  • Horizontal pod autoscaler — HPA
  • Vertical pod autoscaler — VPA, recommendation mode
  • Karpenter — node provisioning
  • Changes
    • Replica count — handles traffic
    • Resource requests — sets what the scheduler reserves
    • Node count and instance shape — supplies the room

Relationships

  • Requests per second, concurrency, queue depth → Horizontal pod autoscaler
  • Observed CPU and memory, over time → Vertical pod autoscaler
  • Pods that cannot be scheduled → Karpenter
  • Horizontal pod autoscaler → Replica count
  • Vertical pod autoscaler → Resource requests
  • Karpenter → Node count and instance shape

Horizontal scaling on a request-rate or concurrency signal is what handles traffic. CPU is a poor proxy for load in most web workloads and reacts late. For queue consumers, scale on queue depth rather than on anything about the pod.

Vertical scaling is mostly useful as a measurement tool. We run VPA in recommendation mode and use what it learns to set honest resource requests, because requests that are set by guesswork are the single largest source of waste on a Kubernetes bill: every over-stated request reserves capacity that nothing uses and that Karpenter then faithfully provisions.

Deployment without a stored credential

A long-lived AWS access key in a CI secret is the most common serious finding in a platform review. It is valid until someone remembers to rotate it, it is readable by every workflow in the repository, and its use is nearly indistinguishable from legitimate traffic in CloudTrail.

OIDC removes the credential entirely. GitHub presents a signed token describing the repository, the branch and the workflow; AWS validates it against the registered identity provider and issues a short-lived role session.

Deployment path, no stored AWS credentials
Deployment from GitHub Actions to AWS using OIDC, with no stored credentialA commit starts two paths. On one, the image is built with a software bill of materials, scanned, signed and pushed to ECR. On the other, GitHub Actions presents an OIDC token describing the repository, branch and workflow, which AWS validates against a registered identity provider and exchanges at STS for a short-lived role session. The deploy is where the two meet, and an admission controller verifies the signature before the image runs.THE ARTEFACTTHE IDENTITYpushpresentexchangepullassumeverifyCommit on a protected branchBuild, test, SBOMScan and signseverity gate blocksECRimage and signatureOIDC tokenrepository, branch, workflowIAM identity providertrust policy in TerraformRole sessionminutes, not monthsApply to the clusterthe two paths meet hereAdmission controlan unsigned image does not run
Stored keys
none, anywhere in CI
Trust pinned to
repository and branch or environment
Reviewed as
Terraform, not a settings page
This diagram as text
  • Commit on a protected branch
  • The artefact
    • Build, test, SBOM
    • Scan and sign — severity gate blocks
    • ECR — image and signature
  • The identity
    • OIDC token — repository, branch, workflow
    • IAM identity provider — trust policy in Terraform
    • Role session — minutes, not months
  • Apply to the cluster — the two paths meet here
  • Admission control — an unsigned image does not run

Relationships

  • Commit on a protected branch → Build, test, SBOM
  • Build, test, SBOM → Scan and sign
  • Scan and sign → ECR — push
  • Commit on a protected branch → OIDC token
  • OIDC token → IAM identity provider — present
  • IAM identity provider → Role session — exchange
  • ECR → Apply to the cluster — pull
  • Role session → Apply to the cluster — assume
  • Apply to the cluster → Admission control — verify

The trust policy is the control. It is pinned to the repository and the branch or environment, so a pull request from a fork cannot assume the production role. It lives in Terraform, which means it is reviewed like code and its history is readable, rather than being a setting somebody changed once.

Inside the cluster, workloads get their AWS permissions through IAM Roles for Service Accounts or EKS Pod Identity. Neither the node role nor a shared secret is the identity, so a compromised pod has the permissions of that service and nothing more.

Secure SDLC, as pipeline stages rather than policy

The controls are only real if they can fail the build:

  1. Dependency and image scanning on every build, with a documented severity threshold that blocks.
  2. A software bill of materials generated and retained per release, so the question “are we exposed to this CVE” is answerable in minutes rather than in a day of archaeology.
  3. Image signing, verified by an admission controller. An unsigned image does not run, which is what makes the registry a control rather than a convenience.
  4. Infrastructure-as-code scanning before apply, because a permissive security group reaches production exactly as fast as application code does.
  5. No direct human access to production. Changes arrive through the pipeline. Break-glass access exists, is separately credentialed, and is alerted on.

Each of these produces an artefact with a timestamp, which is the same evidence an ISO 27001 or SOC 2 auditor asks for. That is deliberate: a pipeline built this way is also the compliance evidence, and nobody has to assemble a screenshot pack at audit time.

The Well-Architected pillars as acceptance criteria

A design review that only asks “is it up” accepts a platform that is expensive, unobservable and impossible to change safely. We accept against all six pillars, with a specific question for each.

Pillar The question we must be able to answer
Operational excellence Can any engineer deploy and roll back without a person who holds special knowledge? Is every change traceable to a commit?
Security Are there any long-lived credentials? Is every data store encrypted and is the key’s lifecycle owned? Is the blast radius of one compromised pod bounded?
Reliability Has a zone loss been exercised, not assumed? Does the restore meet the stated RTO when timed?
Performance efficiency Are resource requests derived from measurement? Is the scaling signal the thing that actually indicates load?
Cost optimisation What is the cost per environment and per service, and who sees it weekly? What is running at 3am that need not be?
Sustainability Is idle capacity being reclaimed? Are ARM instances used where the image supports them?

The sixth pillar is not decoration. Consolidation, right-sized requests and Graviton adoption improve the cost and the carbon number through exactly the same mechanism, which makes sustainability the easiest pillar to satisfy once the others are done properly.

When a second region earns its keep

Multi-region is a different class of system, not a bigger version of the same one. It introduces data consistency decisions that multi-AZ does not have, and it roughly doubles the surface that has to be patched, monitored and tested.

It is justified when the recovery time objective is shorter than a rebuild would take, when a regulator requires geographic separation, or when latency to a distant user base is a product requirement rather than a preference.

Two-region posture, warm standby
Two-region warm standby on AWS with health-checked DNS failoverRoute 53 health checks direct traffic to the primary region, which runs at full capacity and takes all writes. The secondary region runs reduced capacity from the same Terraform modules with a read-only replica that is promoted on failover. Backups are written to both regions.AWSprimary regionAWSsecondary regionactiveon failureasyncRoute 53health-checked failover routingFull capacityserving all trafficPrimary databaseaccepts writesBackupswritten to both regionsReduced capacitysame modules, different variablesReplicaread-only until promotedBackup copyrestorable independently
Built from
the same modules, different variables
Routing
health-checked DNS failover
Proven by
a scheduled, timed failover exercise
Otherwise
a standby that never took traffic is not a capability
This diagram as text
  • Route 53 — health-checked failover routing
  • AWS primary region
    • Full capacity — serving all traffic
    • Primary database — accepts writes
    • Backups — written to both regions
  • AWS secondary region
    • Reduced capacity — same modules, different variables
    • Replica — read-only until promoted
    • Backup copy — restorable independently

Relationships

  • Route 53 → Full capacity — active
  • Route 53 → Reduced capacity — on failure
  • Primary database → Replica — async
  • Backups → Backup copy

The rule we hold to: a standby region that has never taken traffic is not a disaster recovery capability, it is a second estate to patch. If it is built, it gets exercised on a schedule, and the timing is recorded against the stated objective.

How STP approaches this

We design to the agreed RTO and RPO first and select services second, build the whole estate in Terraform so both regions and every lower environment come from the same modules, wire deployment through OIDC so no AWS key exists in a CI system, and hand over the Well-Architected review as a written document rather than as a verbal assurance. Where a platform already exists, the same review is how we start: it finds the stored credentials, the unbounded requests and the untested restore before it proposes anything.

More on Kubernetes platform engineering and cloud, hybrid and multi-cloud, or start a conversation.