A cloud bill that grows faster than traffic is rarely a scaling problem. In almost every estate we review, the largest single line is capacity that was reserved and never used, and no autoscaling policy can reclaim it because the scheduler is doing exactly what it was asked.
Requests, not usage
A Kubernetes scheduler places pods by their requests. It does not place them by what they actually use. A pod that requests two cores and uses a tenth of one has reserved two cores, and the node it sits on is full as far as the scheduler is concerned.
Multiply that across an estate and you get the characteristic picture: nodes at ninety per cent committed and twenty per cent utilised, a cluster autoscaler dutifully adding more, and a bill that tracks requests rather than work.
Autoscaling on top of guessed requests scales the guess.
This is why request hygiene comes before any scaling policy. It is also why the cost conversation has to involve the people writing the application, because requests are written in the deployment manifest and nobody on the infrastructure side can set them honestly.
Measure first
Run a vertical autoscaler in recommendation mode, or read the percentiles out of whatever observability you already have, and compare actual usage to what each workload requests. Two numbers matter per workload: the ninety-fifth percentile of CPU and the maximum of memory.
Set CPU requests near the ninety-fifth percentile. CPU is compressible — a pod that exceeds it is throttled, which is survivable. Set memory requests at or slightly above the observed maximum, because memory is not compressible and exceeding it ends the process.
- Biggest line
- reserved and unused, nearly always
- Commit last
- discounting a wrong baseline locks it in
- Owned by
- whoever writes the manifest
This diagram as text
- 1 · Switch off non-production out of hours — no architecture change, no risk
- 2 · Requests from measurement — p95 CPU, max memory
- 3 · Let a provisioner pick the instance — across families and sizes
- 4 · ARM where the image supports it
- 5 · Spot for anything interruptible — with a disruption budget
- 6 · Committed use discounts — only once the shape has settled
Relationships
- 1 · Switch off non-production out of hours → 2 · Requests from measurement — buys time
- 1 · Switch off non-production out of hours → 3 · Let a provisioner pick the instance
- 2 · Requests from measurement → 4 · ARM where the image supports it
- 3 · Let a provisioner pick the instance → 5 · Spot for anything interruptible
- 4 · ARM where the image supports it → 6 · Committed use discounts
- 5 · Spot for anything interruptible → 6 · Committed use discounts
Scale on the signal that means load
The second common failure is scaling on CPU because it is the default.
For a request-driven service, CPU is a lagging and noisy proxy. By the time it rises enough to trigger a policy, latency has already degraded, and garbage collection or a neighbour on the same node can move it for reasons unrelated to traffic. Scale on requests per second, on concurrency, or on queue depth for a consumer. Those are the things that actually describe the work arriving.
For a queue consumer in particular, scale on the depth of the queue and nothing else. It is the only signal that says work is accumulating faster than it is being done.
Horizontal and vertical do not combine
Running a horizontal autoscaler and a vertical autoscaler on the same deployment against the same metric produces oscillation: one adds replicas, the other grows them, each reacting to the other’s effect.
The arrangement that works is vertical in recommendation mode to set honest requests, and horizontal in control of the replica count. One of them measures, the other acts.
The levers in order
Switch things off. Non-production environments outside working hours, ephemeral environments reaped on merge, and anything running at three in the morning that nobody asked for. This is the cheapest money in any estate and it needs no architectural change.
Fix requests. As above. This is where the largest sustained saving is, and it is slow because it is per-workload and needs application teams.
Let a provisioner choose the instance. Constraining node provisioning to a single instance family removes most of its ability to find capacity cheaply. Constrain on what the workload genuinely needs and let it choose across families and sizes.
ARM where the image supports it. Better price-performance for the same work. The effort is in the build pipeline producing a multi-architecture image and in actually testing the ARM build, not in the platform.
Spot for anything interruptible, with a disruption budget and graceful termination genuinely implemented rather than assumed.
Commit last. A committed use discount applied to a baseline you have not yet corrected locks in the waste for one to three years. Do the four steps above, let the shape settle, then commit to what remains.
Make the number visible to the people who can move it
A bill that arrives monthly to one person in finance changes nothing. Cost per service and per environment, visible weekly to the teams that own them, changes behaviour without anybody having to run a programme.
The attribution is the work: tags or labels applied at creation, enforced by the pipeline rather than by a policy document, so an untagged resource cannot reach production. Retrofitting attribution onto an estate that has none is several times harder than applying it from the start.
How STP approaches this
We measure before changing anything, because the first instinct is usually to add a scaling policy to a cluster whose requests are wrong. Non-production schedules go in first since they are free and buy time for the slower work. Commitments come last, after the shape has settled, and attribution is enforced in the pipeline so the number stays answerable as the estate grows.
More on cloud, hybrid and multi-cloud and Kubernetes platform engineering, or start a conversation.

