A virtualisation cluster is bought for one reason: a host can fail without the business noticing. That property is not supplied by the hypervisor licence. It is supplied by how the cluster is sized, where quorum lives, and whether anyone has ever removed a host on purpose to find out.

Size against the largest host, not the average

The common error is capacity planning on averages. A cluster running at 60% across six hosts looks comfortable until the largest host fails at month-end and the remaining five cannot admit its virtual machines, because admission control is evaluated against reservations, not against yesterday’s graph.

The arithmetic that has to hold:

  • Total committed memory fits on the cluster minus one host, with headroom. Memory is almost always the binding constraint, not CPU.
  • No single virtual machine is larger than the free capacity of one host. If one is, it cannot be evacuated, which means it also blocks patching.
  • Admission control is enabled and set to the failure tolerance you actually claim. A cluster with admission control disabled will happily accept workloads it cannot restart.
N+1 cluster, host loss absorbed
An N+1 virtualisation cluster in normal operation and after losing a hostThree hosts each running at sixty per cent in normal operation, with the remaining capacity held as spare rather than used. When one host fails the cluster restarts its virtual machines on the survivors, which then run at ninety per cent: degraded, and still inside tolerance. The binding constraint is almost always memory.NORMAL OPERATIONDEGRADED, STILL INSIDE TOLERANCErestartedHost 160%Host 260%Host 360%Held spareone host worthHost 190%Host 2failedHost 390%
Tolerates
one host, at peak, not at average
Binding constraint
memory, nearly always
Admission control
enabled, and matched to the claim
This diagram as text
  • Normal operation
    • Host 1 — 60%
    • Host 2 — 60%
    • Host 3 — 60%
    • Held spare — one host worth
  • Degraded, still inside tolerance
    • Host 1 — 90%
    • Host 2 — failed
    • Host 3 — 90%

Relationships

  • Host 2 (Normal operation) → Host 1 (Degraded, still inside tolerance) — restarted
  • Host 2 (Normal operation) → Host 3 (Degraded, still inside tolerance)

Quorum, and where the witness goes

A cluster decides whether it is allowed to run by counting votes. Getting this wrong produces either a cluster that stops when it did not need to, or two halves that both believe they are authoritative and corrupt shared storage between them.

The rule: the witness must not share a failure domain with either side. A file-share witness on a server in the same rack as node one is not a witness, it is a third vote for rack one. In a two-site design the witness belongs in a third location — another building, or a cloud witness — so that losing either site leaves the other with a majority.

Design Where the witness goes What it survives
Single room, three or more hosts Node majority, no separate witness needed Any one host
Two racks, one room Witness in a third rack on separate power A rack, including its power feed
Two buildings Cloud witness, or a third site A building
Two datacentres, stretched Third site only. Never inside either datacentre A datacentre

VMware and Hyper-V do the same job differently

Both give you live migration, automated restart on host failure and rolling patching. The decision is usually about what the organisation already runs and already licenses, not about capability.

vSphere Hyper-V with Failover Clustering
Shared storage VMFS or vVols on FC/iSCSI, or vSAN CSV on FC/iSCSI/SMB3, or Storage Spaces Direct
Live move vMotion Live Migration
Restart on failure vSphere HA Failover Clustering
Load balancing DRS, automatic Manual, or System Center
Fits best when Mixed workloads, existing VMware skills, demanding storage A Windows-heavy estate where the licensing is already owned

Where a Windows Server estate is already licensed, Hyper-V frequently costs less in total and is entirely adequate. Where the estate is mixed, the storage profile is demanding, or the team already operates vSphere, VMware’s automation does more of the work for you. Neither choice rescues a cluster that is sized wrong.

Storage: separate the failure domains, or accept that you have not

With a shared array, a host failure and a storage failure are different events with different recoveries. With hyperconverged storage, the disks live in the hosts, so losing a host removes compute and a copy of the data, and a rebuild competes with production for the same network.

Where we use a distributed file system — GlusterFS for replicated bulk storage, or object storage for backup targets and artefact mirrors — the same quorum logic applies as to the cluster itself. A two-replica volume without an arbiter cannot decide which copy is authoritative after a split, which is the file-system version of the same mistake.

Whatever the layout, two things are specified at build time rather than discovered later: the rebuild time after a disk loss at the array’s real capacity, and whether that rebuild runs over a network that production is also using.

Patching without a maintenance window

This is the day-to-day payoff and the thing most clusters cannot actually do.

Rolling patch cycle
Rolling patch cycle across a virtualisation cluster, one host at a timeA host enters maintenance mode, its workloads live-migrate off with no downtime, and it is patched and rebooted. It rejoins, workloads rebalance, and cluster health is verified before the next host starts. Never two at once.Host enters maintenance modeWorkloads live-migrate offno downtimePatch and rebootHost rejoinsWorkloads rebalanceCluster health verifiedbefore the next hostNext hostnever in parallel
This diagram as text
  • Host enters maintenance mode
  • Workloads live-migrate off — no downtime
  • Patch and reboot
  • Host rejoins
  • Workloads rebalance
  • Cluster health verified — before the next host
  • Next host — never in parallel

Relationships

  • Host enters maintenance mode → Workloads live-migrate off
  • Workloads live-migrate off → Patch and reboot
  • Patch and reboot → Host rejoins
  • Host rejoins → Workloads rebalance
  • Workloads rebalance → Cluster health verified
  • Cluster health verified → Next host

It works only when the capacity arithmetic above holds, because entering maintenance mode means running the whole cluster on N-1 for the duration. A cluster that is too full to lose a host is also too full to patch, which is how estates end up years behind on firmware and hypervisor updates with a documented reason that was never true.

Firmware and drivers are part of this cycle, not separate from it. On HPE hardware we hold hosts to a single validated recipe — server firmware, adapter firmware and hypervisor build moving together — and drive it from Ansible so that what is actually installed is a fact in version control rather than a thing someone remembers doing.

What we hand over

  • The capacity model, with the arithmetic shown, so a future purchase can be argued from it.
  • The quorum design and why the witness sits where it does.
  • A runbook for host failure, host replacement and the rolling patch cycle, written to be executed by an engineer who did not build the cluster.
  • The result of an actual host-loss exercise, with timings, performed before handover rather than described in a document.

That last point is the one that separates a cluster that is highly available from one that is only described that way.

How STP approaches this

We size from the real peak and the largest host, place quorum outside both failure domains, drive firmware and configuration from Ansible so drift is visible, and pull a host before handover so the first failure is one we chose. Where a cluster already exists, the same exercise is the audit: it finds the disabled admission control, the witness in the wrong rack and the virtual machine too large to evacuate.

More on data centre build and relocation and private cloud, or start a conversation.