Terminology

Infrastructure glossary

79 terms used across our technical writing and in proposals, defined once and linked to the articles that explain them properly.

Cabling and the room

What gets installed in the building, and the standards it is tested against.

As-built documentation

As-built documentation is the record of what was actually installed rather than what was designed: port maps, floor plans, cable schedules and the certified test results.

It is the deliverable that lets a different contractor work on the estate in ten years. Its absence is what makes an estate expensive to change.

Cat6A

Cat6A is a twisted pair cabling category rated for 10 Gbit/s over the full 100-metre channel, with tighter alien crosstalk limits than Cat6.

The practical reason to specify it is lifespan: the cable is the part of the estate nobody replaces for fifteen years, and every switch bought in that period inherits it.

Also Category 6A, Augmented Category 6

Certified testing

Certified testing measures every installed cable against the standard with a calibrated tester and produces a per-port pass or fail result that is handed over as a file.

A contractor who tested will give you the results per port. One who did not will offer a verbal assurance, which is the difference worth checking before you sign.

Also link certification, channel testing

Containment

Containment is the tray, basket, ladder and conduit that carries cable through a building, sized and installed before any cable is pulled.

It is set out first because retrofitting a route into an occupied building costs several times what installing it during fit-out does.

Also cable tray, basket tray, ladder rack

ISO/IEC 11801

ISO/IEC 11801 is the international standard for generic customer premises cabling, and is the European counterpart to the TIA-568 series.

OM4 and OS2 fibre

OM4 is multi-mode optical fibre for short reaches inside a building, and OS2 is single-mode fibre for long reaches between buildings or across a campus.

The choice is set by distance and by the optics you will buy, not by preference. Multi-mode optics are cheaper; single-mode reach is far longer.

Also multi-mode fibre, single-mode fibre

Raised access floor

A raised access floor is a levelled grid of pedestals and removable panels that creates a service void under a technical room for cabling, power or cooling air.

TIA-568

TIA-568 is the North American standard series defining structured cabling topology, pin assignments and the test limits a certified installation is measured against.

Networking

How packets move, and where one failure is stopped from becoming two.

802.1X

802.1X is the standard for port-based network access control, requiring a device to authenticate before the switch port it is plugged into carries traffic.

BGP

BGP is the routing protocol that exchanges reachability between independent networks, and gives explicit control over which routes are advertised and accepted.

That control is why it sits at a site boundary: you can fail a site out by withdrawing its routes rather than by unplugging something.

Also Border Gateway Protocol

CDN

A CDN serves cacheable responses from locations close to the user and terminates TLS at the edge, so the origin handles only what genuinely has to reach it.

Also Content Delivery Network, CloudFront, Cloud CDN

EVPN

EVPN is a BGP control plane for carrying Ethernet reachability across a routed network, so MAC addresses are learned by signalling rather than by flooding.

It is what makes a layer-2 adjacency between sites controllable instead of turning two buildings into one broadcast domain.

Also Ethernet VPN, RFC 7432

Load balancer

A load balancer distributes incoming requests across healthy instances of a service and stops sending traffic to one that fails its health check.

Also ALB, Application Load Balancer

OSPF

OSPF is a link-state interior routing protocol that distributes topology within one administrative domain and converges on the shortest path.

Also Open Shortest Path First

SD-WAN

SD-WAN steers traffic across more than one wide area link per application, failing over between them without dropping sessions and reporting what each link actually delivered.

It earns its cost where branches have two links of different kinds, or where voice and conferencing need brownouts routed around rather than reported.

Segmentation

Segmentation divides a network into zones with policy enforced between them, so a compromise in one zone does not reach the rest of the estate.

It is cheapest designed before cable is pulled and most expensive retrofitted into an occupied building.

Also network segmentation, VLAN

Spine and leaf

Spine and leaf is a data centre fabric topology in which every leaf switch connects to every spine switch, so any two endpoints are the same number of hops apart.

TTL

A DNS record time-to-live is how long a resolver may cache that record, which sets how quickly a change to it takes effect everywhere.

Lowering it several days before a planned change turns DNS propagation from a risk into a non-event, and it costs nothing.

Also time to live

Underlay and overlay

The underlay is the physical routed network that carries packets, and the overlay is the virtual network built on top of it that carries tenant traffic.

Keeping the underlay simple and stable is what lets the overlay change often without risking the fabric.

VXLAN

VXLAN is an encapsulation that carries layer-2 frames inside UDP across a routed underlay, allowing the same subnet to exist in more than one place.

Platform and cloud

Where workloads actually run, and what schedules them.

Amazon ECS

Amazon ECS is AWS container orchestration without Kubernetes: less to operate than a cluster, and proprietary, so it runs only on AWS.

The right choice when staying on AWS is a settled business decision rather than an assumption, because moving off it later is a rewrite rather than a migration.

Also Elastic Container Service

Amazon EKS

Amazon EKS is managed Kubernetes on AWS, where AWS operates the control plane and the workloads stay portable because Kubernetes is the same substrate on another cloud or on your own hardware.

The right choice where the business holds to not locking itself into proprietary technology. The portability is only real if the infrastructure code can provision the platform elsewhere and somebody has proved it.

Also Elastic Kubernetes Service

Availability Zone

An Availability Zone is one isolated datacentre within a cloud region, with its own power and cooling, so a failure in one does not take the others with it.

Three is the usual default, because it is the point at which a quorum-based datastore can lose one and keep writing.

Also AZ

Container

A container packages an application with its dependencies into an image that runs the same way on any host with a compatible runtime.

Control plane

The control plane is the set of components that decide what should be running and where, as distinct from the worker nodes that actually run it.

Losing it usually means keeping your workloads but losing the ability to change them, which during an incident is close to the same thing.

Fargate

Fargate runs containers without you provisioning or patching the underlying instances, billed per task rather than per machine.

Suits bursty jobs and per-pod isolation, and costs more per unit of compute than nodes you manage. Available under both ECS and EKS, so it is a capacity choice rather than a lock-in one.

GKE

GKE is managed Kubernetes on Google Cloud, offered as Standard where you manage node pools and Autopilot where Google does.

Also Google Kubernetes Engine

Helm

Helm packages a set of Kubernetes manifests as a versioned chart with values that vary per environment.

Horizontal Pod Autoscaler

The Horizontal Pod Autoscaler adds or removes pod replicas in response to a metric, which is the mechanism that absorbs changes in traffic.

Scale on request rate or concurrency rather than CPU, which is a poor proxy for load in most web workloads and reacts late.

Also HPA

Hyperconverged

Hyperconverged infrastructure places storage in the same hosts that run compute, removing the separate array and SAN fabric.

The trade is that a host failure removes compute and a copy of the data at once, and the rebuild competes with production for the same network.

Hypervisor

A hypervisor runs virtual machines on physical hosts and can move them between hosts while they are running.

Also vSphere, Hyper-V

Karpenter

Karpenter provisions cloud instances shaped to the pods that are waiting to be scheduled, and consolidates workloads onto fewer nodes as demand falls.

The consolidation is the part that produces the saving. Provisioning alone mostly changes how nodes appear.

Kubernetes

Kubernetes is a container orchestrator that schedules workloads across a pool of machines, restarts what fails, and reconciles the running estate towards a declared state.

Also K8s

Live migration

Live migration moves a running virtual machine from one host to another without stopping it, which is what makes patching possible without a maintenance window.

It only works if the cluster has the spare capacity to run on one fewer host for the duration.

Also vMotion

Pod

A pod is the smallest deployable unit in Kubernetes: one or more containers that share a network namespace and are scheduled onto a node together.

Region

A cloud region is a geographic grouping of Availability Zones, and is the boundary that data residency and most regulatory questions are asked about.

Resource requests

A resource request is the CPU and memory a pod reserves on a node, and it is what the scheduler uses to decide whether the pod fits.

Requests set by guesswork are the largest single source of waste on a Kubernetes bill, because every overstated request reserves capacity nothing uses.

Vendor lock-in

Vendor lock-in is the cost of leaving a provider, and it accrues with every proprietary service an application is written against rather than appearing at the moment of the decision.

It is a business decision rather than a technical one. Accepting it is reasonable where staying is settled; the mistake is accepting it by default and discovering the bill years later.

Also exit cost, portability

Vertical Pod Autoscaler

The Vertical Pod Autoscaler observes what a workload actually uses and recommends or applies revised CPU and memory requests for it.

Most useful in recommendation mode, as a measurement tool for setting honest requests.

Also VPA

Delivery and automation

How a change gets from a commit to production.

CI/CD

CI/CD is the automated path a change takes from commit to production, running the checks that can block it and leaving a record of every step.

Also continuous integration, continuous delivery

Ephemeral environment

An ephemeral environment is a complete environment created for one pull request from the same modules as production and destroyed when it merges.

It cannot drift from production the way a long-lived staging environment does, because it does not live long enough to.

Also preview environment, PR environment

GitOps

GitOps makes a Git repository the declared state of an environment, with a controller continuously reconciling the running estate towards it.

Infrastructure as code

Infrastructure as code defines infrastructure in version-controlled files so that an environment is created by applying them rather than by someone repeating a sequence of clicks.

It is also what makes a second region or a recovery site the same system as production rather than a similar one.

Also IaC, Terraform

Running it

What you watch, what you keep, and what wakes someone up.

Cardinality

Cardinality is the number of distinct time series a metric produces, which multiplies with every label value it carries.

A user id or request id used as a metric label is the failure that takes a metrics backend down, and it is written in the application.

Collector

A collector sits between applications and observability backends, and is where data is filtered, sampled, redacted and routed.

It is the single place a backend can be changed, duplicated during a migration, or split by signal.

Latency, errors, saturation, traffic

Latency, error rate, saturation and traffic are the four symptoms worth alerting on, because they describe what a user experiences rather than what a component is doing.

Measure latency at a percentile rather than a mean, which hides the tail that people actually complain about.

Also golden signals

OpenTelemetry

OpenTelemetry is a vendor-neutral standard for emitting traces, metrics and logs, so instrumentation does not belong to whichever backend you chose first.

Instrument once with it and changing backend becomes a collector configuration change rather than a change to every service.

Also OTel, OTLP

Retention

Retention is how long observability data stays searchable, and it is the setting that most often explains a surprising bill.

Thirty to ninety days hot covers nearly all operational need; the longer regulatory obligation is better served by a cheap cold archive.

Tail-based sampling

Tail-based sampling decides whether to keep a trace after it finishes, so every error and every slow request is retained while routine traffic is sampled down.

Head-based sampling decides at the start and therefore discards the interesting traces at random.

Resilience and recovery

What the estate survives, and how that is proven rather than asserted.

Admission control

Cluster admission control refuses to accept workloads the cluster could not restart after the failure it claims to tolerate.

Disabled, a cluster will happily accept more than it can recover, and nobody finds out until a host fails.

Failback

Failback is the planned return of service to the primary site after a failover, and it needs its own runbook and its own rehearsal.

Teams plan the failover and discover during the incident that nobody decided how to return.

Failover

Failover is the transfer of service from a failed component or site to a standby one, either automatically or on a declared decision.

A standby that has never taken traffic is not a recovery capability, it is a second estate to patch.

Isolated restore

An isolated restore rebuilds a system from the real backup into a network with no route to production, proving the backup, the runbook and the timing without touching live service.

It is the highest-value recovery exercise available and can be run quarterly without a maintenance window.

N+1

N+1 means the estate has enough spare capacity to keep running after losing one component, sized against the largest one rather than the average.

The test is arithmetic: committed memory must fit on the remaining hosts, with no single virtual machine larger than one surviving host can take.

Quorum

Quorum is the majority of votes a cluster needs to agree it is allowed to run, which is what stops two halves of a split cluster both believing they are authoritative.

Replication lag

Replication lag is how far behind a replica is from its primary, and it is the number that determines your real recovery point.

It drifts quietly, so it should be graphed and alerted on rather than assumed.

RPO

The Recovery Point Objective is how much data an organisation can afford to lose, measured as a period of time rather than as a volume.

It drives backup and replication frequency. Whatever the plan says, your real RPO is your measured replication lag.

Also Recovery Point Objective

RTO

The Recovery Time Objective is how long a system may be unavailable before the impact on the business is unacceptable.

It drives the recovery architecture. It is a business decision, and deriving it from what the current infrastructure can already do is the most common way to get it wrong.

Also Recovery Time Objective

Runbook

A runbook is the written procedure for executing a recovery or an operational task, written for someone who did not design the system and may have been woken up.

The honest test is to hand it to a competent engineer who has never seen the system and have them execute it while the author stays silent.

Warm standby

A warm standby is a reduced-capacity copy of production kept running and current, which is scaled up when it has to take over.

Witness

A witness is the extra vote that lets a cluster establish quorum, and it must not share a failure domain with either side it is arbitrating between.

A witness in the same rack as one node is not a witness. It is a third vote for that rack.

Also quorum witness, cloud witness

Identity and security

Who is allowed to do what, and what is verified rather than trusted.

Admission controller

An admission controller evaluates every request to the Kubernetes API before it is persisted, and can reject one that violates policy.

An admission policy rejecting any image not from the internal registry is how an air-gapped estate discovers its undocumented dependencies while it still has a network.

Air gap

An air-gapped system has no network route to or from the internet in either direction, so everything it will ever need must be inside the boundary before it needs it.

It is not a private network with a strict firewall. The usual remedy of pulling a fixed image does not exist.

Break-glass access

Break-glass access is a separately credentialed route into production used only during an incident, which is alerted on whenever it is used.

It fails in practice when the credential lives in a vault that depends on the system that is down.

Data obfuscation

Data obfuscation replaces personal data with generated values of the same shape, so a lower environment can carry production-shaped data lawfully.

The order is the whole control: it must happen inside the production boundary, before export. Copying first and anonymising afterwards is the breach it was meant to prevent.

Also masking, anonymisation

Image signing

Image signing attaches a cryptographic signature to a container image so an admission controller can refuse to run anything unsigned.

It is what turns a registry from a convenience into a control.

Internal certificate authority

An internal certificate authority issues and renews certificates inside a boundary that cannot reach a public one, with its own lifecycle to own.

Manual renewal inside an air gap is how a cluster expires on a weekend.

Least privilege

Least privilege grants an identity only the permissions its job requires, so a compromise reaches only what that one identity could reach.

OIDC

OIDC lets a CI system prove its identity to a cloud provider with a short-lived token instead of a stored access key, and exchange it for a scoped, expiring session.

The credential cannot be exfiltrated from the CI system because it does not persist there, and the trust policy is reviewable in code.

Also OpenID Connect, workload identity federation

SAST and DAST

SAST analyses source code for vulnerabilities without running it, and DAST tests a running application by sending it hostile input.

SBOM

A software bill of materials lists every component and version inside a build, so the question "are we exposed to this vulnerability" is answerable in minutes.

Also software bill of materials

WAF

A web application firewall inspects HTTP requests before they reach an application and blocks those matching known attack patterns or rate limits.

Also Web Application Firewall, Cloud Armor

WireGuard

WireGuard is a VPN protocol built on modern cryptography with a deliberately small codebase, used here for administrator and engineer remote access.

Workload identity

Workload identity gives a running workload its own cloud identity, so it calls provider APIs as itself rather than inheriting the permissions of the machine it runs on.

Also IRSA, Pod Identity

Standards and compliance

The frameworks an auditor samples against.

Audit evidence

Audit evidence is the dated record showing a control operated over a period, which is what an auditor samples rather than a description of the control.

Infrastructure built as code emits most of it as a by-product of normal operation, which is why it costs less at audit time.

AWS Well-Architected Framework

The AWS Well-Architected Framework is a review structure with six pillars: operational excellence, security, reliability, performance efficiency, cost optimisation and sustainability.

Used as acceptance criteria rather than as a document, it stops a design review from only asking whether the system is up.

ISO 22301

ISO 22301 is the international standard for business continuity management, which treats exercising recovery as an operating requirement rather than a documentation one.

ISO/IEC 27001

ISO/IEC 27001 is the international standard for an information security management system, certified by audit against documented controls that are shown to have operated over time.

An auditor samples evidence that controls operated, not screenshots proving they exist.

SOC 2

SOC 2 is an attestation report on a service organisation’s controls over security and related criteria, issued by an auditor for a defined period.