Terminology
Infrastructure glossary
79 terms used across our technical writing and in proposals, defined once and linked to the articles that explain them properly.
Cabling and the room
What gets installed in the building, and the standards it is tested against.
- As-built documentation
As-built documentation is the record of what was actually installed rather than what was designed: port maps, floor plans, cable schedules and the certified test results.
It is the deliverable that lets a different contractor work on the estate in ten years. Its absence is what makes an estate expensive to change.
Explained inChoosing a cabling contractor
Related servicesStructured Cabling SystemsData Centre Design & Build
- Cat6A
Cat6A is a twisted pair cabling category rated for 10 Gbit/s over the full 100-metre channel, with tighter alien crosstalk limits than Cat6.
The practical reason to specify it is lifespan: the cable is the part of the estate nobody replaces for fifteen years, and every switch bought in that period inherits it.
Also Category 6A, Augmented Category 6
Explained inChoosing a cabling contractor
Related servicesStructured Cabling Systems
- Certified testing
Certified testing measures every installed cable against the standard with a calibrated tester and produces a per-port pass or fail result that is handed over as a file.
A contractor who tested will give you the results per port. One who did not will offer a verbal assurance, which is the difference worth checking before you sign.
Also link certification, channel testing
Explained inChoosing a cabling contractor
Related servicesStructured Cabling Systems
- Containment
Containment is the tray, basket, ladder and conduit that carries cable through a building, sized and installed before any cable is pulled.
It is set out first because retrofitting a route into an occupied building costs several times what installing it during fit-out does.
Also cable tray, basket tray, ladder rack
Related servicesStructured Cabling SystemsData Centre Design & Build
- ISO/IEC 11801
ISO/IEC 11801 is the international standard for generic customer premises cabling, and is the European counterpart to the TIA-568 series.
Related servicesStructured Cabling Systems
- OM4 and OS2 fibre
OM4 is multi-mode optical fibre for short reaches inside a building, and OS2 is single-mode fibre for long reaches between buildings or across a campus.
The choice is set by distance and by the optics you will buy, not by preference. Multi-mode optics are cheaper; single-mode reach is far longer.
Also multi-mode fibre, single-mode fibre
Related servicesStructured Cabling SystemsData Centre Design & Build
- Raised access floor
A raised access floor is a levelled grid of pedestals and removable panels that creates a service void under a technical room for cabling, power or cooling air.
Related servicesData Centre Design & Build
- TIA-568
TIA-568 is the North American standard series defining structured cabling topology, pin assignments and the test limits a certified installation is measured against.
Explained inChoosing a cabling contractor
Related servicesStructured Cabling Systems
Networking
How packets move, and where one failure is stopped from becoming two.
- 802.1X
802.1X is the standard for port-based network access control, requiring a device to authenticate before the switch port it is plugged into carries traffic.
Related servicesEnterprise & Data Centre Networking
- BGP
BGP is the routing protocol that exchanges reachability between independent networks, and gives explicit control over which routes are advertised and accepted.
That control is why it sits at a site boundary: you can fail a site out by withdrawing its routes rather than by unplugging something.
Also Border Gateway Protocol
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre Networking
- CDN
A CDN serves cacheable responses from locations close to the user and terminates TLS at the edge, so the origin handles only what genuinely has to reach it.
Also Content Delivery Network, CloudFront, Cloud CDN
Explained inProduction on AWSProduction on Google Cloud
Related servicesCloud, Hybrid & Multicloud
- EVPN
EVPN is a BGP control plane for carrying Ethernet reachability across a routed network, so MAC addresses are learned by signalling rather than by flooding.
It is what makes a layer-2 adjacency between sites controllable instead of turning two buildings into one broadcast domain.
Also Ethernet VPN, RFC 7432
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre NetworkingData Centre Design & Build
- Load balancer
A load balancer distributes incoming requests across healthy instances of a service and stops sending traffic to one that fails its health check.
Also ALB, Application Load Balancer
Explained inProduction on AWSProduction on Google Cloud
Related servicesEnterprise & Data Centre NetworkingKubernetes & Platform Engineering
- OSPF
OSPF is a link-state interior routing protocol that distributes topology within one administrative domain and converges on the shortest path.
Also Open Shortest Path First
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre Networking
- SD-WAN
SD-WAN steers traffic across more than one wide area link per application, failing over between them without dropping sessions and reporting what each link actually delivered.
It earns its cost where branches have two links of different kinds, or where voice and conferencing need brownouts routed around rather than reported.
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre Networking
- Segmentation
Segmentation divides a network into zones with policy enforced between them, so a compromise in one zone does not reach the rest of the estate.
It is cheapest designed before cable is pulled and most expensive retrofitted into an occupied building.
Also network segmentation, VLAN
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre Networking
- Spine and leaf
Spine and leaf is a data centre fabric topology in which every leaf switch connects to every spine switch, so any two endpoints are the same number of hops apart.
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre NetworkingData Centre Design & Build
- TTL
A DNS record time-to-live is how long a resolver may cache that record, which sets how quickly a change to it takes effect everywhere.
Lowering it several days before a planned change turns DNS propagation from a risk into a non-event, and it costs nothing.
Also time to live
Explained inHybrid Exchange
- Underlay and overlay
The underlay is the physical routed network that carries packets, and the overlay is the virtual network built on top of it that carries tenant traffic.
Keeping the underlay simple and stable is what lets the overlay change often without risking the fabric.
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre Networking
- VXLAN
VXLAN is an encapsulation that carries layer-2 frames inside UDP across a routed underlay, allowing the same subnet to exist in more than one place.
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre NetworkingData Centre Design & Build
Platform and cloud
Where workloads actually run, and what schedules them.
- Amazon ECS
Amazon ECS is AWS container orchestration without Kubernetes: less to operate than a cluster, and proprietary, so it runs only on AWS.
The right choice when staying on AWS is a settled business decision rather than an assumption, because moving off it later is a rewrite rather than a migration.
Also Elastic Container Service
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- Amazon EKS
Amazon EKS is managed Kubernetes on AWS, where AWS operates the control plane and the workloads stay portable because Kubernetes is the same substrate on another cloud or on your own hardware.
The right choice where the business holds to not locking itself into proprietary technology. The portability is only real if the infrastructure code can provision the platform elsewhere and somebody has proved it.
Also Elastic Kubernetes Service
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- Availability Zone
An Availability Zone is one isolated datacentre within a cloud region, with its own power and cooling, so a failure in one does not take the others with it.
Three is the usual default, because it is the point at which a quorum-based datastore can lose one and keep writing.
Also AZ
Explained inProduction on AWS
Related servicesCloud, Hybrid & Multicloud
- Container
A container packages an application with its dependencies into an image that runs the same way on any host with a compatible runtime.
Related servicesKubernetes & Platform Engineering
- Control plane
The control plane is the set of components that decide what should be running and where, as distinct from the worker nodes that actually run it.
Losing it usually means keeping your workloads but losing the ability to change them, which during an incident is close to the same thing.
Explained inProduction on Google Cloud
Related servicesKubernetes & Platform Engineering
- Fargate
Fargate runs containers without you provisioning or patching the underlying instances, billed per task rather than per machine.
Suits bursty jobs and per-pod isolation, and costs more per unit of compute than nodes you manage. Available under both ECS and EKS, so it is a capacity choice rather than a lock-in one.
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- GKE
GKE is managed Kubernetes on Google Cloud, offered as Standard where you manage node pools and Autopilot where Google does.
Also Google Kubernetes Engine
Explained inProduction on Google Cloud
Related servicesKubernetes & Platform Engineering
- Helm
Helm packages a set of Kubernetes manifests as a versioned chart with values that vary per environment.
Related servicesKubernetes & Platform Engineering
- Horizontal Pod Autoscaler
The Horizontal Pod Autoscaler adds or removes pod replicas in response to a metric, which is the mechanism that absorbs changes in traffic.
Scale on request rate or concurrency rather than CPU, which is a poor proxy for load in most web workloads and reacts late.
Also HPA
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- Hyperconverged
Hyperconverged infrastructure places storage in the same hosts that run compute, removing the separate array and SAN fabric.
The trade is that a host failure removes compute and a copy of the data at once, and the rebuild competes with production for the same network.
Explained inOn-premise clusters
Related servicesSovereign Private CloudData Centre Design & Build
- Hypervisor
A hypervisor runs virtual machines on physical hosts and can move them between hosts while they are running.
Also vSphere, Hyper-V
Explained inOn-premise clusters
Related servicesSovereign Private CloudData Centre Design & Build
- Karpenter
Karpenter provisions cloud instances shaped to the pods that are waiting to be scheduled, and consolidates workloads onto fewer nodes as demand falls.
The consolidation is the part that produces the saving. Provisioning alone mostly changes how nodes appear.
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- Kubernetes
Kubernetes is a container orchestrator that schedules workloads across a pool of machines, restarts what fails, and reconciles the running estate towards a declared state.
Also K8s
Explained inProduction on AWSProduction on Google CloudAir-gapped Kubernetes
Related servicesKubernetes & Platform Engineering
- Live migration
Live migration moves a running virtual machine from one host to another without stopping it, which is what makes patching possible without a maintenance window.
It only works if the cluster has the spare capacity to run on one fewer host for the duration.
Also vMotion
Explained inOn-premise clusters
Related servicesSovereign Private CloudData Centre Design & Build
- Pod
A pod is the smallest deployable unit in Kubernetes: one or more containers that share a network namespace and are scheduled onto a node together.
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- Region
A cloud region is a geographic grouping of Availability Zones, and is the boundary that data residency and most regulatory questions are asked about.
Explained inProduction on AWS
Related servicesCloud, Hybrid & Multicloud
- Resource requests
A resource request is the CPU and memory a pod reserves on a node, and it is what the scheduler uses to decide whether the pod fits.
Requests set by guesswork are the largest single source of waste on a Kubernetes bill, because every overstated request reserves capacity nothing uses.
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
- Vendor lock-in
Vendor lock-in is the cost of leaving a provider, and it accrues with every proprietary service an application is written against rather than appearing at the moment of the decision.
It is a business decision rather than a technical one. Accepting it is reasonable where staying is settled; the mistake is accepting it by default and discovering the bill years later.
Also exit cost, portability
Explained inProduction on AWSPrivate cloud vs public cloud
Related servicesCloud, Hybrid & MulticloudKubernetes & Platform Engineering
- Vertical Pod Autoscaler
The Vertical Pod Autoscaler observes what a workload actually uses and recommends or applies revised CPU and memory requests for it.
Most useful in recommendation mode, as a measurement tool for setting honest requests.
Also VPA
Explained inProduction on AWS
Related servicesKubernetes & Platform Engineering
Delivery and automation
How a change gets from a commit to production.
- CI/CD
CI/CD is the automated path a change takes from commit to production, running the checks that can block it and leaving a record of every step.
Also continuous integration, continuous delivery
Explained inProduction on AWS
Related servicesDevOps & AutomationSecure SDLC & Secure AI-DLC
- Ephemeral environment
An ephemeral environment is a complete environment created for one pull request from the same modules as production and destroyed when it merges.
It cannot drift from production the way a long-lived staging environment does, because it does not live long enough to.
Also preview environment, PR environment
Explained inProduction on Google Cloud
Related servicesDevOps & Automation
- GitOps
GitOps makes a Git repository the declared state of an environment, with a controller continuously reconciling the running estate towards it.
Explained inProduction on Google Cloud
Related servicesDevOps & Automation
- Infrastructure as code
Infrastructure as code defines infrastructure in version-controlled files so that an environment is created by applying them rather than by someone repeating a sequence of clicks.
It is also what makes a second region or a recovery site the same system as production rather than a similar one.
Also IaC, Terraform
Explained inProduction on AWSMulti-datacentre networking
Related servicesDevOps & Automation
Running it
What you watch, what you keep, and what wakes someone up.
- Cardinality
Cardinality is the number of distinct time series a metric produces, which multiplies with every label value it carries.
A user id or request id used as a metric label is the failure that takes a metrics backend down, and it is written in the application.
Explained inObservability without lock-in
- Collector
A collector sits between applications and observability backends, and is where data is filtered, sampled, redacted and routed.
It is the single place a backend can be changed, duplicated during a migration, or split by signal.
Explained inObservability without lock-in
- Latency, errors, saturation, traffic
Latency, error rate, saturation and traffic are the four symptoms worth alerting on, because they describe what a user experiences rather than what a component is doing.
Measure latency at a percentile rather than a mean, which hides the tail that people actually complain about.
Also golden signals
Explained inObservability without lock-in
Related servicesManaged IT Services
- OpenTelemetry
OpenTelemetry is a vendor-neutral standard for emitting traces, metrics and logs, so instrumentation does not belong to whichever backend you chose first.
Instrument once with it and changing backend becomes a collector configuration change rather than a change to every service.
Also OTel, OTLP
Explained inObservability without lock-in
Related servicesManaged IT ServicesDevOps & Automation
- Retention
Retention is how long observability data stays searchable, and it is the setting that most often explains a surprising bill.
Thirty to ninety days hot covers nearly all operational need; the longer regulatory obligation is better served by a cheap cold archive.
Explained inObservability without lock-in
- Tail-based sampling
Tail-based sampling decides whether to keep a trace after it finishes, so every error and every slow request is retained while routine traffic is sampled down.
Head-based sampling decides at the start and therefore discards the interesting traces at random.
Explained inObservability without lock-in
Resilience and recovery
What the estate survives, and how that is proven rather than asserted.
- Admission control
Cluster admission control refuses to accept workloads the cluster could not restart after the failure it claims to tolerate.
Disabled, a cluster will happily accept more than it can recover, and nobody finds out until a host fails.
Explained inOn-premise clusters
Related servicesSovereign Private Cloud
- Failback
Failback is the planned return of service to the primary site after a failover, and it needs its own runbook and its own rehearsal.
Teams plan the failover and discover during the incident that nobody decided how to return.
Explained inMulti-datacentre networking
Related servicesDisaster Recovery & Business Continuity
- Failover
Failover is the transfer of service from a failed component or site to a standby one, either automatically or on a declared decision.
A standby that has never taken traffic is not a recovery capability, it is a second estate to patch.
Explained inMulti-datacentre networkingProduction on AWS
Related servicesDisaster Recovery & Business Continuity
- Isolated restore
An isolated restore rebuilds a system from the real backup into a network with no route to production, proving the backup, the runbook and the timing without touching live service.
It is the highest-value recovery exercise available and can be run quarterly without a maintenance window.
Explained inWhat a real disaster recovery test looks like
Related servicesDisaster Recovery & Business Continuity
- N+1
N+1 means the estate has enough spare capacity to keep running after losing one component, sized against the largest one rather than the average.
The test is arithmetic: committed memory must fit on the remaining hosts, with no single virtual machine larger than one surviving host can take.
Explained inOn-premise clusters
Related servicesData Centre Design & BuildSovereign Private Cloud
- Quorum
Quorum is the majority of votes a cluster needs to agree it is allowed to run, which is what stops two halves of a split cluster both believing they are authoritative.
Explained inOn-premise clusters
Related servicesSovereign Private CloudData Centre Design & Build
- Replication lag
Replication lag is how far behind a replica is from its primary, and it is the number that determines your real recovery point.
It drifts quietly, so it should be graphed and alerted on rather than assumed.
Explained inMulti-datacentre networking
Related servicesDisaster Recovery & Business Continuity
- RPO
The Recovery Point Objective is how much data an organisation can afford to lose, measured as a period of time rather than as a volume.
It drives backup and replication frequency. Whatever the plan says, your real RPO is your measured replication lag.
Also Recovery Point Objective
Explained inWhat a real disaster recovery test looks like
Related servicesDisaster Recovery & Business Continuity
- RTO
The Recovery Time Objective is how long a system may be unavailable before the impact on the business is unacceptable.
It drives the recovery architecture. It is a business decision, and deriving it from what the current infrastructure can already do is the most common way to get it wrong.
Also Recovery Time Objective
Explained inWhat a real disaster recovery test looks likeMulti-datacentre networking
Related servicesDisaster Recovery & Business Continuity
- Runbook
A runbook is the written procedure for executing a recovery or an operational task, written for someone who did not design the system and may have been woken up.
The honest test is to hand it to a competent engineer who has never seen the system and have them execute it while the author stays silent.
Explained inWhat a real disaster recovery test looks like
Related servicesDisaster Recovery & Business ContinuityManaged IT Services
- Warm standby
A warm standby is a reduced-capacity copy of production kept running and current, which is scaled up when it has to take over.
Explained inProduction on AWS
Related servicesDisaster Recovery & Business Continuity
- Witness
A witness is the extra vote that lets a cluster establish quorum, and it must not share a failure domain with either side it is arbitrating between.
A witness in the same rack as one node is not a witness. It is a third vote for that rack.
Also quorum witness, cloud witness
Explained inOn-premise clusters
Related servicesSovereign Private CloudData Centre Design & Build
Identity and security
Who is allowed to do what, and what is verified rather than trusted.
- Admission controller
An admission controller evaluates every request to the Kubernetes API before it is persisted, and can reject one that violates policy.
An admission policy rejecting any image not from the internal registry is how an air-gapped estate discovers its undocumented dependencies while it still has a network.
Explained inAir-gapped Kubernetes
Related servicesKubernetes & Platform Engineering
- Air gap
An air-gapped system has no network route to or from the internet in either direction, so everything it will ever need must be inside the boundary before it needs it.
It is not a private network with a strict firewall. The usual remedy of pulling a fixed image does not exist.
Explained inAir-gapped Kubernetes
Related servicesSovereign Private CloudIT Audit & Compliance Enablement
- Break-glass access
Break-glass access is a separately credentialed route into production used only during an incident, which is alerted on whenever it is used.
It fails in practice when the credential lives in a vault that depends on the system that is down.
Explained inWhat a real disaster recovery test looks like
Related servicesManaged IT ServicesIT Audit & Compliance Enablement
- Data obfuscation
Data obfuscation replaces personal data with generated values of the same shape, so a lower environment can carry production-shaped data lawfully.
The order is the whole control: it must happen inside the production boundary, before export. Copying first and anonymising afterwards is the breach it was meant to prevent.
Also masking, anonymisation
Explained inProduction on Google Cloud
Related servicesSecure SDLC & Secure AI-DLCIT Audit & Compliance Enablement
- Image signing
Image signing attaches a cryptographic signature to a container image so an admission controller can refuse to run anything unsigned.
It is what turns a registry from a convenience into a control.
Explained inProduction on AWSAir-gapped Kubernetes
Related servicesSecure SDLC & Secure AI-DLCKubernetes & Platform Engineering
- Internal certificate authority
An internal certificate authority issues and renews certificates inside a boundary that cannot reach a public one, with its own lifecycle to own.
Manual renewal inside an air gap is how a cluster expires on a weekend.
Explained inAir-gapped Kubernetes
Related servicesSovereign Private Cloud
- Least privilege
Least privilege grants an identity only the permissions its job requires, so a compromise reaches only what that one identity could reach.
Related servicesIT Audit & Compliance EnablementSecure SDLC & Secure AI-DLC
- OIDC
OIDC lets a CI system prove its identity to a cloud provider with a short-lived token instead of a stored access key, and exchange it for a scoped, expiring session.
The credential cannot be exfiltrated from the CI system because it does not persist there, and the trust policy is reviewable in code.
Also OpenID Connect, workload identity federation
Explained inProduction on AWSProduction on Google Cloud
Related servicesDevOps & AutomationSecure SDLC & Secure AI-DLC
- SAST and DAST
SAST analyses source code for vulnerabilities without running it, and DAST tests a running application by sending it hostile input.
Related servicesSecure SDLC & Secure AI-DLC
- SBOM
A software bill of materials lists every component and version inside a build, so the question "are we exposed to this vulnerability" is answerable in minutes.
Also software bill of materials
Explained inProduction on AWS
Related servicesSecure SDLC & Secure AI-DLC
- WAF
A web application firewall inspects HTTP requests before they reach an application and blocks those matching known attack patterns or rate limits.
Also Web Application Firewall, Cloud Armor
Explained inProduction on AWS
Related servicesSecure SDLC & Secure AI-DLCEnterprise & Data Centre Networking
- WireGuard
WireGuard is a VPN protocol built on modern cryptography with a deliberately small codebase, used here for administrator and engineer remote access.
Explained inMulti-datacentre networking
Related servicesEnterprise & Data Centre Networking
- Workload identity
Workload identity gives a running workload its own cloud identity, so it calls provider APIs as itself rather than inheriting the permissions of the machine it runs on.
Also IRSA, Pod Identity
Explained inProduction on AWSProduction on Google Cloud
Related servicesKubernetes & Platform Engineering
Standards and compliance
The frameworks an auditor samples against.
- Audit evidence
Audit evidence is the dated record showing a control operated over a period, which is what an auditor samples rather than a description of the control.
Infrastructure built as code emits most of it as a by-product of normal operation, which is why it costs less at audit time.
Explained inWhat an ISO 27001 auditor asks for
Related servicesIT Audit & Compliance Enablement
- AWS Well-Architected Framework
The AWS Well-Architected Framework is a review structure with six pillars: operational excellence, security, reliability, performance efficiency, cost optimisation and sustainability.
Used as acceptance criteria rather than as a document, it stops a design review from only asking whether the system is up.
Explained inProduction on AWS
Related servicesCloud, Hybrid & Multicloud
- ISO 22301
ISO 22301 is the international standard for business continuity management, which treats exercising recovery as an operating requirement rather than a documentation one.
Explained inMulti-datacentre networkingWhat a real disaster recovery test looks like
Related servicesDisaster Recovery & Business Continuity
- ISO/IEC 27001
ISO/IEC 27001 is the international standard for an information security management system, certified by audit against documented controls that are shown to have operated over time.
An auditor samples evidence that controls operated, not screenshots proving they exist.
Explained inWhat an ISO 27001 auditor asks for
Related servicesIT Audit & Compliance Enablement
- SOC 2
SOC 2 is an attestation report on a service organisation’s controls over security and related criteria, issued by an auditor for a defined period.
Related servicesIT Audit & Compliance Enablement

