An air-gapped cluster is not a hardened cluster with the firewall turned up. It is a cluster for which the usual remedy — pull the fixed image, apply the upstream manifest — does not exist. Every assumption about reaching out has to be removed before isolation, not after.
This is what the build looks like, and the order it has to happen in.
Discover the dependencies while you still have a network
The most common failure is an image reference nobody knew about. Helm charts pull sidecars, operators fetch their own controllers during reconciliation, init containers reference upstream registries by digest. Each of these works perfectly until the day there is no route, and then fails during an incident.
So the first step happens connected: stand the full platform up in a normal environment, then apply an admission policy that rejects any image whose registry is not the internal mirror. Everything that breaks is a dependency that was not on the list. Iterate until nothing breaks.
This diagram as text
- Connected build of the full platform
- Admission policy: internal registry only — every rejection is a dependency nobody wrote down
- Add to the mirror manifest
- Re-apply
- Repeat until clean
Relationships
- Connected build of the full platform → Admission policy: internal registry only
- Admission policy: internal registry only → Add to the mirror manifest — rejected
- Add to the mirror manifest → Re-apply
- Re-apply → Repeat until clean
The output is a manifest of every image, chart, operator, binary and CRD bundle the platform needs, pinned by digest. That manifest is a deliverable in its own right, and it is what makes the mirror reproducible rather than a thing that accumulated.
Inside the boundary: a registry, not a cache
A pull-through cache still needs a route when it misses. Inside the gap you need a real registry holding real copies.
- Registry
- internal, authoritative, not a cache
- Trust
- the signature is verified inside, the origin is not trusted
- Time
- internal NTP, because certificates depend on it
- Admission
- internal registry only, enforced
This diagram as text
- Outside connected build environment
- Build
- Scan
- Sign, with provenance — images by digest, charts, SBOM
- Controlled one-way transfer — reviewed, logged, no return path
- Inside air-gapped
- Verify the signature against the expected key — admitted only if it passes
- Internal registry — authoritative, not a cache
- Object store — charts and binaries
- Internal time source — certificates depend on it
- Kubernetes cluster — pulls only from here
- Observability — local retention, sized
- Internal certificate authority
Relationships
- Build → Scan
- Scan → Sign, with provenance
- Sign, with provenance → Controlled one-way transfer — bundle
- Controlled one-way transfer → Verify the signature against the expected key
- Verify the signature against the expected key → Internal registry
- Verify the signature against the expected key → Object store
- Internal registry → Kubernetes cluster — pull
- Internal time source → Internal certificate authority
An object store inside the boundary — MinIO serves this well — holds the Helm charts, the OS packages, the node images and the backup targets. It also gives the cluster somewhere to put backups that is not the cluster.
Two services are easy to forget and both cause confusing failures:
- An internal certificate authority. Nothing can reach a public CA to issue or to check revocation. Certificate lifecycle has to be owned, with issuance and renewal automated, because manual renewal inside an air gap is how a cluster expires on a weekend.
- An internal time source. Certificate validation, token expiry and log correlation all depend on clocks that agree. There is no public NTP, so the boundary needs its own stratum source.
Crossing the gap
The transfer mechanism matters less than the discipline around it. What has to be true:
- It is one-way by design. A bidirectional link that is firewalled is not an air gap; it is a firewall, and it will eventually be misconfigured.
- Nothing is admitted because of where it came from. It is admitted because a signature was verified on the inside, against a key held on the inside. This is the difference between a boundary and a habit.
- Every transfer is logged with its content hashes, who authorised it and when. This log is the evidence an auditor will ask for, and it is also how you answer “when did that version arrive”.
- The process is rehearsed. An emergency patch should follow a procedure people have performed, not one they are reading for the first time at 2am.
Operating without the usual escape hatches
Several everyday practices do not survive isolation and need replacements decided in advance:
| What normally happens | Inside the gap |
|---|---|
kubectl apply -f an upstream URL |
Manifests are mirrored and version-controlled inside |
| Pull a fixed image tag | The fix must be built, signed and transferred first |
| Vulnerability feeds update themselves | Feeds are part of the scheduled transfer, with their own freshness alert |
| Managed control plane upgrades | Upgrades are rehearsed offline, on a replica of the cluster |
| Vendor support takes a remote session | Diagnostics are exported, sanitised and carried out |
The last two are where projects stall. An upgrade path that has only ever been performed with a network is not a tested upgrade path, so we build a rehearsal cluster inside the boundary and prove the upgrade there first. And a support agreement that assumes remote access needs renegotiating before it is needed, not during an incident.
Observability has to be local, and sized
Logs and metrics cannot be shipped to a hosted platform. The stack lives inside, and its retention is a capacity decision made at design time rather than discovered when a disk fills. Retention also tends to be the longest in exactly these environments, because the regulatory regimes that require air gaps usually require long log retention as well.
How STP approaches this
We discover the dependency list in a connected environment first, so isolation does not surface it one failure at a time; we make signature verification on the inside the admission rule rather than trusting the transfer; we build the internal certificate authority and time source as part of the platform rather than as follow-ups; and we rehearse the upgrade and the emergency patch before handover, because in an air gap the rehearsal is the only thing that makes either of them fast.
More on private cloud, Kubernetes platform engineering and audit and compliance, or start a conversation.

