An air-gapped cluster is not a hardened cluster with the firewall turned up. It is a cluster for which the usual remedy — pull the fixed image, apply the upstream manifest — does not exist. Every assumption about reaching out has to be removed before isolation, not after.

This is what the build looks like, and the order it has to happen in.

Discover the dependencies while you still have a network

The most common failure is an image reference nobody knew about. Helm charts pull sidecars, operators fetch their own controllers during reconciliation, init containers reference upstream registries by digest. Each of these works perfectly until the day there is no route, and then fails during an incident.

So the first step happens connected: stand the full platform up in a normal environment, then apply an admission policy that rejects any image whose registry is not the internal mirror. Everything that breaks is a dependency that was not on the list. Iterate until nothing breaks.

Dependency discovery, before isolation
Discovering undocumented image dependencies before a cluster is isolatedBuild the full platform in a connected environment, but enforce an admission policy that permits the internal registry and nothing else. Every rejection is an undocumented dependency. Add it to the mirror manifest, re-apply, and repeat until nothing is rejected.rejectedConnected build of the full platformAdmission policy: internal registry onlyevery rejection is a dependency nobody wrote downAdd to the mirror manifestRe-applyRepeat until clean
This diagram as text
  • Connected build of the full platform
  • Admission policy: internal registry only — every rejection is a dependency nobody wrote down
  • Add to the mirror manifest
  • Re-apply
  • Repeat until clean

Relationships

  • Connected build of the full platform → Admission policy: internal registry only
  • Admission policy: internal registry only → Add to the mirror manifest — rejected
  • Add to the mirror manifest → Re-apply
  • Re-apply → Repeat until clean

The output is a manifest of every image, chart, operator, binary and CRD bundle the platform needs, pinned by digest. That manifest is a deliverable in its own right, and it is what makes the mirror reproducible rather than a thing that accumulated.

Inside the boundary: a registry, not a cache

A pull-through cache still needs a route when it misses. Inside the gap you need a real registry holding real copies.

The boundary, and what sits either side of it
An air-gapped Kubernetes platform, with signing outside and verification insideOutside the boundary, a connected environment builds, scans and signs with provenance, producing a bundle of images by digest, charts, a software bill of materials and signatures. The transfer across the gap is controlled, reviewed, logged and one-way, with no return path. Inside, the signature is verified against an expected key before anything is admitted. An internal registry is authoritative rather than a cache, an object store holds charts and binaries, and an internal time source and certificate authority keep the cluster working without the outside world.OUTSIDEconnected build environmentINSIDEair-gappedbundlepullBuildScanSign, with provenanceimages by digest, charts, SBOMControlled one-way transferreviewed, logged, no return pathVerify the signature against the expected keyadmitted only if it passesInternal registryauthoritative, not a cacheObject storecharts and binariesInternal time sourcecertificates depend on itKubernetes clusterpulls only from hereObservabilitylocal retention, sizedInternal certificate authority
Registry
internal, authoritative, not a cache
Trust
the signature is verified inside, the origin is not trusted
Time
internal NTP, because certificates depend on it
Admission
internal registry only, enforced
This diagram as text
  • Outside connected build environment
    • Build
    • Scan
    • Sign, with provenance — images by digest, charts, SBOM
  • Controlled one-way transfer — reviewed, logged, no return path
  • Inside air-gapped
    • Verify the signature against the expected key — admitted only if it passes
    • Internal registry — authoritative, not a cache
    • Object store — charts and binaries
    • Internal time source — certificates depend on it
    • Kubernetes cluster — pulls only from here
    • Observability — local retention, sized
    • Internal certificate authority

Relationships

  • Build → Scan
  • Scan → Sign, with provenance
  • Sign, with provenance → Controlled one-way transfer — bundle
  • Controlled one-way transfer → Verify the signature against the expected key
  • Verify the signature against the expected key → Internal registry
  • Verify the signature against the expected key → Object store
  • Internal registry → Kubernetes cluster — pull
  • Internal time source → Internal certificate authority

An object store inside the boundary — MinIO serves this well — holds the Helm charts, the OS packages, the node images and the backup targets. It also gives the cluster somewhere to put backups that is not the cluster.

Two services are easy to forget and both cause confusing failures:

  • An internal certificate authority. Nothing can reach a public CA to issue or to check revocation. Certificate lifecycle has to be owned, with issuance and renewal automated, because manual renewal inside an air gap is how a cluster expires on a weekend.
  • An internal time source. Certificate validation, token expiry and log correlation all depend on clocks that agree. There is no public NTP, so the boundary needs its own stratum source.

Crossing the gap

The transfer mechanism matters less than the discipline around it. What has to be true:

  • It is one-way by design. A bidirectional link that is firewalled is not an air gap; it is a firewall, and it will eventually be misconfigured.
  • Nothing is admitted because of where it came from. It is admitted because a signature was verified on the inside, against a key held on the inside. This is the difference between a boundary and a habit.
  • Every transfer is logged with its content hashes, who authorised it and when. This log is the evidence an auditor will ask for, and it is also how you answer “when did that version arrive”.
  • The process is rehearsed. An emergency patch should follow a procedure people have performed, not one they are reading for the first time at 2am.

Operating without the usual escape hatches

Several everyday practices do not survive isolation and need replacements decided in advance:

What normally happens Inside the gap
kubectl apply -f an upstream URL Manifests are mirrored and version-controlled inside
Pull a fixed image tag The fix must be built, signed and transferred first
Vulnerability feeds update themselves Feeds are part of the scheduled transfer, with their own freshness alert
Managed control plane upgrades Upgrades are rehearsed offline, on a replica of the cluster
Vendor support takes a remote session Diagnostics are exported, sanitised and carried out

The last two are where projects stall. An upgrade path that has only ever been performed with a network is not a tested upgrade path, so we build a rehearsal cluster inside the boundary and prove the upgrade there first. And a support agreement that assumes remote access needs renegotiating before it is needed, not during an incident.

Observability has to be local, and sized

Logs and metrics cannot be shipped to a hosted platform. The stack lives inside, and its retention is a capacity decision made at design time rather than discovered when a disk fills. Retention also tends to be the longest in exactly these environments, because the regulatory regimes that require air gaps usually require long log retention as well.

How STP approaches this

We discover the dependency list in a connected environment first, so isolation does not surface it one failure at a time; we make signature verification on the inside the admission rule rather than trusting the transfer; we build the internal certificate authority and time source as part of the platform rather than as follow-ups; and we rehearse the upgrade and the emergency patch before handover, because in an air gap the rehearsal is the only thing that makes either of them fast.

More on private cloud, Kubernetes platform engineering and audit and compliance, or start a conversation.