Engineering
How we build, and why.
We do not invent our own patterns. We follow the ones the frontier engineering organisations have published and proven at scale, because battle-tested beats clever, and we write down how we apply them here. Each piece is meant to be useful to someone making a decision, including where the honest answer is that you do not need what we sell.
Where a piece builds on published work from Netflix, Google SRE, LinkedIn, Zalando, AWS or the CNCF, the source is cited in it. The reasoning is theirs and proven at a scale we are not claiming; the writing, and how it applies to an estate of normal size, is ours.
- Cloud & private cloud
- Data centre
- Networking
- Operations
- Platform engineering
- Security & compliance
- Structured cabling
Air-gapped Kubernetes: running with no route out
An air-gapped Kubernetes cluster works when every image, chart, operator and binary it needs is mirrored into an internal registry inside the boundary, when artefacts cross the gap through a signed, reviewed and logged transfer rather than by a person with a USB drive, and when the cluster's upgrade path has been rehearsed offline, because the usual remedy of pulling a fixed image is unavailable.
Written forTeams running regulated, classified or operational-technology workloads that cannot have an internet route
Autoscaling that actually reduces the bill
Autoscaling reduces cost only after resource requests are derived from measurement, because a scheduler provisions capacity to satisfy requests rather than usage, so an overstated request reserves machines that nothing runs on and no scaling policy can reclaim them.
Written forEngineering and finance leaders looking at a cloud bill that grew faster than the traffic did
Production on AWS: EKS, ECS and Fargate
A production AWS platform spreads every tier across at least three Availability Zones, provisions nodes with Karpenter rather than fixed node groups, fronts traffic with CloudFront and an Application Load Balancer, deploys through GitHub Actions using OIDC so no long-lived AWS keys exist anywhere, and is accepted against all six AWS Well-Architected pillars rather than against uptime alone.
Written forEngineering leaders and platform teams moving a production workload onto AWS, or paying more for one than they expected
Chaos engineering when you cannot break production
Chaos engineering is applicable without a Chaos Monkey in production because its value comes from stating a steady state in business metrics, forming a falsifiable hypothesis about a specific failure, and bounding the blast radius, all of which can be done in an isolated environment or in a scheduled window.
Written forInfrastructure and operations teams who have read about chaos engineering and concluded, correctly, that they cannot run it the way the talks describe
The golden path, for a team of twelve
A golden path is the one supported route from commit to production that a platform team paves and keeps working, and a small organisation gets most of its value by writing that route down and automating it rather than by building the self-service portal that large platform teams are known for.
Written forEngineering leads at organisations too small for a platform team but large enough that every service is set up differently
Production on Google Cloud: GKE and ephemeral environments
A production Google Cloud platform runs a regional GKE cluster across three zones, authenticates CI and workloads through Workload Identity Federation so no service account keys exist, provisions a complete ephemeral environment per pull request and destroys it on merge, and refreshes lower environments from production on a schedule through an obfuscation job that removes personal data before it ever leaves the production boundary.
Written forPlatform and delivery teams running on Google Cloud who need realistic test data and environments that do not drift from production
Hybrid Exchange: one identity, two mail systems
A hybrid Microsoft Exchange deployment keeps one authoritative identity by synchronising the on-premise directory to the cloud, routes mail through a chosen point rather than both, and retains at least one on-premise Exchange server for as long as the directory is synchronised from on-premise, because recipient attributes remain owned by the local directory and cannot be edited in the cloud.
Written forIT managers running Exchange on-premise who are moving to Microsoft 365, or who moved and still have servers they cannot decommission
The incident review that changes something
An incident review changes something only when each action has a named owner and a date and is tracked to completion, because a blameless culture removes the fear that hides causes but does nothing on its own to ensure the causes are fixed.
Written forTeams who hold incident reviews and keep seeing the same incident
Multi-datacentre networking and the recovery site
A two-datacentre estate with a recovery site works when each site runs its own routing domain joined by a routed interconnect rather than a stretched layer-2 broadcast domain, when the recovery site is built from the same automation as production rather than by hand, and when the failover is exercised on a schedule and timed against the stated recovery objective.
Written forNetwork and infrastructure leads designing a second datacentre, a recovery site, or the links between branches and both
Observability without lock-in: OpenTelemetry first
An observability stack avoids lock-in by instrumenting applications once with OpenTelemetry and sending data through a collector, because the collector is the point at which a backend can be changed, duplicated or split by signal without touching application code or redeploying a single service.
Written forPlatform and operations teams choosing a monitoring stack, or paying more for one than the systems it watches
On-call that does not burn the team
An on-call rotation becomes sustainable when every page is a symptom a user would notice and requires a human decision, because an alert that fires on a cause, or that nobody acts on, trains the responder to ignore the channel the real page will arrive on.
Written forEngineering and IT managers whose on-call rotation is losing people, or who are about to introduce one
On-premise clusters: VMware, Hyper-V and shared storage
An on-premise virtualisation cluster survives a host loss only if it is sized N+1 against the largest host rather than against the average, has a witness placed outside both failure domains so quorum cannot be lost by a single event, and is patched by rolling hosts through maintenance mode with live migration, which requires that no single virtual machine is larger than the spare capacity of one host.
Written forInfrastructure managers running or buying virtualisation on their own hardware, and anyone who has a cluster that has never lost a host on purpose
Secrets, and the places they leak
A secret management programme reduces risk only when it removes long-lived credentials rather than relocating them, because a vault still issues a value that reaches a process, and the leak paths that matter are CI logs, container images, Git history and the process environment rather than the store.
Written forPlatform and security teams who have deployed a secret store and want to know what it did and did not fix
When active-active is the right answer
Multi-region active-active is the right architecture when the business is genuinely fault-intolerant, meaning it cannot accept the outage that any failover implies, and it is one of several ways to meet that requirement rather than a maturity level every estate should reach.
Written forEngineering and business leaders deciding how much a minute of downtime costs, and what architecture that justifies
When event-driven is the wrong shape
Event-driven architecture is the wrong shape when a caller needs an answer before it can continue, because turning a synchronous question into an asynchronous flow does not remove the wait, it only removes the caller's ability to see where the wait is happening.
Written forTeams deciding between a queue and an API call, or debugging a system where that decision was already made
Schema migrations that do not take the site down
A schema migration avoids downtime by expanding the schema first so old and new code both work, deploying the code, backfilling in batches, and only then contracting, because every step in that sequence is independently reversible and none requires the application to stop.
Written forEngineering teams who have taken an outage for a migration, or who are about to run one on a table that has grown
Zero trust without rebuilding the network
Zero trust can be adopted incrementally by moving authorisation decisions from network location to verified identity and device state one application at a time, starting with the applications a compromised laptop would reach first, rather than by replacing the network.
Written forIT and security leads who have been told to adopt zero trust and have an existing network, a VPN and a budget
Securing the AI development lifecycle
Secure AI-DLC applies secure development lifecycle discipline to systems built on AI models: controlling the model supply chain, treating prompts and retrieved content as untrusted input, governing what data crosses the inference boundary, restricting access to model endpoints, and testing for prompt injection and data leakage.
Written forCTOs, security leads and platform teams shipping features built on large language models
What a real disaster recovery test looks like
A real disaster recovery test restores actual systems from actual backups into an isolated environment, measures how long it truly took against the stated RTO, verifies the restored data against the stated RPO, and records what failed: anything short of that is a plan review, not a test.
Written forIT managers, heads of infrastructure and risk officers who own a disaster recovery plan they have never executed
How to choose a structured cabling contractor
Choose a structured cabling contractor on three verifiable things: the certified test results they will hand over per port, the as-built documentation they produce, and the warranty covering installation workmanship as well as materials, because everything else they tell you is unverifiable before the walls close.
Written forFacilities managers, IT managers and project managers commissioning an office fit-out or building refurbishment
What ISO 27001 auditors ask for from infrastructure
ISO 27001 auditors do not primarily test whether your controls exist. They sample evidence that the controls operated over time, which means an asset inventory matching reality, joiner-mover-leaver access records, periodic access reviews, production change records, vulnerability scan and remediation history, restore test results, and log retention proof.
Written forCISOs, IT managers and compliance leads preparing for a first ISO 27001 certification audit
Private or public cloud for Armenian financial institutions
Armenian financial institutions generally choose private cloud over public cloud for three reasons: data residency requirements over customer data, simpler auditability of an owned platform under regulatory scrutiny, and lower steady-state cost for predictable always-on workloads, while accepting that they carry the capacity planning and hardware refresh burden in return.
Written forCIOs, heads of IT and risk officers at banks, universal credit organisations and other regulated financial institutions
Have a question we have not written about?
Ask it. If the answer is useful to more than one person, we will probably publish it.

