Standing up a new environment used to mean a person working through a written runbook — networking, permissions, cluster add-ons, secrets wiring — with hand-offs between people at several points. It took about two days, most of which was waiting. I turned it into a set of reusable modules a product team runs themselves, and the same build now takes under an hour.
Platform team of six engineers and a manager, supporting four product teams across 15+ AWS accounts. I owned the module layer — the network and cluster modules, the per-environment composition, and the pipeline that applies them. In use since late 2023 — every environment built since has come from them. What each team then builds inside their environment is theirs; my job stopped at handing over something consistent to build on.
A new environment went from about two days to under an hour.
wall-clock from terraform apply to a namespace serving traffic, averaged over the last three environment builds. The old two days were not compute time — they were hand-offs. Networking, IAM roles, cluster add-ons and secrets wiring each sat in a different person's queue.
The same modules build every environment across 15+ accounts.
the account count is the AWS Organizations tree — three environments across four product and platform domains, plus shared services, security tooling, log archive and networking. It's 15 or 16 depending on whether sandbox accounts are counted, so I say 15+.
The change that mattered more than the time saved: differences between environments became declared. Built from the same modules with different variables, staging differing from production is a line someone wrote on purpose rather than something that accumulated. Most "but it worked in staging" incidents are this problem.
One set of modules; one composition per environment, each with its own state backend and variables. Applied through CI, never from a laptop.
Two decisions on that path are worth stating, because both traded convenience for something else.
Non-production doesn't get a NAT gateway unless it asks. At roughly $32 a month before data charges, per gateway per environment, it's the kind of cost not worth a conversation individually and worth real money across an estate. Gated on a variable, dev and staging run without one by default.
The cluster doesn't hand admin rights to whoever created it. AWS grants cluster-admin to the identity that ran the create — in a CI-provisioned world that's a pipeline role, a grant nobody chose and nobody can see in the manifests. Turning it off makes every grant explicit:
# Access entries only — no aws-auth ConfigMap.
access_config {
authentication_mode = "API"
# Declining the automatic admin grant. Whoever ran `apply` should not
# silently become cluster-admin; every grant is an explicit access entry.
bootstrap_cluster_creator_admin_permissions = false
}
Subnets are built per availability zone from one variable, and carry the tags the load balancer controller discovers them by — the sort of detail that is invisible when it's right and produces a very confusing afternoon when it's missing:
resource "aws_subnet" "public" {
count = length(var.azs) # one per AZ, 3 by default
availability_zone = var.azs[count.index]
cidr_block = local.public_subnet_cidrs[count.index]
tags = merge(
{
Name = "${var.name}-${var.environment}-public-${count.index + 1}"
# How the AWS load balancer controller finds where to put
# public load balancers. Untagged subnets fail silently.
"kubernetes.io/role/elb" = "1"
"kubernetes.io/cluster/${var.cluster_name}" = "shared"
},
var.tags,
)
}
A module hardcoded its storage class, and the assumption held right up until it didn't.
gp3 is the correct default — newer, cheaper, faster — so it went into the module as a constant, and every environment built from it worked. Then it met an older cluster without that storage class, and the volume claims sat in Pending. Nothing errored. The workload just never started.
The failure is specific to the thing that makes shared modules valuable. A module encodes decisions so that nobody has to make them again — which is exactly the point, and also means a decision that isn't universally true gets propagated everywhere with confidence. The wider the reuse, the further a wrong assumption travels before anyone questions it.
I'd standardize on one storage class again. What I'd change is the difference between a default and a constant: a default says "this is what you want unless you know otherwise" and can be overridden; a constant says "this is always true," which for an infrastructure assumption is worth being suspicious of. It's now a variable with gp3 as its default.