Giving a platform a reliability number it could defend

Fifty services, and no way to answer "is the platform healthy?" without opening three tools and forming an opinion. Now there is one number, agreed in advance.

a month's error budget at 99.9%43.2 min
12 min
A twelve-minute incident spends roughly a quarter of the month. That is the whole point of measuring a budget rather than counting outages — it turns "was that bad?" into a number.

What changed

Static thresholds fire on a moment. Burn-rate rules fire on a trend, so a spike that recovers in thirty seconds never pages anyone.
How these were measured

MTTR: median (not mean) sev-2 resolution time, the quarter before burn-rate alerting against the quarter after, measured from impact rather than detection. One quarter against another, not a controlled experiment, and few sev-2s in a quarter — directional, not precise.

Alert volume: alerts actually delivered by Alertmanager per week, before and after. The drop came from changing what the rules ask, not from deleting rules. That the burn-rate conditions still catch what the thresholds caught was my judgement at the time, not something I proved.

Error budget: arithmetic. 0.1% of a rolling 30-day window is 43.2 minutes.

The best evidence isn't on this page: a credential rotation later took out part of a service's capacity, and the burn-rate alert fired before anyone reported it. That incident is written up here.

How it works

How telemetry reaches a dashboard and an alert Ingress metrics, container and node metrics, and application metrics are scraped by Prometheus every thirty seconds with thirty days of retention. Prometheus feeds Grafana, which also reads AWS CloudWatch directly using an IRSA role rather than stored credentials. Prometheus also feeds Alertmanager, which routes alerts by severity and owning team to the team that owns the affected service. ingress traffic containers · nodes application metrics AWS CloudWatch Prometheus 30s · 30-day retention Alertmanager routes by severity + team Grafana SLO · budget · deploys the owning team not a shared channel IRSA, no stored keys
Everything is scraped from what the cluster already emits. The only external source is CloudWatch, reached with a scoped IRSA role rather than a stored credential.

Availability over a rolling 30-day window, 5xx as failure:

# 30-day availability. 5xx counts as failure; 4xx is treated as client error.
# Namespace redacted.
sum(increase(nginx_ingress_controller_requests{
      namespace="<app>", status!~"5.."
    }[30d]))
/
sum(increase(nginx_ingress_controller_requests{
      namespace="<app>"
    }[30d]))
* 100

Replacing Datadog for our services with a stack we ran ourselves was the decision worth defending: the signals already existed in Prometheus format, so self-managing meant scraping what was there instead of paying per host to re-ingest it — and retention and cardinality stayed ours. Datadog stayed in use elsewhere in the company. We paid for this by carrying the stack.

What went wrong

The SLI counts some of our own failures as successes, and I shipped it that way on purpose. It excludes everything that isn't a 5xx — but a 429 is a 4xx, and a 429 is us. When the platform is rejecting requests under load, the SLI scores every rejection as a success. The measurement is most optimistic exactly when the platform is least healthy.

I'd ship the same v1 again — a defensible simplification that exists beats a perfect definition still being argued about. What I'd change is treating the exclusion as finished rather than provisional, and writing the bias into the SLO document instead of carrying it in my head.

Scope, ownership, and the operational cost

Platform team of six engineers and a manager. Three environments, 40–50 services across 10–15 namespaces, 20–30 nodes, 15+ production deploys a day. I owned this end to end — the Helm chart, Prometheus configuration, SLO definition, dashboards, alerting rules and runbooks. The 99.9% target has held since early 2024. IAM and networking were shared team ground.

The operational cost arrived during a staging rollout when Alertmanager crash-looped — the base chart enabled HA gossip unconditionally while staging ran a single replica, so it started up hunting for peers that did not exist. The chart fix was small; the change that mattered was a config-validity check in CI, so a config that cannot start now fails before merge.

All projects · Home