Fifty services, and no way to answer "is the platform healthy?" without opening three tools and forming an opinion. Now there is one number, agreed in advance.
MTTR: median (not mean) sev-2 resolution time, the quarter before burn-rate alerting against the quarter after, measured from impact rather than detection. One quarter against another, not a controlled experiment, and few sev-2s in a quarter — directional, not precise.
Alert volume: alerts actually delivered by Alertmanager per week, before and after. The drop came from changing what the rules ask, not from deleting rules. That the burn-rate conditions still catch what the thresholds caught was my judgement at the time, not something I proved.
Error budget: arithmetic. 0.1% of a rolling 30-day window is 43.2 minutes.
The best evidence isn't on this page: a credential rotation later took out part of a service's capacity, and the burn-rate alert fired before anyone reported it. That incident is written up here.
Availability over a rolling 30-day window, 5xx as failure:
# 30-day availability. 5xx counts as failure; 4xx is treated as client error.
# Namespace redacted.
sum(increase(nginx_ingress_controller_requests{
namespace="<app>", status!~"5.."
}[30d]))
/
sum(increase(nginx_ingress_controller_requests{
namespace="<app>"
}[30d]))
* 100
Replacing Datadog for our services with a stack we ran ourselves was the decision worth defending: the signals already existed in Prometheus format, so self-managing meant scraping what was there instead of paying per host to re-ingest it — and retention and cardinality stayed ours. Datadog stayed in use elsewhere in the company. We paid for this by carrying the stack.
The SLI counts some of our own failures as successes, and I shipped it that way on purpose. It excludes everything that isn't a 5xx — but a 429 is a 4xx, and a 429 is us. When the platform is rejecting requests under load, the SLI scores every rejection as a success. The measurement is most optimistic exactly when the platform is least healthy.
I'd ship the same v1 again — a defensible simplification that exists beats a perfect definition still being argued about. What I'd change is treating the exclusion as finished rather than provisional, and writing the bias into the SLO document instead of carrying it in my head.
Platform team of six engineers and a manager. Three environments, 40–50 services across 10–15 namespaces, 20–30 nodes, 15+ production deploys a day. I owned this end to end — the Helm chart, Prometheus configuration, SLO definition, dashboards, alerting rules and runbooks. The 99.9% target has held since early 2024. IAM and networking were shared team ground.
The operational cost arrived during a staging rollout when Alertmanager crash-looped — the base chart enabled HA gossip unconditionally while staging ran a single replica, so it started up hunting for peers that did not exist. The chart fix was small; the change that mattered was a config-validity check in CI, so a config that cannot start now fails before merge.