Making Secrets Manager the source of truth for live applications

What it is

Application credentials lived in hand-maintained Kubernetes Secrets — base64, not encryption, readable by anyone with namespace access. I moved the source of truth into AWS Secrets Manager, auditable and centrally rotatable, with each workload on a scoped identity that reads only its own. All under production traffic.

Scope and ownership

Platform team of six engineers and a manager. I owned the consumer layer end to end: the inventory, the cutover sequencing, the per-app IAM scoping, the pattern itself, and the migration pull requests. I deliberately did not take ownership of the CSI driver install — it had been installed by hand from another workstream, and adopting it would have meant a second release owner on a live DaemonSet that another team's database already depended on. I documented who actually owned it instead.

What changed

The authoritative copy of every credential moved out of the cluster and into a store that can be audited and rotated, with a scoped IAM role per application rather than one shared credential.

How this was measured

the security claim is structural rather than measured. Each app's role grants read on exactly one secret path, so the blast radius of a compromised workload is that app's own credentials — that's a property of the policy, checkable by reading it, not a number I collected.

Zero application code changes. 40+ services moved onto the pattern, the first two on consecutive days with no rollback.

How this was measured

verifiable in the migration pull requests — they touch manifests only, no application source. The first two were sequenced by ascending risk to prove the pattern before the rest of the estate followed.

How it works

The CSI driver runs on every node and fetches secrets from AWS at pod start, authenticating through IRSA — so there is no long-lived credential anywhere in the path. Each application declares what it needs, and gets a service account annotated with a role scoped to just its own secret.

How a credential reaches a running application At pod start, the CSI driver reads the application's SecretProviderClass, assumes an IAM role scoped to that application's secret path via IRSA, and fetches the secret from AWS Secrets Manager. The driver mounts it as a volume and also syncs it into a Kubernetes Secret using the key names the application already expects, so the application reads its configuration from environment variables exactly as it did before. pod starts app unchanged CSI driver reads SecretProviderClass IRSA role one secret path only Secrets Manager auditable · rotatable env vars same key names assumed at pod start
The application is the one part of this diagram that did not change.

The detail that decided the design: every application read config as environment variables through envFrom.secretRef — none mounted files. The obvious approach, a plain volume mount, would have delivered files to applications that never read files. Nothing would break loudly; they would simply start with nothing.

So each declaration carries a secretObjects block that syncs the fetched values into a Kubernetes Secret using the key names the application already expected. That mapping is the entire reason no application code had to change:

# Redacted. The secret ARN and app name are placeholders.
apiVersion: secrets-store.csi.x-k8s.io/v1
kind: SecretProviderClass
spec:
  provider: aws
  parameters:
    objects: |
      - objectName: "arn:aws:secretsmanager:<region>:<account>:secret:<app>-db"
        objectType: "secretsmanager"
        jmesPath:                      # pull individual keys out of the JSON
          - path: '"username"'
            objectAlias: "username"
          - path: '"password"'
            objectAlias: "password"

  # The part that made this a zero-code-change migration: sync into a
  # Kubernetes Secret under the SAME key names the app already read.
  secretObjects:
    - secretName: <app>-db-credentials
      type: Opaque
      data:
        - key: username
          objectName: username
        - key: password
          objectName: password

The trap, which caught me: the volume mount is required even when the app only uses envFrom. The mount isn't how the secret is delivered — it's what triggers the fetch at all. Leave it out and the driver never runs, the Secret is never created, and the app starts anyway with empty variables.

I sequenced the cutover by ascending risk, not ticket order — first the app that already had its own service account. Before any of it, a throwaway pod proved the driver actually fetched. Doing three at once would have turned one lesson into three incidents.

What went wrong

Later, a credential rotation landed on one of these services — and this driver fetches a secret when a pod starts, not when the secret changes. A rotation only reaches a pod when that pod restarts.

A deploy went out shortly after. New pods failed their readiness probe, so the rolling update stalled — some replicas still serving on the old credential, others unable to start. Error rate climbed on the paths that had lost capacity, and the burn-rate alert fired before anyone reported it.

the fleet, mid-rolloutrolling update stalled
old credential, still serving new pod, failed readiness
This is what made it confusing. The service was neither up nor down — it was both, and only the paths that had lost capacity showed errors.

My first assumption was a bad deploy — that is what a dashboard suggests when error rate rises next to a deployment marker. It was wrong. The deploy was a symptom: it just happened to be what made pods restart and pick up the new value.

I stopped the rollout before I understood the cause. That kept the healthy replicas serving and stopped it getting worse, which bought the time to actually look.

The cause was the shape of the value. Extraction is type-sensitive: a port stored as 3306 rather than "3306" fails the jmesPath lookup, so the Secret was never built and readiness failed. Nothing in that path errors loudly — it quietly doesn't work. Twelve minutes of degraded capacity, about a quarter of that month's error budget.

Correcting the value took a minute. The change that mattered was validating a secret's shape before it can be rotated in, rather than discovering it whenever a pod next restarts — possibly days later, when nobody connects the two events.

I'd make the same call again — secretObjects mapping is what let production apps move without touching their code. What I'd change is how I verified: I watched pods come up healthy, and on this path healthy is exactly what the failure looks like. The check has to be that the credential arrived, not that the process started.