Why Operations Teams Eventually Hate Client Secrets
The first client credentials integration takes about five minutes.
The first client credentials integration takes about five minutes.
You register a client, copy the secret, paste it into an environment variable, and a service is authenticating against your API. It works immediately. It's so simple that nobody writes a design doc about it, and nobody puts it on the architecture diagram. It's just how services talk to each other now.
The thousandth one is an operations problem, and it announces itself in a very specific way: someone sends a calendar invite titled "Secret Rotation — Prod" with a two-hour block and eleven people on it.
That meeting is the subject of this article. Not the cryptography — the meeting.
Nobody's secret is in one place
Here's the thing about a shared secret: its entire security model depends on both parties knowing it, which means the moment it exists, it starts spreading.
It goes in the environment variable. Then the deployment manifest, so the environment variable gets set. Then the secrets manager, because someone correctly pointed out that plaintext in a manifest is bad. Then the CI/CD variable store, because the pipeline needs it for integration tests. Then a second secrets manager, because the team that owns the batch jobs is in a different cloud account. Then a developer's .env file, because they were debugging something in November and never deleted it. Then a Slack DM, once, in 2023, because someone was blocked and it was faster.
Ask any platform engineer to enumerate every location where a specific client secret currently exists, and watch how long the pause is. That pause is the entire problem. Not "is the secret strong" — nobody's brute-forcing a 40-character random string. The problem is that a shared secret has a distribution graph, and nobody maintains that graph.
flowchart TD
S["Client Secret"] --> V["Vault / Secrets Manager"]
S --> K["Kubernetes Secret"]
S --> CI["CI/CD variables"]
S --> ENV["Service env vars"]
S --> Dev["Someone's .env file"]
S --> Cust["A customer's own config"]
S --> Slack["That one Slack message"]
Every arrow on that diagram is a place that must be updated when the secret changes. Every arrow you don't know about is an outage waiting for the next rotation.
Rotation is easy in the documentation
The official process is four steps. Generate a new secret. Distribute it. Update the consumers. Revoke the old one.
Here's the same process in a production environment with real customers.
Generate. Fine. This part works.
Distribute. To where? You have twelve services using this client. Three are in Kubernetes, two are Lambda functions, one is a cron job on a VM someone set up before the migration, and one belongs to an enterprise customer who calls your API directly and whose security team requires a change request with two weeks' notice for any credential update.
Update the consumers. Most identity servers support two active secrets during a transition — that's the only thing that makes this survivable, and if yours doesn't, you're looking at a coordinated simultaneous cutover across every consumer, which is a euphemism for downtime. With overlap, you flip services one at a time and watch for errors.
Revoke the old one. And this is where it actually goes wrong. Because to revoke safely you need to know that nothing is still using the old secret. Not "nothing should be." Nothing is. So you go looking for evidence — auth logs filtered by client and secret ID, if your platform even distinguishes which of the two secrets was used, which many don't.
And you find something using the old secret from an IP you don't recognize. Is it a forgotten service? A customer who didn't update? A test harness? Nobody knows. So you extend the overlap window "just to be safe," and the old secret stays valid, and the ticket moves to next quarter, and now you have two valid secrets indefinitely — which is exactly the state rotation was supposed to end.
I've watched this specific sequence play out at multiple organizations. The pattern is so consistent it's almost funny: rotation doesn't fail at the cryptography. It fails at the inventory.
The uncomfortable follow-up question
Here's what makes it worse. That rotation you just spent two hours on — what did it actually accomplish?
If the secret had leaked, rotating it helps. But you don't know whether it leaked, because a shared secret used correctly and a shared secret used by an attacker who copied it produce identical authentication requests. There's no signal. That's not a monitoring gap you can close; it's inherent to the mechanism. Both parties know the secret, so possession proves nothing about identity.
So you rotate on a schedule, as a hedge, because the alternative is never rotating. And every rotation carries real outage risk in exchange for a security benefit you can't measure.
This is the point where a certain kind of engineer starts asking whether the whole approach is wrong.
The change that's smaller than it sounds
Here's the part that surprised me the first time I did this migration: switching from client secrets to private key JWT changes almost nothing about your system.
I mean that quite literally. Here's what a client credentials token request looks like today:
POST /oauth2/token
grant_type=client_credentials
client_id=my-service
client_secret=s3cr3t-value-shared-with-the-server
scope=orders:read
And here's the same request using private key JWT:
POST /oauth2/token
grant_type=client_credentials
client_id=my-service
client_assertion_type=urn:ietf:params:oauth:client-assertion-type:jwt-bearer
client_assertion=eyJhbGciOiJSUzI1NiIsImtpZCI6...
scope=orders:read
That's it. That's the migration.
The client_assertion is a short-lived JWT that the client builds itself and signs with its private key. The authorization server verifies the signature using the client's public key. Everything downstream is unchanged:
flowchart LR
A["Client authenticates\n(secret OR signed assertion)"] --> B["Access Token"]
B --> C["Resource Server"]
C --> D["Same JWT validation,\nsame scopes,\nsame everything"]
style A fill:none,stroke-dasharray: 4 4
The access token you get back is identical. The scopes work the same. The resource server validates it exactly as before and has no idea anything changed. Your API surface doesn't change. Your authorization logic doesn't change. Your token caching doesn't change.
Only the handful of lines that construct the token request change — and in most languages, that's a library call.
For a change with this much operational payoff, the blast radius is remarkably contained. That asymmetry is the actual argument.
What actually changed
The mechanism swap is small. The operational consequences are not.
| Client Secret | Private Key JWT | |
|---|---|---|
| Who knows the credential | Both parties | Only the client |
| What the server stores | The secret itself (or its hash) | A public key — harmless if leaked |
| Distribution | Secret must travel to every consumer | Public key is published; nothing sensitive moves |
| Rotation | Coordinated update across every holder | Client generates a new key, starts signing |
| Leak detection | Impossible — usage looks identical | Private key never transmitted, so exposure surface is one location |
| What's on the wire | The secret, every single request | A short-lived signature; the key itself never transmits |
The row that matters most is the last one. With a shared secret, the credential itself crosses the network on every token request. TLS protects it, so this is mostly fine — until it isn't: a proxy that logs request bodies, a debug dump, a mis-scoped APM trace. With private key JWT, the private key never leaves the client. Ever. Not in transit, not in the server's database, not in a log. The only thing transmitted is a signature that's already expired by the time anyone could reuse it.
And notice what happens to the distribution problem. You still have something to distribute — the public key — but you no longer care who sees it. A public key in a Slack message is a non-event. A public key in a git repo is a non-event. The entire category of "where did this credential end up" anxiety just evaporates, because the sensitive half of the pair was generated on the client and never left.
Where it gets genuinely better: JWKS
Everything so far still involves handing the authorization server a public key when the client registers. Better than a secret, but still a manual step at rotation time — generate new keypair, upload new public key, wait, remove old one.
Unless the client publishes a JWKS endpoint.
sequenceDiagram
participant C as Client
participant AS as Authorization Server
C->>C: Generate new keypair
C->>C: Publish both keys at /.well-known/jwks.json
C->>AS: Token request signed with new key (kid=new)
AS->>C: Fetch JWKS (cache miss on new kid)
AS->>AS: Verify signature with new key
Note over C: Later: drop old key from JWKS
Now rotation is: generate a keypair, add it to your published JWKS, start signing with it, and remove the old one after in-flight assertions have expired.
Read that again and notice what's absent. No coordination meeting. No distribution step. No "did every consumer update." No revocation window where you're hunting for stragglers in auth logs. The client rotates its own credential, unilaterally, and the authorization server discovers the new key on its own because the kid in the assertion header tells it to look.
Rotation stops being an operational event. It becomes a cron job.
That's the aha moment, and it's worth sitting with: the reason secret rotation is painful isn't that rotation is inherently hard. It's that shared secrets require coordination, and coordination between systems owned by different teams is where operational cost comes from. Asymmetric cryptography doesn't make rotation more secure so much as it makes rotation unilateral — and unilateral operations are the ones that actually get automated.
Security is almost a side effect
The conventional pitch for private key JWT leads with cryptography: asymmetric is stronger, the credential never transits, possession is provable. All true, all somewhat abstract.
The argument that actually lands with a platform team is different: this eliminates a recurring, manual, outage-prone process from your operational calendar. You stop having the rotation meeting. You stop maintaining a mental inventory of where a secret lives. You stop extending overlap windows because you can't prove nothing's using the old credential.
The security improvement is real, but it arrives as a consequence of the operational improvement rather than a trade-off against it. That's rare enough to be worth noticing. Most security hardening makes operations harder — more steps, more friction, more ways to break production. This one makes operations simpler and happens to be more secure. When you find one of those, take it.
Where client secrets still make sense
I'm not arguing they're always wrong. A secret is genuinely the right tool when the client can't hold a key securely, when you're integrating with a partner whose stack doesn't support assertion-based auth, or when you have exactly three internal services and rotating them is a ten-minute job you do once a year. The operational cost I've described scales with the number of clients and the number of teams involved — at small scale, it barely exists.
The threshold isn't a specific client count. It's the first time someone asks "where else is this secret used?" and nobody in the room can answer with confidence.
That's the day the model stopped working. The mechanism didn't fail — the inventory did, and no amount of rotation discipline fixes an inventory you can't enumerate. Asymmetric keys sidestep it entirely, not by being cleverer cryptography, but by removing the reason the inventory needed to exist.