Designing for Key Rotation Before You Need It

Here's a question worth asking your team today, while nothing is on fire:

Here's a question worth asking your team today, while nothing is on fire:

Can we rotate our token signing key right now, on a Tuesday afternoon, without telling anyone?

If the answer is yes, this article isn't for you. If the answer involves a maintenance window, a coordination email, or a pause while somebody thinks about it — you don't have key rotation. You have a key rotation project, and the day you need it will be the day a key is compromised, which is the worst possible day to be running an untested multi-step procedure.

Rotation is one of a small set of capabilities that costs almost nothing to design in and is genuinely painful to retrofit. It's worth spending an afternoon on before you need it.

Why it's harder than it looks

The naive model: generate a new key, start signing with it, delete the old one. Three steps, obviously correct.

It breaks immediately, because of a property specific to token-based systems: tokens outlive the moment they were signed.

At the instant you swap keys, there are tokens in flight — sitting in browser storage, in mobile app keychains, in service-to-service caches, mid-request — all signed with the old key, all still within their validity period. Delete the old key and every one of them fails validation at once. Users who were happily authenticated thirty seconds ago are now getting 401s, and your error rate goes vertical for the length of your token lifetime.

So rotation isn't a swap. It's an overlap, and the overlap has a minimum duration you don't fully control.

The four things you need

Rotation requires four mechanisms working together. Miss any one and rotation means downtime.

1. Multiple simultaneously-valid keys. Your validators must accept tokens signed by any key in a current set, not just one. This has to be true of your own services and of every third-party library your customers use to validate your tokens.

2. A kid on every token. The key ID in the JWT header tells a validator which key to use. Without it, a validator holding three keys has to try each one until something verifies — which works, is slower, and gets genuinely ugly during an incident when one of those keys has been revoked. Emit kid from day one; it costs nothing and it's nearly impossible to add retroactively to tokens already in circulation.

3. A discovery endpoint. JWKS, published at a stable URL, listing every currently-valid public key. This is how validators learn about a new key without you telling them.

4. A caching contract you've actually thought about. Every consumer caches your JWKS. You don't control their TTLs. This is the constraint that shapes everything else.

The timeline that works

Rotation is a phased operation, and the phases exist because of consumer caching.

flowchart LR
    P1["Phase 1<br/>Publish new key<br/>(don't sign with it)"] --> P2["Phase 2<br/>Wait > longest<br/>consumer cache TTL"]
    P2 --> P3["Phase 3<br/>Start signing<br/>with new key"]
    P3 --> P4["Phase 4<br/>Wait > longest<br/>token lifetime"]
    P4 --> P5["Phase 5<br/>Retire old key"]

Phase 1: publish, don't sign. Add the new public key to your JWKS. Keep signing with the old one. Nothing changes for anybody.

Phase 2: wait for propagation. Consumers refresh their JWKS cache on their own schedule. You need to wait longer than the longest plausible TTL out there. For internal services you control, this might be minutes. For third parties, assume hours — some libraries default to caching for 24 hours, and some cache until they see an unknown kid.

Phase 3: start signing with the new key. By now every consumer already has it. A token arrives with a new kid, and the validator finds it in cache without a network call. Nobody notices anything.

Phase 4: wait for old tokens to die. The old key must stay published until every token signed with it has expired. That's your maximum token lifetime plus clock skew tolerance.

Phase 5: retire the old key. Remove it from JWKS. Now it's genuinely gone.

The whole thing might take a day. That's fine — it's a background process, not an outage window, and no user experiences anything at any point. That property is the entire goal.

The part that catches people: emergency rotation

Everything above assumes you're rotating on schedule. Emergency rotation — a key is compromised, and it must stop being trusted now — is a different operation, and conflating the two is a common design error.

The scheduled timeline has a phase where you deliberately keep the old key valid. In an emergency, that's exactly what you can't do. You need the old key to stop being accepted immediately, and you have to accept the consequence: every token signed with it becomes invalid at once. Everyone gets logged out. Every service-to-service call fails until it re-authenticates.

That's a real, user-visible outage — and it's the correct behavior. The mistake is not having decided in advance that it's the correct behavior, so that when it happens, nobody is arguing about it in an incident channel.

Two things make this survivable:

Have the emergency path built and tested. It's a different runbook from scheduled rotation, and it's the one that will run under pressure. Practice it in a lower environment.

Understand your recovery shape. When every client is simultaneously rejected, every client simultaneously re-authenticates. That's a thundering herd against your token endpoint at exactly the moment you're already handling an incident. Jittered client backoff matters here more than usual.

The failure modes worth knowing about

Validators that cache forever. Some client libraries fetch JWKS once at startup and never refresh. These break at phase 3 and stay broken until restarted. You'll discover them during your first rotation, which is an argument for rotating early and often rather than waiting until it's urgent — the first rotation is a discovery exercise, and you want to run it when the stakes are low.

Validators that refetch on every unknown kid. The opposite problem. A malformed or malicious token with a random kid triggers a JWKS fetch. Send a thousand of those and you've got a thousand fetches. Well-behaved libraries rate-limit this and cache negative lookups briefly; not all libraries are well-behaved, and if you're writing the validator, this is on you.

Rotating the key without rotating anything that references it. Signing keys are often not the only key material in play. SAML signing certificates, encryption keys, client certificates for mTLS — each has its own rotation story, and each tends to be forgotten until it expires. Certificate expiry in particular is a scheduled outage you've agreed to in advance and then forgotten about.

Nobody knows when things expire. Every certificate and key should have its expiry as a monitored metric with alerting well before the date — 60 days is a reasonable first alert, since that's roughly the lead time for coordinating with an external party. An expiry date sitting only in a certificate file is not a monitoring strategy.

The organizational half

The technical mechanism is the easy part. Two things reliably make rotation harder than it should be:

External coordination. For SAML federation, certificate rotation requires the customer's IdP administrator to update their configuration. That person may be in a different company, on a change-freeze, or simply unreachable. This is why SAML certificate rotation is genuinely harder than JWKS rotation — there's no automatic discovery mechanism, so it's a human process. Support dual certificates where the protocol allows it, and start the conversation months early.

Nobody owns it. Rotation is a task with no natural owner. It isn't a feature, it doesn't have a customer asking for it, and it appears on no roadmap. So it gets deferred until it's urgent. Making it automatic — a scheduled job that runs the phased timeline without human involvement — is the only reliable fix, because anything requiring a human to remember will eventually not happen.

The counterargument

There's a reasonable position that rotation is over-emphasized. If your keys are stored in an HSM or a managed KMS, never leave it, and are only accessible to a small set of well-audited services, then the actual probability of compromise is low, and frequent rotation is ceremony that introduces operational risk in exchange for a security benefit you can't measure. There's a real echo of the password-rotation argument here: mandatory periodic change of a strong, well-protected secret is not obviously net-positive.

I think that's partly right and it leads to the wrong conclusion. The argument for building rotation isn't that you should rotate constantly. It's that the capability must exist and be exercised, because the day you need it is not a day you get to choose. A rotation path that has never been run is a rotation path that doesn't work — the caching consumer nobody knew about, the service with a hardcoded key, the library that fetches JWKS once at boot.

Rotating quarterly isn't primarily about limiting key exposure. It's about ensuring that when you have to rotate in an hour, the procedure is boring.

The test, again

Can you rotate a signing key right now, on a Tuesday afternoon, without telling anyone?

If yes, you have four mechanisms in place and a timeline that respects consumer caching, and the emergency version is a documented variant of something you do routinely.

If no, the work to get there is maybe a week — and it's a week that's dramatically cheaper now than during the incident where you first need it.