The Cost of Key Rotation
The security standard said it in one line: all cryptographic key material shall be rotated at least every 90 days. It became one ticket, with a checklist, assigned to one engineer, sized at three points.
Item one was the token signing keys. A cron job ran the phased rotation across all 200 tenants, took about forty seconds of wall clock, and produced no dashboard movement of any kind. Item two was the TLS certificates, which had been renewing themselves via ACME every sixty days for two years and were on the checklist only because somebody had copied the checklist from the standard.
Item seven was the SAML signing certificate — one certificate, on their side, consumed by 190 customer identity providers.
That item took fourteen weeks. It consumed most of one engineer and a slice of four account managers. It produced six customer-visible SSO outages, three of which were caused by an IdP that could only hold one certificate at a time and one of which was caused by a customer's admin pasting the certificate with the PEM header intact into a field that didn't want it. When Q2's identical ticket opened, eleven counterparties from Q1 still hadn't switched, and the old certificate was still trusted, which is to say the Q1 rotation had not actually finished when the Q2 rotation was supposed to begin.
The same policy. The same cadence. The same ticket. Two operations whose costs differed by more than two orders of magnitude and whose elapsed times differed by a factor of eight thousand.
And the coda, which is the part I think about most: the key that actually leaked that year was an API key compiled into a mobile app binary shipped in 2021. It was on no checklist at all, because the rotation register had been built by listing the things the platform could rotate. A key with no rotation path generates no rotation task, so it silently disappears from the document whose entire purpose is to make sure nothing is forgotten.
This is the credential-lifecycle chapter of a series doing arithmetic on identity infrastructure. The Hidden Cost of Audit Logs set the method; The Cost of Session Storage found a memory bill that was really a write bill; The Cost of Password Hashing priced a security control as a deliberate purchase; The Cost of Token Introspection vs Self-Validating Tokens found payroll dominating both options; The Cost of SCIM Synchronization found the cost was never infrastructure.
The mechanics of rotating a key correctly are covered in Designing for Key Rotation Before You Need It — the four mechanisms, the phased timeline, why the overlap exists. Assume all of it. This article is about what each of those operations costs, why the answer varies by five orders of magnitude across key classes that a compliance document treats as identical, and what you should actually buy with a rotation budget.
The short version, up front, because it takes the rest of the article to defend: a graceful scheduled rotation is not a revocation, so the security benefit most organizations think they're buying with cadence isn't the benefit they're getting. What cadence buys is rehearsal. And almost everyone rehearses the operation that already works.
The assumptions, published so you can re-run it
| Parameter | Value | Note |
|---|---|---|
| Enterprise tenants | 200 | Same scenario as the rest of the series |
| Identities under management | 600,000 | |
| SAML federations | 190 | Each with its own IdP administrator |
| Registered OAuth clients | 3,000 | 2,400 internal, 600 customer-held |
| Customer-held API keys | 500 | |
| Validating services / pods | 40 services, 200 pods | Per the introspection piece |
| Token validations | 12,000/s | |
| Access token lifetime | 15 minutes | |
| Consumer JWKS cache TTL, observed p99 | 24 hours | Plus an unbounded tail — see below |
| Fully-loaded engineer | $96/hour | $200k/yr |
| vCPU | $25/vCPU-month | |
| Managed KMS asymmetric key | $1/key-month, $0.03 per 10k ops | AWS-shaped |
| Counterparty change lead time | p50 3 weeks, p95 14 weeks, p99 26 weeks | Enterprise change management |
| Counterparty failure rate, per round | 3% | They miss it, or their IdP can't hold two |
| Modelled cost of one broken federation | $2,500 | Incident + support + relationship |
| Modelled cost of a single-tenant signing-key compromise | $250,000 fixed, plus $8,000/day | Derived below |
Two of those deserve flagging now. The consumer cache TTL p99 is the number that sets your overlap window, and it has a tail you cannot measure: some validator libraries fetch JWKS once at process start and never again, which makes their effective TTL "until someone restarts that pod," which could be November. And the counterparty lead time p99 is the number that decides whether a cadence is even arithmetically possible, which turns out to be the most under-used figure in this entire subject.
Every key class, priced
The core claim of this article is that "rotate everything every 90 days" is a category error, because it assigns one price to operations whose costs are not in the same universe. Here is the enumeration, for a realistic identity platform.
| Key class | Count | Parties who must act | Discovery mechanism | Marginal cost per operation |
|---|---|---|---|---|
| TLS server certificates (ACME) | 40 | 1 — you | TLS handshake | ~$0.02 |
| Internal client credentials via private key JWT | 2,400 | 1 — the client | Client-published JWKS | ~$0.50 |
| Per-tenant OIDC signing keys | 200 | 1 — you | Your JWKS | ~$1 |
| Session / cookie encryption key | 1 | 1 — you | None; key ring in every node | ~$5 with a ring |
| KEK over envelope-encrypted data | 1 | 1 — you | None; re-wrap | ~$50 |
| Customer-held client secrets | 600 | 2 | None | ~$180 |
| Customer-held API keys | 500 | 2 | None | ~$220 |
| SAML signing certificate (yours) | 1 → 190 relationships | 191 | Metadata, for ~60% of them | ~$326 per relationship |
| mTLS client certs to vendor endpoints | 14 | 2, via a vendor portal | None | ~$400 |
| Data encryption key with no envelope | 1 | 1 — you | None; rewrite the data | ~$7,700 |
| HSM / KMS root | 2 | 3+ custodians, plus a witness | N/A | ~$8,000 |
| Key compiled into a shipped mobile binary | 1 | Every install, forever | None | unrotatable |
Five orders of magnitude between the top of that table and the bottom, before you reach the row where the cost is undefined.
Three of those rows deserve their arithmetic shown, because they're the ones people get wrong.
Per-tenant OIDC signing keys, ~$1. Keygen for P-256 is a fraction of a millisecond. The real cost is KMS rent, and the interesting part is that it doubles during the overlap: 200 keys at $1/key-month is $200, and holding two keys per tenant for a 30-day overlap inside a 90-day period raises the average population to about 266 keys, so $266/month. The $66/month delta divided by 200 rotations per quarter is about a dollar apiece. There is no human in this operation at any point, which is why its cost has a floor measured in cents.
SAML signing certificate, ~$326 per relationship. One certificate, 190 bilateral changes. Per counterparty: roughly 1.5 engineer-hours to identify the right contact, send a notice with the actual thumbprint and dates, track it, and verify the switch happened — $144. Plus the expected cost of breakage: 3% × $2,500 = $75. Plus program management across a round that runs for months: about 0.2 FTE for six months, $20,000, spread over 190 relationships is another $105. Total $324, call it $326.
Data encryption key with no envelope, ~$7,700. Rotating the key that directly encrypts 4 TB of stored data means reading, decrypting, re-encrypting and writing 4 TB. The CPU is nothing — AES-GCM at a couple of GB/s per core is about 2,000 CPU-seconds, five cents. The cost is that you are rewriting your entire user table safely, online, with rollback, which is two engineer-weeks of build and supervision plus a week of elevated write load on your most important store. Compare it to the row above: re-wrapping 600,000 envelope DEKs under a new KEK is 1.2M KMS calls at $0.03 per 10,000, which is $3.60, plus a batch of row updates.
That comparison is worth pausing on, because it contains the whole theory of this article. Envelope encryption is not primarily a security technology. It is a rotation-cost-reduction technology. It inserts a level of indirection so that the thing you rotate is small and the thing that is expensive to touch is not the thing you rotate. It converts a $7,700 operation into a $50 one, permanently, at the cost of one extra hop on every decrypt.
The variable that actually sets the price
Look back at the table and ask what predicts the cost. It isn't the algorithm. It isn't the key length. It isn't how the key is stored, and it isn't how important the key is.
It's two questions:
- How many parties must take an action for the change to complete?
- Is there an indirection between the key and the parties who trust it — and does that indirection have a discovery mechanism?
JWKS is an indirection with discovery: relying parties trust a URL, not a key, so you can change the key without renegotiating trust. Envelope encryption is an indirection without discovery, but with a single party, which is enough. A SAML certificate pasted into a customer's IdP configuration is neither: the counterparty trusts the key material itself, holds a copy, and there is no channel by which they learn you've changed it. A key compiled into a mobile binary is the limiting case — the party holding it is a shipped artifact that cannot be told anything, ever.
flowchart TD
K["A key you want<br/>to rotate"] --> Q1{"Can you complete<br/>the change alone?"}
Q1 -->|no| B["Coordinated class<br/>floor: one human<br/>interaction per counterparty<br/>~$150–400 each"]
Q1 -->|yes| Q2{"Do consumers<br/>discover it, or hold<br/>their own copy?"}
Q2 -->|"hold a copy"| C["Distribution class<br/>cost = every place the<br/>key currently lives"]
Q2 -->|"discover it"| A["Automated class<br/>cents per operation<br/>no human in the loop"]
B --> Q3{"Can the counterparty<br/>be reached at all?"}
Q3 -->|no| Z["Unrotatable<br/>cost is undefined;<br/>the only lever is scope"]
The practical consequence is that rotation cost is a property of your implementation, not of your key. The clearest demonstration is the session cookie encryption key. If your runtime holds a key ring — decrypt with any key in the set, encrypt with the newest — rotating it costs about five dollars and nobody notices. If your runtime holds a single key read from config, rotating it invalidates every cookie in existence simultaneously and you have bought a mass logout.
That mass logout is worth pricing, because its cost is not where people look. A million forced re-authentications is about $1.40 of hashing CPU and $4.80 of audit ingest. Trivially cheap. What it actually costs is the peak: a million logins arriving inside ten minutes is 1,667/s against a pool sized for a 90/s peak, and scale-up takes three to six minutes while the herd arrives in seconds. You either hold eighteen times your normal capacity permanently, or you accept twenty minutes of degraded authentication as part of a routine key rotation. Same key, same cadence, same policy line in the standard — and a cost that varies by five orders of magnitude depending on whether one data structure in your runtime is a key or a set of keys.
The coordinated class is not expensive. It's infeasible.
The $326-per-relationship figure understates the problem, because the binding constraint on the coordinated class is not money. It's elapsed time, and elapsed time behaves badly when N is large.
Breakage becomes certainty. At a 3% per-counterparty failure rate, the probability that at least one of 190 federations breaks during a round is 1 − 0.97^190 = 99.7%, with an expectation of 5.7 breakages. This reframes the whole conversation. You are not asking whether a rotation round might cause an outage. You are choosing how many outages to buy, and the answer scales linearly with your customer count while the benefit does not.
The tail sets your maximum cadence. A round is not complete when the median counterparty finishes; it's complete when the last one does. With 190 independent counterparties, the probability that at least one exceeds its own 99th-percentile lead time is 1 − 0.99^190 = 85%. If that p99 is 26 weeks — which is unremarkable for a large enterprise with a quarterly change window and a December freeze — then a rotation round has an 85% chance of taking longer than six months.
A 90-day mandate against a population whose rounds take six months does not produce quarterly rotation. It produces permanently overlapping rounds, an old certificate that is never retired because the previous round never closed, and an audit answer of "yes, quarterly" that is not true. This is the arithmetic reason the SAML row in the opening story was still open when the next ticket arrived, and no amount of process discipline fixes it, because the distribution belongs to somebody else's change advisory board.
Here is the same idea stated as the rule I'd want on a wall: the feasible cadence for a coordinated key class is set by the p99 of your counterparties' change lead time, multiplied by nothing and divided by nothing. For 190 enterprise federations, the first feasible cadence is annual. Everything shorter is a fiction that shows up in your audit evidence as an exception you'll argue about later.
The overlap window costs more than you think, and it has a floor
Everyone knows rotation needs an overlap. Almost nobody prices the overlap or notices that its length is set by a variable in someone else's budget.
The window has a hard lower bound:
W ≥ max consumer JWKS cache TTL
+ longest lifetime of any artifact validated by signature
+ clock skew tolerance
For the scenario above, that's 24 hours of cache plus 15 minutes of access token plus skew — call it 26 hours, with an unbounded tail for the validator that never refetches. But look at what happens if any signature-validated artifact is long-lived. If your refresh tokens are self-validating JWTs with a 30-day lifetime, your overlap window is 31 days, because retiring the key before then invalidates a month's worth of refresh tokens. Artifact lifetime and rotation cost are directly coupled, and nobody budgets the coupling. The team that extended refresh token lifetime from 14 to 90 days for a mobile UX improvement tripled the cost of every future signing key rotation and moved the exposure floor from two weeks to three months, in a ticket that mentioned neither.
Now the result that I think is genuinely non-obvious. The number of simultaneously valid keys in steady state is:
K = 1 + ceil(W / P)
…where P is the rotation period. Which means rotating faster does not shrink your set of valid keys — past a threshold it grows it:
| Overlap W | Period P | Simultaneously valid keys | Per-tenant KMS rent at 200 tenants |
|---|---|---|---|
| 26 h | 90 days | 2 | $266/month |
| 26 h | 30 days | 2 | $270/month |
| 26 h | 7 days | 2 | $290/month |
| 31 days | 90 days | 2 | $266/month |
| 31 days | 30 days | 3 | $400/month |
| 31 days | 7 days | 6 | $1,000/month |
| 31 days | 1 day | 32 | $6,400/month |
The bottom rows are absurd on purpose, but the direction is real and it's the honest version of the folk claim that frequent rotation is "worse than linear." It is worse than linear, in two specific ways.
Key population. At P < W, every additional rotation adds a key to the valid set rather than replacing one. Your JWKS documents grow, every validator's key cache grows by the same factor, and a validator with a fixed-size cache across 200 tenants starts evicting keys it will need in a second. Rotating a per-tenant key daily with a 31-day overlap means 32 keys per tenant × 200 tenants = 6,400 entries in every one of 200 pods.
Fetch amplification at the boundary. Steady-state JWKS traffic is set by cache TTL, not by rotation period — 200 pods × 200 tenants ÷ 24 hours is about 0.5 fetches per second, noise. The cost lands entirely in the propagation gap. A validator that sees an unknown kid fetches synchronously, and if it lacks negative caching, every in-flight request for that tenant fetches. One tenant at 50 rps with a 2-second gap is 100 fetches — fine. Rotate all 200 tenants in the same maintenance window and it's 20,000 fetches in two seconds, roughly 10,000 rps against an endpoint that normally serves 0.5. If that endpoint sheds load, the gap lengthens, which produces more fetches, which lengthens the gap. That's the feedback loop, and it's the actual mechanism behind "rotation caused an outage" in systems where the rotation itself was correct.
The fix is boring and costs nothing: stagger tenant rotations rather than batching them, and require negative caching plus a rate limit on the unknown-kid path in every validator you own. But notice what the batched version really is — an operator scheduling all 200 rotations in one window because that's how you'd schedule work that has a human in it. The stampede is what happens when a cheap automated operation inherits the scheduling habits of an expensive coordinated one.
The last cost of the overlap is the one that decides the security argument, so it gets its own section.
The uncomfortable part: graceful rotation is not revocation
Here is the justification that appears in nearly every rotation policy, in some phrasing: periodic rotation limits the window during which a compromised key is useful to an attacker.
Now put that sentence next to the mechanism. A phased rotation deliberately keeps the old key published and trusted for the entire overlap window, because that is the only way to avoid an outage. During the overlap, a token signed with the old key validates. After the overlap, the old key stops being accepted — but every token minted with it before then remains valid until it expires, and an attacker who held that key for even one second minted whatever they needed on the first day.
So a graceful scheduled rotation does not remove an attacker. By construction, it cannot: the property that makes it non-disruptive is precisely the property that makes it non-revoking. The only rotation that revokes is the abrupt one — the emergency path, with a hard cutover and everyone logged out, which the mechanics article is right to treat as a categorically different operation.
Which leaves scheduled rotation defending exactly one scenario: a compromise you never detect, where the attacker's access path has also closed on its own, so they cannot simply take the new key too. That's a real scenario. It's narrower than the policy language implies, and it has a probability term nobody estimates.
Let's estimate it anyway, because the alternative is arguing about it with adjectives.
Model a per-tenant signing key compromise as $250,000 of fixed cost — incident response, forced re-authentication, notification, legal, credits — plus $8,000 per day of continued attacker access, which represents additional accounts reached, additional data in scope, and the widening notification boundary. Take a compromise rate of λ = 0.08/year across the key population, detection in 60% of cases at a mean lag of 3 days, and put the probability that an undetected attacker's access path self-closes at 0.3.
| Posture | Undetected branch (40%) | Detected branch (60%) | Expected annual rate-cost |
|---|---|---|---|
| Annual rotation, no emergency path | 182 d × 0.3 effective → $436k | (3 d + 3 d) × $8k = $48k | $16,300 |
| Quarterly rotation, no emergency path | 45 d × 0.3 effective → $108k | $48k | $5,760 |
| Annual rotation, 20-minute emergency path | $436k | (3 d + 20 min) × $8k = $24k | $15,100 |
| Quarterly rotation, 20-minute emergency path | $108k | $24k | $4,610 |
Read that honestly and it does not say "cadence is worthless." For the automated class, going from annual to quarterly is the single largest term in the table, and it costs about $800 a year in KMS rent. Buy it. That's the correct conclusion when the marginal cost of a rotation is a dollar.
What the table also says, and what matters more, is where the leverage is not. The gap between rows 2 and 4 — the value of an emergency path, priced purely in exposure-days — is about $1,150 a year, against a build cost of three engineer-weeks plus quarterly rehearsal. On exposure-days alone, latency loses.
Exposure-days are the wrong metric, and this is the part I most want to argue.
An organization that cannot hard-rotate quickly does not merely rotate three days later. It hesitates, and the hesitation changes the shape of the incident rather than its length. The conversation I've watched happen, more than once, goes: "we could invalidate the key now, but that logs out every tenant and we've never done it, so let's contain it another way and rotate in the maintenance window." That decision converts a bounded, single-tenant, containable event into an unbounded one — which moves the $250,000 fixed term, not the $8,000/day rate term, and moves it by a multiple. The value of low rotation latency is not that the exposure is shorter. It's that the containment option is actually on the table during the meeting where someone decides what to do.
And the fixed term is where the money is. If having a rehearsed 20-minute path changes the fixed cost of one event in four from $250,000 to $60,000 — because the blast radius stayed at one tenant instead of becoming a platform-wide notification event — that's 0.08 × 0.25 × $190,000 = $3,800/year of expected value, from a capability that also covers every key class where cadence is not purchasable at any price.
Which is the real argument, and it doesn't need the numbers to be precise:
Cadence is only purchasable for the automated class. Latency is purchasable for every class. So when the marginal cost of a rotation is near zero, buy frequency; when it is high, buy latency — and notice that most organizations do the exact opposite.
They mandate quarterly rotation of the coordinated class, where it's infeasible and low-value, and they have never once rehearsed an emergency rotation of anything.
The automation break-even, and why it's the wrong question
Standard reasoning: automate when the build cost divided by the per-round saving is a payback period you can live with. Run it for the automated class.
Manual rotation of 200 per-tenant keys, 5 minutes of babysitting each = 16.7 h = $1,600/round
Quarterly = $6,400/year
Automation build: 3 engineer-weeks = $11,520
Automation maintenance: 2 h/month = $2,304/year
Payback = $11,520 ÷ ($6,400 − $2,304) = 2.8 years
Nearly three years. By the usual rule, that's a marginal investment you'd defer for something with a faster return — which is exactly why rotation automation sits in the backlog at most organizations, and why the business case, argued on labour savings, keeps losing.
The business case is being argued on the wrong quantity. You do not automate rotation to save the sixteen hours. You automate rotation because automation is what makes the emergency path exist, and the emergency path is the thing that was worth buying in the first section. A manual runbook executed under pressure at 02:00 by whoever is on call is not a 20-minute capability at any level of documentation quality. The labour saving is a rounding error attached to a capability purchase.
Now run the same analysis on the coordinated class, where it produces a different and sharper answer.
What is there to automate in a SAML certificate rotation? Not the counterparty's change process. Not their change advisory board. You can automate your side of the tracking, and you'll save perhaps 20 minutes of the 90 per relationship. The other 70 minutes are on the far side of a company boundary and no engineering of yours reaches them.
There is exactly one automation that works, and it isn't an automation of the rotation — it's an automation of the trust relationship. Scheduled metadata refresh with add-only semantics converts a counterparty from "holds a pasted copy of your certificate" to "discovers your certificate from a URL," which moves them from the coordinated class to the automated class permanently. Price it:
Build: metadata fetch, diff, add-only, signed-metadata validation,
last-known-good on failure = 4 engineer-weeks = $15,360
Applies to the ~60% of counterparties who publish usable metadata
Cost per round before: 190 × $326 = $61,940
Cost per round after: 76 × $326 = $24,776
Saving per round = $37,164
Payback in well under one round. But notice the ceiling: your automation investment in the coordinated class is bounded by your counterparties' capabilities, not by your engineering. The remaining 76 relationships cost $326 every round forever, and no amount of money spent on your side changes that number by a cent.
Which produces the single most actionable line in this article, and it's a scheduling insight rather than a technical one. A rotation round is the one time each year when you have a named administrator at each of 190 customers actually engaged with your certificate configuration. That coordination is already paid for. Spend it on eliminating future coordination. While you have them, get the metadata URL, confirm their IdP can hold two certificates, and record which ones can't. Paying $326 once to move a counterparty out of the coordinated class beats paying $326 every round, and the marginal cost of asking is a sentence in an email you were already sending.
Most rotation rounds are run as pure operations, close their tickets, and leave the population exactly as expensive as they found it.
Per-tenant keys: multiplying operations to divide blast radius
Worth pricing explicitly, because it's the design decision where the two axes of this article pull in opposite directions.
A shared platform signing key gives you one rotation operation per round. Per-tenant keys give you two hundred. If each operation cost a hundred dollars, that would settle it. Each operation costs about a dollar, which changes the answer entirely.
| Shared platform key | Per-tenant keys (200) | |
|---|---|---|
| Rotation operations per round | 1 | 200 |
| Marginal cost per round | ~$2 | ~$200 |
| KMS rent including overlap | ~$2/month | ~$266/month |
| JWKS cache entries across the fleet | 200 | 40,000 |
| Blast radius, tenant-scoped compromise | all 200 tenants | 1 tenant |
| Blast radius, platform-scoped compromise | all 200 tenants | all 200 tenants |
| Emergency rotation | platform-wide mass invalidation | one tenant, routine |
| Realistic time-to-untrust under duress | hours to days | minutes |
Two things in that table, and the second is the important one.
The blast-radius rows need an honest caveat that the marketing version of this argument usually omits. Per-tenant keys reduce blast radius only against tenant-scoped compromise vectors — a scoped export bug, a support engineer's mistake, a per-tenant backup landing somewhere it shouldn't. They do nothing whatsoever against a platform-scoped vector, because whatever compromised the process that holds one signing key holds all two hundred. The isolation is real and it is narrower than it sounds.
The row that carries the real value is the last one. Per-tenant keys convert an emergency rotation from a platform-scale event into a tenant-scale routine. A shared-key emergency rotation logs out every customer, so it requires a decision nobody wants to make alone, which is why it takes days. A per-tenant emergency rotation affects one customer, which means an on-call engineer can make it at 02:00 without waking a VP — and that is precisely the containment option we priced two sections ago. The isolation argument for per-tenant keys is usually made on blast radius. The stronger argument is on rotation latency, and it's the one nobody makes.
I should disclose an interest: ClavionX, the platform I work on, issues per-tenant OIDC issuers and per-tenant signing keys, so this is a trade-off I have opinions about rather than a neutral survey. The bill for that choice is the 40,000-cache-entry row and the tenant-multiplied JWKS coherence problem, which the introspection economics piece already prices and I won't re-derive. It's also worth being precise about what per-tenant keys don't buy: in a design where the runtime works only from projected state and fails secure when that state is missing, a validator that hasn't yet received a new key rejects rather than reaching back for it — which is the correct behaviour, and it means propagation time is a real availability parameter rather than something the runtime can paper over. That is a cost, and it belongs in the overlap-window arithmetic above.
The keys you cannot rotate
Back to the mobile binary from the opening.
Every organization has some. A key compiled into a shipped application. A secret hardcoded by a partner whose integration was built by a contractor in 2019 and whose vendor has since been acquired. A trust anchor in the firmware of an appliance whose manufacturer is end-of-life. A client secret embedded in a customer's compiled connector.
The register never lists them, for the structural reason named at the top: rotation registers are built by enumerating rotation tasks, and a key with no rotation path generates no task. Your inventory of key material is silently filtered to exclude precisely the entries with the worst risk profile.
Price one. An API key in a mobile binary, 3% of 400,000 installs on devices that will never take an update — 12,000 devices. You have three options and all of them cost money:
| Option | Cost | What you're actually buying |
|---|---|---|
| Keep trusting it | λ × full blast radius, with exposure window = ∞ | Nothing. You have pre-purchased the entire loss at whatever probability the key eventually leaks. |
| Force-retire it | Churn on 12,000 accounts plus a support wave | A closed exposure, paid for in customers |
| Run a scoped-down legacy path | ~0.1 FTE forever = $20,000/year | Time, indefinitely, at a fixed annual price |
Most organizations choose the third without ever deciding to, which is how a key in a binary becomes a permanent operating expense. The general form is worth stating because it applies to every row of that class:
When rotation cost is infinite, the only substitute good is scope reduction. You cannot change the key, so change what the key can do: narrow its scopes, bind it to a network path or a client certificate, rate-limit it hard, make it useless outside a single endpoint, and put an expiry on the population rather than on the key.
And the preventive rule that follows: the moment you are about to put key material somewhere it cannot be retrieved from — a binary, a partner's compiled artifact, firmware — you are choosing an infinite rotation cost, and that decision deserves the same scrutiny as choosing an unrecoverable data store. It almost never gets it, because it happens in a build script.
A per-class cadence framework you can defend to an auditor
The control objective an auditor is actually testing, once you strip the 90-day number off it, has two halves: key material of unknown or excessive age is not trusted, and any key can be untrusted within a bounded time. The first is cadence. The second is latency. A policy that specifies only the first is answering half the objective, and it's the cheaper half.
So write the policy per class, with both numbers, and with the reason:
| Class | Cadence | Time-to-untrust target | The sentence that defends it |
|---|---|---|---|
| Automated, discoverable (TLS, JWKS-backed signing keys, private key JWT clients) | 30–90 days | < 30 min, per tenant | Fully automated, no human in the path, cost per operation under $1 — so cadence is bought at its cheapest and rehearsed continuously |
| Unilateral, non-discoverable (cookie keys, KEKs) | 90–365 days | < 2 h | Key-ring or envelope construction means rotation is non-disruptive; rehearsed twice a year |
| Coordinated, bilateral (SAML, customer secrets, API keys, vendor mTLS) | Annual, or on evidence | < 24 h for a single relationship | Counterparty p99 change lead time is 26 weeks; a shorter cadence cannot complete. Compensated by overlap support, expiry monitoring, and per-relationship emergency revocation |
| Ceremonial (HSM/KMS roots) | 3–5 years, or on custodian change | < 8 h, documented ceremony | Rotation requires a quorum ceremony; subordinate material is re-wrapped rather than re-issued |
| Unrotatable | N/A — register separately | N/A | Scope-reduced, rate-limited, path-bound, with a dated retirement plan |
Four things make this land in an audit conversation rather than start an argument, and they're the same four that make the compensating-control case for password expiry work:
Name the control objective, not the standard. You are meeting "key material is not trusted beyond a defined age, and can be untrusted within a bounded time." You are meeting it differently per class because the classes have different mechanisms.
Bring evidence for the latency half. This is the piece almost nobody has, and it's the piece that makes the cadence argument winnable: a rehearsal log, with dates, of emergency rotations performed in a lower environment, per class. An auditor who can see that you rotated a signing key under simulated duress in 18 minutes last month is being handed a workpaper. "We rotate quarterly" without that is a schedule, not a capability.
Publish the unrotatable register. It reads as bad news and lands as rigour. A team that can name its four unrotatable keys, their scopes, their compensating controls and their retirement dates is demonstrably managing key material. A team whose register is complete and clean has, in nearly every case, simply not looked.
Be able to state your overlap window per class, because it is the floor on your exposure and nothing you do to the cadence goes below it. If your exposure floor is 31 days because of a refresh-token lifetime decision, say so — and note that shortening the artifact is the only way to move it, at the refresh-amplification price the introspection piece computes.
And rehearse the expensive path, not the cheap one
One more consequence, because it falls straight out of the argument and inverts common practice.
If the security value of graceful scheduled rotation is smaller than the policy language claims — and I've argued it is, because graceful rotation doesn't revoke — then a large part of what cadence genuinely buys is proof that the path works. That's real value. A rotation path that has never run does not work; every operator knows this, and it's the strongest argument in the mechanics article.
But then look at what gets rehearsed. Quarterly rotation of the automated class exercises a cron job that has run 200 times this year and has never failed. It rehearses the operation with the lowest probability of failure and the smallest blast radius when it does fail. Meanwhile, the cookie encryption key rotation that would log out a million users, the emergency single-tenant signing key revocation, the SAML rotation with the customer whose IdP holds one certificate, the KEK re-wrap over 600,000 DEKs — those get rehearsed never, and they are the entire population of operations whose failure would matter.
If cadence is rehearsal, spend the rehearsal budget where the uncertainty is. Two rehearsed emergency rotations a year, in a lower environment, of the two classes that would actually hurt, cost about sixteen engineer-hours — $1,536 — and produce the only evidence in this entire subject that's worth anything: a date, a stopwatch time, and a list of what broke.
What to measure
Six numbers. Most teams running an identity platform know none of them.
- Your overlap window
W, per key class, derived rather than assumed — longest consumer cache TTL plus longest signature-validated artifact lifetime plus skew. It is the floor on your exposure and no cadence goes below it. K = 1 + ceil(W / P): your count of simultaneously valid keys. If it's above 2, you are accumulating keys rather than rotating them, and rotating more often is making that worse.- The p50, p95 and p99 of counterparty change lead time, from your own rotation history. The p99 is your maximum feasible cadence for the coordinated class, and it is the number that turns a policy argument into an arithmetic one.
- Measured time-to-untrust, per class, from a rehearsal with a date on it. Not the runbook's estimate. The stopwatch.
- Your class-B-to-class-A conversion rate — the fraction of counterparties on discovery (metadata URL, JWKS) rather than a pasted copy. This is the only number in the coordinated class that ever improves, and only if someone is deliberately moving it.
- The unrotatable register, with scope and a retirement date for each entry. If it's empty, it's wrong.
The closing thought
The team in the opening story eventually rewrote the standard. Not to remove rotation — to split one line into five, one per class, each with a cadence, a time-to-untrust target, and a sentence explaining why. The automated class went to 30 days, because it costs a dollar. The SAML class went to annual, with metadata refresh built and a program to collect metadata URLs during the round they were running anyway. Two emergency rotations a year got put on the calendar as rehearsals, with named owners and a stopwatch.
The security posture measurably improved and the annual cost of the rotation program fell by about two thirds, which is the shape of result you get whenever a uniform price is applied to a non-uniform population and someone finally separates them.
The uncomfortable general lesson is the same one that keeps turning up in this series. A single number in a policy document — 90 days, seven years of retention, 15-minute tokens — is nearly always a price applied to a population whose costs differ by orders of magnitude. The number isn't wrong. It's unpriced, which means it will be either ignored or ruinously expensive, and which of those you get depends on nothing but which key class you happened to be looking at when you wrote it down.
Rotation is worth doing. Most of what it's worth, though, isn't the rotation. It's being able to do it in twenty minutes on the day you find out you have to.