The Deploy That Logs Everyone Out
The change was a version bump in a Helm chart. Patch release, no schema, no migration, no feature flag. It went out at 14:05 on a Thursday because Thursday afternoon was the boring window.
At 14:06 the first pod cycled and a small number of users were bounced to the login screen. Nobody noticed, because a small number of users are always bounced to the login screen. At 14:08, with four of twelve pods rolled, support had six tickets and the channel had a theory: something flaky, probably the load balancer, probably transient. At 14:11 the rollout completed and every session in the estate — four hundred thousand of them — was gone at once.
The cause was three lines in a Kubernetes secret template. The chart generated the cookie-encryption key with randAlphaNum 32 and no lookup guard, so every helm upgrade minted fresh key material and wrote it over the old. The key had never been in a secret manager. It had never been rotated deliberately, because nobody knew it existed to rotate. It had survived eleven previous deploys only because those eleven deploys had gone out through a path that didn't re-render that template.
Then somebody rolled back, which re-rendered the template, which generated a third key. The rollback logged out everyone who had managed to log back in during the intervening six minutes.
By 14:20 the platform was not down in any sense a health check could detect. Every pod was ready. Every dependency was green. The error rate at the load balancer was flat, because 401 is a successful response. What was actually happening was that four hundred thousand people were trying to authenticate at the same time into a password-verification pool sized for ten logins a second, and the queue in front of it had a drain time of roughly two hours.
This article is about that shape: what actually causes a deploy to invalidate every live session, why the second-order failure is reliably worse than the first, how to make session state survive a deploy on purpose, and — the part people skip — how to tell in the first sixty seconds whether you are in an outage or in the middle of something you meant to do.
A scope note, because there is a neighbouring article and I don't want to re-run it. Blue-Green Deploying an Identity Platform covers the compatibility discipline: readers before writers, version bytes in the envelope, expand-and-contract migrations, publishing keys before you sign with them, and the honest question of whether rollback is available at all. That's the how not to break it half. This is the other half — the specific mechanisms that produce a total invalidation rather than a partial one, and the arithmetic of what happens next.
Little's Law is why the herd is unsurvivable
Start with the parameter that governs everything and that almost nobody connects to capacity: session lifetime.
Take a workforce SaaS with 400,000 concurrent sessions at peak. Sliding idle timeout of eight hours, absolute cap of twenty-four, which works out to a mean session lifetime of about six hours before the user has to establish a new one.
Little's Law gives you the login rate for free. If the population P is 400,000 and the mean time in system W is 21,600 seconds, then the arrival rate is:
λ = P / W = 400,000 / 21,600 ≈ 18.5 new sessions/second
Daily average 18.5/s; the Monday-morning peak runs about 3× that, call it 55/s. But not all 55 are equal. Most arrivals at the authorization endpoint are silent — an existing IdP session satisfies the request, or a refresh token is exchanged, or a remembered-device record short-circuits the challenge. Those cost a few milliseconds. Only the residue is a full credential verification.
| At peak hour | Rate | Unit cost | Cores |
|---|---|---|---|
| Silent re-auth (live IdP session, refresh, remembered device) | 45/s | ~3ms | 0.14 |
| Full credential verification (bcrypt cost 12 ≈ 250ms) | 10/s | 250ms | 2.5 flat out |
| Provisioned hashing pool at 40% target utilization | ~6 cores |
Six cores. That is the entire credential-verification capacity of a platform serving four hundred thousand people, and it is not under-provisioned — it is correctly provisioned, because sessions are long and long sessions amortize the expensive path across hours of cheap ones.
Which is the trap, stated precisely:
Your hashing pool is small because your sessions are long. Which makes the hashing pool exactly the component that cannot absorb your sessions ending.
Now invalidate all 400,000. Say 40% of the population is actively in the product and re-enters within ten minutes — a conservative number for a workday tool people keep open in a tab:
160,000 re-authentications / 600 s = 267 credential verifications/second
267/s against a pool sized for 10/s. That is 27× the provisioned capacity, and it is worse than a 27× traffic spike, because the workload mix inverted at the same time: the cheap path that carried 82% of normal arrivals no longer exists. There are no live IdP sessions to satisfy silently. There are no valid refresh tokens if the invalidation touched signing material. Every single arrival is now the expensive one. This is the phenomenon The Art of Load Testing makes the case for testing — mix shifts faster than volume — in its most violent form.
The drain time is the number that ends the argument. Six cores at four verifications per core-second is 24/s of real capacity:
160,000 verifications ÷ 24/s = 6,667 seconds ≈ 1 hour 51 minutes
Nearly two hours to serve a queue whose clients time out in ten seconds. Every one of those requests will be abandoned before it is served, and the server will hash it anyway, because a queued request doesn't know its socket closed. That is goodput collapsing to zero, and from there the retry storm makes it self-sustaining in exactly the way Why Identity Systems Fail Gradually, Then All at Once describes. I won't re-derive the metastability; I'll only note that the trigger here is not a failover or a burst of attack traffic. You did this to yourself, deliberately, through a change-management process, with an approval on it.
The lever you don't have
Here is the part that surprises people during the incident: you cannot turn down the work factor to get out of it.
The cost of verifying a password is determined by the parameters embedded in the stored hash, not by your configuration. Lowering bcrypt.cost or Argon2's m and t changes what you write on the next successful login — it does nothing whatsoever to the 160,000 verifications currently queued against hashes written at the old parameters. There is no runtime dial that makes credential verification cheaper. It is the one line item on the login path that is deliberately, structurally, irreducibly expensive (What Does One Login Actually Cost You? puts a number on it).
Which means during a mass re-authentication your only levers are admission control and capacity, and capacity arrives cold.
The second-factor failure runs on a different clock
Now add MFA, because the herd hits it too, and this one produces a genuinely self-sustaining loop that most teams have never modelled.
Suppose 30% of that population is on SMS OTP — a share that refuses to go away whatever your roadmap says. 48,000 messages in ten minutes is 80/s into a gateway provisioned, quite reasonably, at 50/s.
The queue grows at 30/s. After ten minutes it holds 18,000 messages, and delivery latency is 18,000 ÷ 50 = 360 seconds. Your OTP validity window is five minutes.
Every code now arrives after it expires. Users who receive one and type it in are rejected, so they press "resend," which appends another message to the queue, which lengthens the delay, which guarantees the next code also arrives dead. The SMS path has entered its own metastable state, running on entirely different infrastructure from the one you're staring at, and it will not recover when the hashing pool does. The tell in your telemetry is a rising ratio of otp_verify_expired to otp_verify_ok with a falling absolute success count — and if you only alert on OTP delivery success rate, the gateway will report that it delivered everything perfectly.
And the herd comes back
The last piece of arithmetic is the one I have never seen written down, and it explains a class of mystery spike that appears days after an incident is closed.
A mass invalidation doesn't only cost you one herd. It phase-locks your session population. Before the event, session start times were spread across the working day. Afterwards, 400,000 sessions were all established inside the same twenty minutes — which means they will all reach their absolute expiry inside the same twenty minutes, six hours later.
If session lifetime were a constant, that pulse would repeat forever. It isn't constant, so it spreads: if the effective lifetime has a standard deviation σ, the pulse width after k cycles is roughly σ√k, and the peak amplitude decays like 1/√k.
| Cycle | Elapsed | Pulse width (σ = 1h) | Peak rate |
|---|---|---|---|
| 1 | 6h | ~1h | ~110/s |
| 4 | 24h | ~2h | ~55/s |
| 9 | 54h | ~3h | ~37/s |
| 36 | ~9 days | ~6h | indistinguishable from baseline |
With an hour of natural variance it takes about a week for the echo to wash out. With a rigid absolute timeout and no idle-driven variation — which is what a strict compliance-driven session policy produces — σ is small, and you have installed a recurring six-hourly spike into your own traffic that outlives the incident review.
The mitigation is cheap and worth doing before you need it: jitter absolute session lifetimes by ±15–20% at creation. It costs nothing, it is invisible to users, it de-synchronizes the population continuously, and it converts a resonant system into a damped one. It is the same argument the load-testing article makes about token expiry, applied one layer up and with a decay rate you can calculate.
What actually causes it
The arithmetic above is the consequence. Here are the mechanisms, roughly in order of how often I have seen each one be the actual cause.
1. A secret that is born with the process. The single largest category. A framework generates session-signing or cookie-encryption material at startup when none is configured, and the default is a working default, so nothing ever fails in development. Then the process boundary changes shape underneath it.
| Where it hides | What triggers the invalidation |
|---|---|
| ASP.NET Core Data Protection key ring on the container filesystem | Any new image or pod — the ring is ephemeral, so every auth cookie dies with the container |
Rails secret_key_base generated into ephemeral credentials |
Regenerated credentials, or a container without the persisted file |
Flask SECRET_KEY / Express session({secret}) set from a random default |
Every restart, and independently per replica |
Helm randAlphaNum in a secret template without a lookup guard |
Every helm upgrade, including the rollback |
Terraform-managed KMS key or secret version with a create_before_destroy gap |
Any plan that replaces rather than updates the resource |
The rule this produces is short enough to put on a wall: any secret whose loss logs users out must be born outside the process that consumes it, and must outlive every deployment artifact that references it. If you cannot name the system of record for your session-signing key, you do not have one.
There is a nasty sub-case worth calling out because it disguises itself. When each replica generates its own secret, you don't get an outage — you get a load-balancer lottery. With twelve pods, a user's cookie validates on one in twelve requests. Symptoms: intermittent logouts, unreproducible, "works for me," blamed on the client. It stays that way indefinitely if you have sticky sessions masking it, and it converts into a total outage the day you remove session affinity or scale the fleet. The first four minutes of the story at the top of this article looked exactly like this, which is why the initial hypothesis was "something flaky."
2. A serialization change the new binary can't read — or reads wrongly. The failure-to-deserialize version is well understood and covered by the two-phase rule in the blue-green article. The version worth your attention is the one that succeeds and is wrong:
- Enum ordinals. A codec that stores enums by position, and someone inserts a value in the middle of the declaration.
AuthMethod.PASSWORDdeserializes asAuthMethod.WEBAUTHN. Nothing throws. Sessions keep working, and step-up decisions are now made against a factor the user never presented — a security bug that presents as no bug at all. - A class rename or package move. The stored blob names a type that no longer resolves. This one at least fails loudly.
- Field reordering in a positional/tuple codec, where
created_atandlast_seen_atswap and your idle timeout starts measuring the wrong interval. - A unit change.
auth_timewidened from seconds to milliseconds. Old sessions are now interpreted as having authenticated in 1970; new sessions read by old code claim a timestamp in the year 56,000. Which direction breaks depends on which side of the comparison you're on, and one of the two directions silently skips a required re-authentication rather than forcing one.
The distinction that matters operationally: a decode failure logs people out, which is loud and self-limiting. A decode mis-read keeps people logged in with corrupted authentication state, which is quiet and unbounded. The authentication state vs. session state split tells you which fields are which — anything feeding an authorization or step-up decision belongs in the second category, and belongs under a stricter compatibility contract than the rest of the envelope.
3. Cookie attribute changes, including the ones that are improvements. Changing the cookie's Name, Path, or Domain is a rename, and a rename is a mass logout. The uncomfortable case is that adopting __Host- — which is the correct thing to do, and which Sessions, Tokens, and Cookies argues for on the grounds that it structurally eliminates cookie tossing — is also a rename. The security hardening and the outage are the same change. Ship it as a dual-read: accept both session and __Host-session, issue only the new one, and delete the old reader after a full session lifetime has elapsed.
Two subtler ones:
SameSite=Lax→Strictdoesn't invalidate anything. It makes the cookie conditionally invisible, on cross-site top-level navigations — arriving from an email link, from a partner portal, from a SAMLHTTP-POSTbinding. The result is a partial, traffic-shaped logout that correlates with referrer rather than with time, which makes it nearly impossible to attribute to the deploy.- Adding
Secureor narrowingPathwhile any client is still on the old scheme or path. The browser keeps the old cookie and accepts the new one, both named the same; RFC 6265 doesn't define the ordering, most servers take the first match, and you get a coin flip per request.
4. Signing key continuity. A kid that disappears from JWKS before the tokens signed with it expire. The mechanics of overlapping validity are in Designing for Key Rotation Before You Need It, so the deploy-specific point is narrower: most teams put a kid in their JWTs and no key identifier in their session cookies. The JWT path has a key ring, an overlap window, and a rotation runbook. The cookie path has one secret, in one environment variable, with no version marker and no way to accept the previous one. That asymmetry is why the cookie is what breaks.
5. Session-store changes that are a flush wearing a costume. Changing the Redis key prefix from sess: to session:v2:. Changing the store's serializer. Resharding a cluster without migrating slots. Switching from one library's key layout to another's. In each case the old data is intact, addressable, and unreachable — which is operationally identical to having deleted it, but reassuring on the dashboard because memory usage barely moves.
6. TTL and clock semantics. The delayed-detonation category. A refactor that stops refreshing the store TTL on read converts a sliding eight-hour session into an absolute one. Nothing happens at deploy time. Eight hours later, everyone whose session was established before the deploy expires together — and by then the deploy is closed, the change is off the top of the timeline, and nobody connects the two. Same shape when a session-lifetime policy reduction (24h → 8h, usually for a compliance answer) is evaluated against existing sessions rather than applied to new ones. That one detonates instantly, and it is a policy change, so it doesn't go through deploy review at all. Session Timeout Policy Across a Product Suite covers why these numbers get chosen badly; the operational point here is that a retroactive tightening is a mass invalidation and must be scheduled like one.
7. Sticky-session loss on rolling restart. Not really a cause so much as a revealer. If anything load-bearing lives in process-local memory — a decrypted session cache, a PKCE verifier, a partially-completed MFA challenge — affinity has been hiding it, and a rolling restart removes affinity for everything simultaneously.
8. Non-human clients, which behave differently. If the invalidation touched signing keys or refresh tokens rather than browser cookies, your herd includes machine identities — and in most estates those outnumber the humans. They have a different shape: no cognitive pacing, no coffee break, no giving up. A human who fails to log in twice will wander off for five minutes, which is unintentional but real load shedding. A service with a fixed 1-second retry and no jitter will hold its share of the load flat until you fix the problem or it exhausts its file descriptors. Human herds decay; machine herds don't.
Making sessions survive a deploy on purpose
The compatibility rules are in the blue-green article. What follows is the mechanism I'd actually build, which is one idea applied consistently.
Give the session envelope a key ring and a kid, exactly like your JWTs. Outside the ciphertext, in cleartext: format version and key identifier. Verification accepts any key in the ring; issuance uses the current one. This costs a handful of bytes and eliminates two of the eight causes above outright.
Re-wrap on read. This is the part worth building deliberately, because it makes both key rotation and format migration free and, more importantly, observable. When a request arrives carrying a session sealed with key n-1 or encoded at format v1, you decrypt it with the old key, re-encrypt with the current one, and set the cookie on the way out. Nobody is logged out. The population migrates at the rate of real traffic. And because you now have a counter of how many sessions still carry the old marker, you have an actual gate for the contract phase:
Don't delete the old reader on a schedule. Delete it when the gauge of sessions still using it has been zero for one full maximum session lifetime.
This is the difference between "we think everyone's migrated by now" and knowing. It also converts the blue-green article's deploy readers before writers rule from a discipline you have to remember into a state machine you can watch.
Separate the invalidation authority from the crypto. You want a way to say "these sessions are no longer valid" that has nothing to do with which key sealed them, because otherwise the only tool you have for a deliberate mass logout is destroying key material — which is precisely the accident this article is about, performed on purpose. The standard shape is a monotonic epoch: a counter at global, tenant, and user scope, carried in the session envelope and compared on read. Bumping a user's epoch is "log this user out everywhere." Bumping the tenant's is "log out one customer." Bumping the global one is a decision with a name on it.
Crucially, the epoch check should support a deadline rather than an instant. Instead of "every session below epoch 7 is invalid now," express it as "every session below epoch 7 must re-authenticate before T," and assign each session an individual deadline drawn uniformly from the window. Four hundred thousand sessions spread over ninety minutes is 74 re-authentications per second — a bad hour, not an outage. Same policy outcome, different physics, and the only code required is a random offset computed at check time from a hash of the session ID so it's stable across requests.
Decouple secret lifecycle from process lifecycle. From the table above: the key comes from a secret manager, it has a version history, rotation is an explicit operation with an audit record, and no code path anywhere generates one as a fallback. The single highest-value change most teams can make here is deleting the fallback — make an absent session key a hard startup failure. A process that refuses to boot is an incident lasting four minutes. A process that boots with a fresh random key is an incident lasting two hours, and it looks healthy the whole time.
Testing for it before you ship
The question you want a machine to answer, on every build, without anyone remembering to ask it:
Will yesterday's session validate against tomorrow's binary?
Three layers, cheapest first.
A golden artifact corpus. On every release, CI emits a fixture set: a sealed session cookie, a serialized store blob, a JWT signed by each key in the current ring, a refresh-token record, a remembered-device record. Commit them, tagged by release. Then a test in every subsequent build loads the fixtures from the last N releases — where N covers your longest artifact lifetime, not your release cadence — and asserts each one decodes to the expected values, not merely that it decodes without throwing. The value assertion is what catches enum-ordinal drift and unit changes; a decode-success assertion catches only the loud half.
Run the same test in reverse for rollback: can release N−1's binary read fixtures produced by release N? That's the mechanical version of the blue-green article's question nine, and it's the one people skip because it requires keeping the old binary around. Keep the old binary around.
Shadow decode in production. Deploy the new codec in a mode where it attempts to decode live sessions, reports the outcome, and acts on nothing. Emit session_decode_shadow_failure_ratio and session_decode_shadow_mismatch_ratio — failure and disagreement with the current codec are different signals, and the second is the dangerous one. Cut over when both have been zero across a full session lifetime, which by definition means every live session shape has been exercised.
A deploy gate with a real session. A synthetic user establishes a session against the current version, and the rollout will not proceed past the first canary pod unless that pre-existing session still authenticates against the new one. It is about forty lines of code and it would have caught every single cause in the list above except the delayed-TTL one. Make it a gate, not an alert — an alert during a rollout is a notification that you are already halfway through the incident.
For the TTL and clock cases, which no static check catches, the answer is a time-accelerated environment: a staging system with a realistic session population and a clock you can advance. That's a chaos experiment rather than a test, and it belongs in the catalogue alongside the others in Chaos Engineering for Identity Systems. Rotate the cookie key in staging against 100,000 live sessions and observe whether anything at all happens. If the answer is "nothing," you have a key ring. If the answer is "everything," you have found your next quarter's work in an environment where it doesn't matter.
The first sixty seconds
Suppose it happens anyway. The signature is distinctive and the triage is one question.
What you see: 401s up sharply, 5xx flat, load-balancer error rate flat because 401 is a successful HTTP transaction. A step change in session_decode_failure or invalid_session_key. New-session creation rate climbing to a multiple of baseline. Every health check green, every dependency green, and the login endpoint's p99 climbing past the client timeout.
The question that determines what you do next:
flowchart TD
A["Sessions failing<br/>+ login rate spiking"] --> B{"Do sessions created<br/>AFTER the deploy<br/>also fail?"}
B -->|Yes| C["Live bug — the platform is<br/>still destroying sessions.<br/>Stop the bleeding first:<br/>halt rollout, restore key/codec"]
B -->|No| D["One-shot invalidation.<br/>Population is already lost.<br/>Rollback restores the cause,<br/>not the state."]
D --> E["Survive the herd:<br/>bulkhead, admit below capacity,<br/>jitter the clients"]
C --> E
The right branch is the one people get wrong under pressure, and it's worth saying plainly: when the invalidation is a one-shot event, rolling back does not undo it. The sessions are already dead in four hundred thousand browsers. Rollback restores the code that stopped destroying sessions, which is necessary, and restores exactly zero sessions, which is the part nobody expects. In the opening story the rollback actively made things worse, because re-rendering the chart minted a third key and invalidated the sessions people had just finished re-establishing. Before you roll back a session-invalidation incident, ask what the rollback does to key material and stored format. If the answer is "changes it again," you are about to fire the second barrel.
Then survive the herd. The general playbook is in the metastable-failure article; three moves are specific to this situation:
Bulkhead by cost class. Credential verification gets its own bounded pool with its own admission control, separate from token validation, introspection, and session reads. Without this, a herd of password logins starves the read path — and the read path is 90%+ of your traffic and belongs to the users who are still logged in and currently fine (Identity Is Mostly Read Traffic). One bulkhead is the difference between "logged-out users are queuing" and "the platform is down."
Admit below capacity and reject the rest fast. Twenty-four verifications a second completing inside the client timeout is worth infinitely more than 267 a second completing after everyone has left. Denominate the limit in CPU cost rather than requests, for the reasons in Rate Limiting an Identity Platform Is a Capacity Problem, and enforce it per tenant so one customer's population can't consume the whole pool.
Jitter the clients, if you own them. This is the highest-leverage thing available and almost nobody has built it. The default SPA 401 handler redirects to the authorization endpoint immediately — which means your own frontend is the transport for the retry storm. A handler that waits a random 2–20 seconds, shows "reconnecting," and then redirects converts a 267/s wall into a 60/s ramp. Users experience a few seconds of delay instead of a login page that doesn't work. Build it before you need it; you cannot ship a frontend change to logged-out users during the incident.
Do not, incidentally, expect scaling out to rescue you. New instances arrive with cold caches and immediately take a full share of traffic, and the pool you need to grow is CPU-bound on a cost you cannot reduce. Capacity is how you avoid this. It is not how you climb out of it.
When logging everyone out is the right answer
All of the above is about mass invalidation as an accident. Sometimes it's the correct action, and the design should make it easy rather than making it impossible.
The cases are narrow and they are all compromise cases: signing key or session-key exposure, a dump of the session store, a vulnerability that allowed sessions to be minted or escalated, an authentication bypass of any kind. In each, the argument is the same — you cannot enumerate which sessions are attacker-controlled, and a partial invalidation based on a guess leaves the attacker in while logging out everyone else. When you can't establish scope, the scope is everything.
Note what this does to the staggering advice. A ninety-minute jittered deadline is exactly right for a routine policy change and exactly wrong for a live intrusion, because for ninety minutes the attacker still has a session. Emergency invalidation is instant, unstaggered, and you accept the herd. That's not a failure of the design; it's the design working — you built the graceful path for the ordinary case and you keep the sharp one for the case that needs it. The emergency-rotation discussion makes the same trade in the key domain.
Which leaves the question the title is really asking. Operationally, what separates a deliberate mass logout from an outage? Not the mechanism — the physics are identical, and your users cannot tell them apart. The difference is entirely preparation:
| Deliberate | Accidental | |
|---|---|---|
| Trigger | A named epoch bump with an operator and an audit record | A template, a refactor, a chart |
| Timing | Chosen — off-peak, not Friday, not during a customer's month-end | 14:05 on a Thursday |
| Shape | Staggered over a chosen window, or instant by explicit decision | Instant, by accident |
| Capacity | Hashing pool and MFA delivery pre-scaled; admission control armed before the bump | Discovered mid-incident |
| Clients | In-app notice ahead of time; jittered re-auth path deployed | Every SPA redirecting at once |
| Comms | Sent before, so support isn't triaging "is the site down" | Written at minute forty |
| Verification | You can prove every pre-T session is gone | You're not sure which ones died |
Everything in the left column is a capability, not a decision. Build the deliberate path — the epoch, the stagger, the pre-scale runbook, the jittered client — and you have simultaneously built your recovery tooling for the accidental one. That's the argument for doing it before you have a reason to: the mass-logout button and the mass-logout incident response are the same machinery, and you will use one of them eventually whether or not you chose to.
Where this lands
The reason this failure keeps happening to competent teams is that it lives in the gap between two things that are each individually well managed. The compatibility work — versioned envelopes, key rings, readers before writers — is understood and mostly done. The capacity work — hashing pools, admission control, retry budgets — is understood and mostly done. What's missing is the arrow between them: the recognition that a session-format detail and a CPU capacity plan are the same subject, connected by Little's Law and a factor of twenty-seven.
Session lifetime is the coupling constant. Longer sessions make steady-state login capacity cheaper and make the blast radius of losing them larger, in exact proportion. That's not an argument for shorter sessions — it's an argument for knowing which side of the trade you're on, and for measuring your hashing pool against your session population rather than against your login rate.
The one-line test, before any deploy that touches cookies, keys, session storage, or timeouts: if every live session died the moment this ships, how long would it take to let everyone back in? If you can't answer in a number, the answer is longer than the incident review will find acceptable.