Blue-Green Deploying an Identity Platform

The cutover took four seconds. The load balancer moved 100% of traffic from the blue fleet to the green fleet, connection draining finished cleanly, and the deploy dashboard turned a satisfying shade of green. Error rate: flat. Latency: flat. Somebody typed "clean cutover" in the channel and started closing tabs.

Eleven minutes later the first ticket arrived: "I got signed out and now it won't let me back in." Then forty more. Then a customer's API integration started failing with invalid_token on tokens that had another six hours of validity left. Nobody had deployed anything since the cutover. Nothing was down. Every health check was passing, and the system was rejecting credentials it had issued itself.

The rollback made it worse.

This is the failure mode that makes blue-green deployment strange for identity platforms specifically, and it comes from one property that stateless services don't have. When you drain a web server, you're waiting for in-flight requests. When you drain an identity platform, you're waiting for in-flight artifacts — and their lifetimes are measured in hours to months, not milliseconds.

That sentence is most of the article. The rest is what follows from it.

The asymmetry: what actually needs to drain

A stateless HTTP service drains in seconds because all of its state lives inside the request. When the last response is written, the old version holds nothing anyone will ever need again. You can delete it. That's what makes blue-green feel safe: the old fleet becomes irrelevant the moment its connections close.

An identity platform is a factory for durable artifacts. Every one of them is a promise that some future version of your system will honor something this version produced:

Artifact Typical lifetime Who must still understand it
Authorization code 30–60 seconds Whichever fleet handles the token request
Access token (JWT) 5–60 minutes Every resource server, plus introspection
ID token 5–60 minutes Every RP, plus your own nonce/replay state
Session cookie 8–24 hours (idle-extended) Every fleet that reads a cookie
Refresh token 14–90 days The token endpoint, forever, plus reuse detection
Device-flow / CIBA pending request 5–15 minutes Whichever fleet the poll lands on
Remembered-device / "trust this browser" record 30–180 days The risk engine at every login
Enrollment or recovery link minutes to days The redemption endpoint

The green fleet must honor every artifact blue issued. That's the obvious half, and most teams get it right because it's the direction the deploy actually goes.

The half that causes incidents is the reverse: after a rollback, blue must honor everything green issued. Green was live for eleven minutes. In those eleven minutes it minted sessions, refresh tokens, and signed tokens using whatever the new version does. Roll back, and those artifacts are now being presented to a fleet that predates them.

So the real drain window for an identity deploy is not your connection drain setting. It is:

The longest lifetime of any artifact either version issues.

For a platform with 90-day refresh tokens, "blue and green must be mutually compatible" is a statement about ninety days, not about the ninety seconds your deploy tool waits. You don't have to keep both fleets running for ninety days — you have to keep both versions' artifacts interpretable for that long, which is a versioning discipline, not an infrastructure one.

Most of what follows is about making that true on purpose instead of discovering which parts of it were accidentally true.

Signing keys: the cutover is not the event that matters

The first thing people get wrong is treating the signing key as part of the deploy. It isn't. It's a separately-clocked process that must straddle the deploy, and if you tie it to the cutover you have built an outage with a countdown timer.

The mechanics of phased rotation are covered in Designing for Key Rotation Before You Need It, and I won't rerun that argument. What matters here is how rotation interacts with a fleet cutover, which is where three specific traps live.

Trap one: green signs with a key blue can't validate. If green ships new key material and starts using it immediately, every token it issues is unverifiable by blue. Roll back and every green-issued token becomes garbage — instantly, for its whole remaining lifetime, across every resource server that fetched JWKS during the green window. This is the one-way door in its purest form: the rollback is technically available and functionally catastrophic.

Trap two: RP-side JWKS caching means "published" and "fetchable" are different states. You control when a key appears in your JWKS document. You do not control when a relying party's cache picks it up — some libraries refresh hourly, some daily, some once at process start and never again. Publishing the green key at cutover means a meaningful fraction of validators will see a kid they've never heard of on the very first token green issues. The good libraries refetch and recover. The rest return invalid_signature until someone restarts them, which is a customer-visible incident you caused with a deploy that never failed a health check.

Trap three: no kid. Every token needs a key ID in its header, or validators holding two keys during an overlap are reduced to trial verification — slower, and genuinely ambiguous once one of those keys is revoked. This is bug five in Seven Identity Bugs Every Team Eventually Ships for a reason: it's invisible until the first rotation and impossible to retrofit onto tokens already in circulation.

The rule that resolves all three: publish keys N cache-TTLs before you sign with them, and retire them N token-lifetimes after you stop. Both waits are decoupled from the deploy.

flowchart LR
    A["T-24h<br/>Publish green key<br/>in JWKS<br/>(blue still signing)"] --> B["T-0<br/>Cutover<br/>blue → green<br/>both keys published"]
    B --> C["T+24h<br/>Green begins<br/>signing with<br/>green key"]
    C --> D["T+24h..+N<br/>Both keys valid<br/>rollback still safe"]
    D --> E["T + max token<br/>lifetime<br/>Retire blue key"]

Read the diagram for what it doesn't show: there is no moment where a fleet cutover and a signing-key change happen together. Rotation and deployment are two independent slow processes that overlap. Serialize them and each one stays individually reversible.

There's a per-tenant wrinkle worth naming, because it changes the arithmetic. If your platform issues per-tenant signing keys under per-tenant issuers — which ClavionX does, and I'll disclose that as the platform I work on — a "rotation" isn't one event, it's N events with N independent RP cache populations. You cannot reason about "the JWKS cache TTL" as a single number, and a cutover that assumes a uniform propagation window is assuming something false for your slowest tenant. The same shape appears in the multi-region case, where the propagation boundary is geographic instead of tenant-scoped (Multi-Region Identity).

SAML federation is worse still, because there's no discovery mechanism at all and the propagation step involves a human at another company. If your deploy touches SAML signing material, the coordination window is measured in weeks, not TTLs — see SAML Certificate Rotation Without an Outage.

Session format: the two-phase rule

Sessions are where blue-green quietly turns into a data migration problem.

Whatever your session representation — an encrypted cookie envelope, a serialized blob in Redis, a row with a JSON column — a deploy that changes it creates a window where two versions of your code are reading each other's writes. During the cutover that window is brief. After a rollback it is the entire remaining lifetime of every session green touched.

The rule is simple to state and constantly violated:

Deploy the version that can read both formats before you deploy the version that writes the new one.

Readers before writers. Always. Which means a session format change is a minimum of two deploys, and often three:

  1. Deploy N: read both, write old. Green understands v1 and v2. It still emits v1. Fully reversible — blue never sees anything unfamiliar, because nothing new was written.
  2. Deploy N+1: read both, write new. Green now emits v2. Reversible only back to N, not to N−1. This is the deploy where the one-way door closes behind you.
  3. Deploy N+2: read new, write new. The v1 reader is deleted, once no v1 session can still exist — that's max session lifetime after step 2, not "a few days after."

Compress those into a single deploy and you get the failure at the top of this article. Green writes v2 cookies. Eleven minutes of users get them. You roll back. Blue receives a cookie with a version byte it doesn't recognize — or worse, a field layout it does recognize and misparses — and the honest implementations reject it as tampered and clear it. Every user green touched is logged out mid-workflow, and the "clean rollback" is the thing that logged them out.

Three implementation details make the difference between a rollback and an incident:

Put a version byte in the envelope, before the ciphertext. Not inside the encrypted payload — you have to know how to parse before you can decrypt safely. An unrecognized version must be a clean, logged, single-line "unknown session format version 3, expected ≤2" rejection, not a decryption exception that reads identically to an attack.

Additive-only within a version. New fields must be optional with defaults; removed fields must be tolerated as absent. Old code must ignore unknown fields rather than throwing — the same forward-compatibility contract you'd apply to a wire protocol, applied to your own session. If your serializer's default behavior is "unknown field is an error," you have chosen fail-closed on rollback without knowing it.

Never change the meaning of an existing field. Widening auth_time from seconds to milliseconds is a format change wearing a disguise, and the old version will happily interpret the new value as a timestamp thirty thousand years in the future — or in the past, depending on which direction the comparison runs. That's how a session that should have required re-authentication silently doesn't. The authentication state vs. session state distinction matters here: fields that feed authorization decisions (AMR, ACR, auth_time, step-up markers) are the ones where a misparse becomes a security bug rather than a logout.

In-flight protocol state: the 30-second artifacts

Authorization codes have a lifetime of under a minute, which makes them feel exempt. They aren't — they're the artifacts most likely to be mid-flight at the exact instant of cutover, which is a different risk than long lifetime and needs a different answer.

The scenario: blue issues an authorization code at T−5s. The user's browser follows the redirect. At T+0 the cutover happens. At T+25s the client POSTs to the token endpoint, and green has to redeem a code it never issued.

That's fine if — and only if — every piece of state that redemption depends on lives in shared storage both fleets read, with a representation both understand:

  • The code record itself, with its client binding, redirect URI, scopes, and single-use marker. Consumption must be atomic across fleets; if green and a straggling blue instance can both redeem the same code, you've built a replay window into your deploy process.
  • The PKCE code_challenge and method. If this rides along in the code record, fine. If a previous engineer optimized it into instance-local memory or a sticky-session cache, the redemption on green fails invalid_grant and the user sees a login loop with no error anywhere that names the deploy.
  • Nonce and replay state, which the same argument covers.
  • Device-flow and CIBA pending requests. These are polled repeatedly over minutes by a client that has no idea a cutover happened. Each poll may land on a different fleet. The pending-request record, its approval state, and its polling interval must be shared and version-compatible — and a poll that arrives at a fleet which can't read the record must return authorization_pending, not an error that makes the client give up.

The general test is worth writing on the wall: any protocol state that must survive a cutover cannot live in process memory, and cannot be sticky-routed as a substitute for being shared. Sticky routing hides this class of bug beautifully in staging and then fails during the one deploy where the sticky mapping is exactly what you're changing.

Projected config and version skew

If you've split runtime from control plane — and most identity platforms eventually do, whether they meant to or not (Why Identity Platforms Become Distributed Systems) — blue-green acquires a second skew axis. You're no longer just running two versions of the runtime. You're running two versions of the runtime against config projections that may have been written by a control plane of a third version.

The failure is specific and nasty. The control plane deploys a projection schema with a new field. The runtime that consumes it is older, its deserializer is strict, and an unknown field is an error. The runtime doesn't crash — it fails to load the tenant's projection. And if the runtime is correctly designed to fail secure on missing projected state (What Should an Identity Platform Do When Its Database Is Down?), the result is a runtime that is healthy, serving traffic, passing every liveness probe, and refusing to authenticate anyone for that tenant. A strict parser turned a config change into an outage with a perfect health dashboard.

This is where the reader/writer ordering rule generalizes past sessions into an invariant for the whole platform:

Deploy readers before writers. The consumer of a format must understand version N+1 before any producer emits it.

For a runtime/control-plane split that means: runtime first, control plane second. Always, in that order, regardless of which side the feature "belongs" to.

Three properties make projection skew survivable:

Ignore unknown fields, unconditionally. Forward compatibility is not optional for a projection consumed by a fleet you don't deploy atomically with.

Version-stamp the projection and let the runtime decide. A runtime that reads schema_version: 4 when it understands up to 3 should serve the fields it recognizes and emit a loud metric — not fail, and not guess.

Never let a runtime resolve skew by calling the control plane. The instinct when a projection looks wrong is to fetch the authoritative version. That edge is exactly the one that must not exist: it makes the read path inherit the write path's availability, and it will be exercised for the first time during the incident where hammering your administrative database is the worst available move. ClavionX makes this an absolute rule — the runtime never synchronously calls the control plane during request execution, under any condition including failure; state arrives only as events and projections, and a missing projection fails secure rather than reaching back for a fresh read. I cite it as the system I work on and as one concrete instance of a general constraint; the same rule shows up wherever people build identity like Kubernetes.

Database migrations: expand and contract, never rename

Blue-green means both versions run against one schema simultaneously. Not two schemas. One. Every migration therefore has to be compatible with the code on both sides of the cutover, in both directions.

That rules out the entire category of in-place changes:

Don't Do
RENAME COLUMN a TO b Add b, dual-write both, backfill, later drop a
Change a column's type in place Add a new column of the new type, migrate, drop
Add NOT NULL without a default Add nullable, backfill, then constrain in a later deploy
Drop a column the old code reads Stop reading it, deploy, then drop
Add a unique constraint Verify uniqueness first, add the index concurrently, then constrain

Expand/contract is standard practice and I'm not claiming otherwise. What's specific to identity is the contract phase's timing. Ordinarily you can drop the old column once the old code is gone. Here you can't, because "the old code is gone" isn't the binding condition — the binding condition is that no artifact issued by the old code is still valid. Drop a session table column while 20-hour sessions issued by blue are still in circulation and you've broken sessions on a fleet that isn't even running any more.

The other identity-specific hazard: do not run a migration that rewrites credential or key material during a blue-green window. Password hash upgrades, credential re-encryption under a new KEK, token record re-encoding — these are one-way transforms of the very rows a rollback would need in their original form. Run them as their own deploys with their own windows, on the rehash-on-verify pattern where possible, and never inside the transaction that also flips fleets.

Rollback is the hard part, and often you don't have one

Here's the reframing I'd most like a platform team to take away.

Blue-green's promise is a fast, safe rollback. For identity workloads that promise is conditional, and the condition is rarely checked. Every artifact green creates that blue cannot interpret is a brick in a wall behind you. After enough minutes of green traffic, "roll back" stops meaning "return to the previous state" and starts meaning "invalidate everything the new version produced."

So run the test before you deploy, not during the incident:

Can blue serve every artifact and every row green will create?

Go through it concretely. Tokens signed with a key blue holds? Sessions in a format blue parses? Refresh tokens blue's reuse-detection can chain? Rows blue's ORM can deserialize? Projections blue's schema version tolerates?

If any answer is no, you do not have a blue-green deploy. You have a roll-forward-only deploy wearing blue-green's infrastructure, and the correct response is not to cancel it — it's to know that before you start, so that:

  • the runbook says "fix forward" instead of listing a rollback step that will make things worse,
  • the on-call engineer isn't discovering this at 02:00 with a dashboard full of red,
  • and the change is scheduled when a fix-forward is survivable.

A named one-way door with a prepared fix-forward path is a fine deploy. An unnamed one is how a ten-minute blip becomes a two-hour outage — because the first instinct under pressure is to roll back, and rolling back is the action that logs everyone out.

What "the deploy that logs everyone out" looks like

Briefly, because this deserves its own article and will get one. The signature is distinctive enough to recognize in thirty seconds:

  • Error rate at the load balancer: flat. The requests are succeeding. They're succeeding at returning 401.
  • A step change in session_decrypt_failed or unknown_token_format, starting at the rollback, not at the cutover.
  • Login rate spiking to several times baseline — that's not attack traffic, it's every logged-out user re-authenticating at once, and it will hurt, because everything that made your read path fast is now cold (Identity Is Mostly Read Traffic).
  • Token endpoint 400s with invalid_grant, concentrated in refresh rather than authorization-code grants.
  • Every health check green throughout.

The tell is the timing: the damage begins at the rollback, which is why the first hypothesis in the channel is always wrong.

Blue-green vs. canary for this workload

Blue-green is not obviously the right strategy here, and I'd argue that for most identity deploys it isn't.

Blue-green Canary / rolling
Exposure before you notice 100% of traffic instantly 1–5%, growing
Artifacts created by the bad version All of them, immediately A bounded, small fraction
Version skew during deploy Two versions, briefly Two versions, for a long window
Compatibility work required Two-way (rollback needs it) Two-way (both run concurrently)
Good for Big schema or crypto changes needing one clean switch Almost everything else
Bad for Anything that creates artifacts you can't un-create Changes that can't tolerate mixed versions at all

The decisive column is the second one. The core risk of an identity deploy is artifact contamination — the new version minting things you'll later be unable to honor. Canary bounds that contamination to the canary's traffic share. Blue-green does not bound it at all; it maximizes it, instantly, which is precisely the wrong shape for this workload.

Note the row that surprises people: canary requires the same two-way compatibility work as blue-green, because both versions serve live traffic side by side for far longer. The compatibility discipline in this article isn't a blue-green tax. It's the price of deploying an identity platform without a maintenance window at all, and blue-green just makes skipping it fail more spectacularly.

Blue-green still wins in two cases. First, when the change genuinely cannot tolerate mixed versions — a crypto envelope change, a storage engine swap — and you'd rather have one sharp, short skew window than a long one. Second, when you need a full-fleet warm state before traffic arrives: caches primed, projections loaded, JIT warm. Cold-start latency on a freshly-provisioned identity fleet is real, and blue-green lets you pay it before the cutover instead of during it.

The pre-deploy checklist

Ten questions. If you can't answer all ten, the deploy isn't ready — and answering them takes about twenty minutes.

  1. What is the longest-lived artifact either version issues? That number, not the drain timeout, is your compatibility horizon.
  2. Does every token carry a kid, and is every key either fleet might use published in JWKS now?
  3. Was the new key published at least one max-JWKS-cache-TTL ago? If not, delay signing, not the deploy.
  4. Does the session format change? If yes: is this the read-both-write-old deploy, or the write-new deploy? Say which out loud.
  5. Can the old version parse everything the new version will write — sessions, tokens, DB rows, projections?
  6. Is all in-flight protocol state shared? Codes, PKCE verifiers, nonces, device-flow and CIBA pending requests. Nothing instance-local, nothing depending on sticky routing.
  7. Is the migration expand-only? No renames, no in-place type changes, no drops of anything the old code reads.
  8. Does the projection schema version change, and are readers deployed before writers?
  9. Given all of the above — is rollback actually available? Write the answer down as yes or no.
  10. If no: is the fix-forward path written, and is now a good time to walk through a one-way door?

Question nine is the one that matters. Everything else is a way of arriving at an honest answer to it.

Where this lands

Blue-green deployment assumes the old version becomes irrelevant when its connections close. Identity platforms violate that assumption by design: they exist to issue durable artifacts, and every artifact is a promise that some future version will still understand what this one produced.

That doesn't make blue-green wrong. It makes the drain window the wrong unit of measurement. Once you measure in artifact lifetimes instead of connection lifetimes, the rest of the discipline falls out on its own — publish keys before you sign with them, deploy readers before writers, expand before you contract, and know before you start whether the door behind you is still open.

The deploys that go badly are almost never the ones where somebody got the compatibility work wrong. They're the ones where nobody asked whether it was needed.