Migrating Off an Identity Provider Without Downtime

Nobody plans this project. It arrives — a provider gets acquired and the roadmap changes, pricing moves somewhere the finance team won't follow, a compliance requirement lands that the current platform can't satisfy, or the architecture stopped fitting three years ago and everyone finally admits it.

Nobody plans this project. It arrives — a provider gets acquired and the roadmap changes, pricing moves somewhere the finance team won't follow, a compliance requirement lands that the current platform can't satisfy, or the architecture stopped fitting three years ago and everyone finally admits it.

Then someone asks how long it'll take, and the honest answer is a quarter or more, which nobody believes until they're in it.

The technical problems here have known solutions. The reason migrations take quarters is that the hardest part isn't technical at all.

The password hash problem

Start with the constraint that shapes everything else: you cannot re-hash a password you don't have.

Hashes are one-way. Your old provider stores bcrypt(alice's password). You want argon2id(alice's password). There is no transformation between them, because recovering the plaintext is precisely what the hash is designed to prevent.

Three strategies, in ascending order of how often they actually work.

Export the hashes. If the old provider lets you export them, and the new one supports that exact algorithm and encoding, you bulk-import and you're done. Both conditions matter — bcrypt with a $2a$ prefix and bcrypt with $2y$ are the same algorithm with different encodings, and a provider that supports one may reject the other.

Worth stating plainly: many providers do not export password hashes at all. That is a lock-in mechanism, whether or not it's described as one, and it's the single most useful thing to check before you sign a contract rather than after.

Rehash on login. This is the technique that makes migration tractable and it's oddly under-documented.

Import whatever you can get — old hashes in their original format — and store them alongside a marker for which algorithm they use. When a user logs in for the first time on the new system, you have something you'll never have again: the plaintext, in memory, for a few milliseconds. Verify it against the old hash. If it matches, immediately compute a new hash with your target algorithm, store that, drop the old one.

if user.hash_type == "legacy":
    if legacy_verify(password, user.legacy_hash):
        user.hash = argon2id(password)      # only moment plaintext exists
        user.hash_type = "argon2id"
        user.legacy_hash = None
        return SUCCESS

Users notice nothing. Every successful login converts one account. After a few weeks, most of your active population has migrated itself.

The same technique upgrades work factors within an algorithm — it's worth building even if you never migrate providers, because it's how you raise a bcrypt cost factor without resetting anyone's password.

Wrap the hashes. Store argon2id(bcrypt(password)) — apply the new algorithm on top of the old hash. You get an immediate bulk import with the new work factor applied, and you keep it forever: every verification runs both algorithms, and you can never remove the old implementation from your codebase. It's a real option for regulated environments that can't tolerate legacy hashes sitting in the database during a migration window, and it's a permanent tax otherwise.

The dormant tail

Rehash-on-login has an obvious limit: it only converts accounts that log in.

Three months after cutover, some fraction of accounts hasn't returned. In a consumer product that can be 40% or more. Those accounts still hold legacy hashes, which means you're still running the old verification path, which means you haven't finished.

Decide this policy before the migration, not after:

  • Keep the legacy path indefinitely. Simplest, and you never finish. The old algorithm stays in your code and in your threat model forever.
  • Force a reset after a deadline. Clean, and it converts a silent migration into a visible event for a population that will contact support confused about why their password stopped working.
  • Expire the accounts. Only defensible if they were genuinely abandoned, and "genuinely abandoned" is a product decision, not an engineering one.

There's no correct answer. There's only making the choice deliberately, with the support team in the room, rather than discovering in month six that nobody decided.

What else has to move

Passwords get the attention. These are the ones that surprise people.

MFA enrollments. TOTP secrets are just shared seeds and can often be exported and imported. SMS and email factors are contact details — usually straightforward.

Passkeys generally cannot move at all. This one is worth understanding precisely, because it's a hard constraint rather than a difficulty. A WebAuthn credential is bound to a Relying Party ID — effectively your domain. The private key lives in the user's authenticator and cannot be exported by anyone, including you. If your RP ID stays identical and the new provider can import credential IDs and public keys, migration is possible. If the RP ID changes, every passkey is dead and every user must re-enroll.

Plan the re-enrollment campaign as part of the project, and check the RP ID question early, because it may constrain your domain strategy.

Federation connections. Every enterprise customer's SAML or OIDC configuration must be rebuilt: new metadata, new certificates, new ACS URLs. Each one requires their identity team to make a change on their side. More on this below, because it's the whole schedule.

API clients and secrets. Every machine integration needs new credentials. Client secrets can't be migrated in any meaningful sense — you're issuing new ones and coordinating the swap with whoever owns each integration, including customers you may not have direct contact with.

Groups, roles, and consent records. Structural data that usually transfers, but rarely with identical semantics. Group nesting, role inheritance, and scope definitions tend to differ enough that a mapping exercise is required.

Identifier stability

Here's the failure that turns a migration into a data project.

Your application stores records keyed on the identity provider's subject identifier. The new provider issues its own sub values, which are different. Every foreign key pointing at the old ones is now dangling.

Two ways out:

Map old to new. Build a translation table during migration, keyed on something stable — usually email — and rewrite your references. Workable, and it's an unbounded amount of work proportional to how many places sub leaked into.

Key on your own identifier. If your application has its own user ID and treats the provider's sub as one attribute of that user rather than the primary key, migration is a matter of updating one column. This is the version you want, and it's only available if you did it before you needed it.

flowchart LR
    subgraph Fragile["Keyed on provider sub"]
        R1["orders.user_sub"] --> P1["old IdP sub"]
        R2["prefs.user_sub"] --> P1
        R3["audit.actor_sub"] --> P1
    end
    subgraph Portable["Keyed on your own ID"]
        U["users.id (yours)"] -.->|"one column"| P2["idp_subject"]
        R4["orders.user_id"] --> U
        R5["prefs.user_id"] --> U
    end

The dual-run pattern

The way to do this without downtime is to run both providers simultaneously and move users in cohorts.

flowchart TD
    L["Login request"] --> R{"Routing layer:<br/>which provider<br/>for this user/tenant?"}
    R -->|"not yet migrated"| O["Old provider"]
    R -->|"migrated"| N["New provider"]
    O --> S["Application session"]
    N --> S

A routing layer decides, per user or per tenant, which provider handles authentication. You migrate a cohort, watch the error rates, and keep rollback available by flipping the flag back.

Be honest about the cost. During the transition you are operating two identity platforms, paying for both, and maintaining a routing layer that is now the most critical component in your authentication path. That layer needs the same care as anything else on the login path — it is a single point of failure you introduced deliberately.

Session and token bridging is the related decision. Sessions issued by the old provider validate against its keys. You can either accept both issuers during the transition — validating against both JWKS endpoints, which is more code and more surface — or force a re-login at cutover.

Forced re-login is simpler and it's a visible event. Support will hear about it. For most teams that's the right trade, especially if you schedule it rather than letting it happen to people mid-task.

The order that works

  1. Stand up the new provider in parallel. No traffic. Get federation, branding, and policy configured.
  2. Migrate internal users and test tenants. You'll find the first three problems here, cheaply.
  3. Migrate self-service users in cohorts, smallest and least critical first, with rehash-on-login doing the work.
  4. Migrate enterprise federated tenants last. These need scheduling with someone else's IT department.
  5. Execute the dormant-account policy you decided on in advance.
  6. Decommission — and only now remove the legacy verification path.

Why it takes a quarter

None of the above is what makes migrations slow.

Each enterprise customer running SAML has to update their identity provider configuration. That means a ticket in their system, a change request, a security review of the new metadata and certificate, and a maintenance window on their calendar. For a customer with a quarterly change freeze, you're waiting a quarter. For one where the person who understands their IdP left last year, you're waiting longer.

Multiply by the number of federated customers, run them mostly in parallel, and accept that the slowest one sets your completion date. This is the schedule. The hashes are a week of engineering; the federation coordination is the project.

Start those conversations at the beginning, not when you're ready. The lead time is the deliverable.

When not to do this

Migration consumes a large fraction of a platform team's capacity for a quarter or more, and it puts sustained risk on the one system where failure is total. That cost is worth paying for structural reasons and rarely worth paying for irritation.

Good reasons: the architecture genuinely doesn't fit and won't; lock-in that's actively constraining your product; the provider is exiting your market or your segment; a compliance or data-residency requirement you cannot satisfy where you are.

Weak reasons: cost, unless the delta is large enough to fund the migration several times over — and remember you'll be paying both bills during the transition. A missing feature, unless it's blocking revenue. General frustration with a vendor's support, which is real but is usually cheaper to solve commercially.

Run the arithmetic honestly, including the quarter of roadmap you won't ship.

The part you can do today

The best time to make a migration possible is when you aren't having one.

Own your user identifiers, and treat the provider's sub as an attribute rather than a primary key. Keep identity data out of your application database and application data out of your identity store, so the two can be separated later. Be deliberate about building on provider-proprietary features you couldn't reproduce elsewhere — sometimes they're worth it, but that should be a decision rather than a default. And check, before signing, whether you can export your password hashes.

None of that costs anything meaningful up front. All of it is the difference between a migration that takes a quarter and one that never gets approved at all.