Migrating Password Hashes Without Resetting Every User
The query that starts this project is four lines long:
The query that starts this project is four lines long:
SELECT substring(password_hash from 1 for 7) AS prefix, count(*)
FROM users GROUP BY 1 ORDER BY 2 DESC;
And the result, on a system that's been alive for eleven years, looks something like this:
$2a$08$ 4,412,806
$2a$10$ 883,145
$2b$12$ 210,334
$2y$08$ 41,208
(32 hex) 612
Four million accounts at bcrypt cost 8 — a work factor chosen in 2014 when it was defensible and never touched since. Two hundred thousand at cost 12, because someone raised the default for new signups in 2021 and correctly didn't try to fix the old ones. A few thousand $2y$ rows from a PHP service absorbed in an acquisition. And 612 rows of bare hex that nobody can identify, from before the current schema existed.
You want all of it on Argon2id. You cannot do it, because you don't have anyone's password.
That constraint is total. There is no function mapping bcrypt(p) to argon2id(p); if there were, bcrypt would be broken. You can't batch-convert overnight, you can't do it in a maintenance window, and there's no vendor tool either — anyone selling one is selling a login proxy that waits for users to type their passwords, which is the technique below with a bill attached.
What you can do is exploit the one moment plaintext exists: the few milliseconds during a successful login when it sits in a request-scoped variable on its way to a verify function. That moment is the entire migration. Everything else is bookkeeping around it, plus one genuinely hard decision about the accounts that never come back.
The mechanics of rehash-on-login are sketched in the provider-migration article on this blog; the point there was that migration is possible. The point here is what happens after the easy 70% converts.
Self-describing hashes are the whole reason this is tractable
Before the algorithm, the storage format — because a migration is only survivable if every record can say what it is.
bcrypt's output is already self-describing. $2b$12$LQv3c1yqBWVHxkd0LHAkCO... encodes the variant, the cost, a 22-character radix-64 salt, and the 31-character digest, in exactly 60 bytes. The variants are a common source of confusion: $2a$ is the original, $2y$ is PHP's crypt_blowfish marking a fixed implementation, $2b$ is OpenBSD's fix for a length-counter overflow on passwords over 255 bytes. For any password you'd plausibly see, all three produce identical digests — but a library that refuses to parse a prefix it doesn't recognize will throw on 41,208 of your rows, and that exception gets caught somewhere as "wrong password."
Argon2 uses the PHC string format, which is the same idea generalized:
$argon2id$v=19$m=19456,t=2,p=1$c29tZXNhbHQ$RdescudvJCsgt3ub+b+dWRWJTmaaJObG
Variant, version, all three cost parameters, salt, digest — everything a verifier needs, in one string. Two consequences shape the whole project.
Store one column, not five. The temptation is to normalize: algorithm, cost, memory_kib, salt, digest. Resist it. A single password_hash TEXT column holding a complete PHC-style string means a rehash is one atomic write of one value. Split across five columns, a rehash is a multi-column update, and a crash partway through leaves a row whose declared algorithm doesn't match its digest. That record isn't "wrong" — it is unverifiable. No algorithm will ever accept the user's correct password again, and no support flow will diagnose it, because from the outside it's indistinguishable from a person who forgot.
Tag the untaggable. Those 612 bare-hex rows have no prefix, and you must not infer the algorithm from length — 32 hex characters is MD5, and md5(md5(p)) looks identical. Find whichever ancient service wrote them, then do a one-time UPDATE to prefix them with an explicit marker of your own ($legacy-md5-2011$...). Guessing at verification time is how you end up with a default: branch in a hash dispatcher, which is a way of saying "some users will be locked out and we won't know which."
The ordering, and the two details that aren't obvious
stored = load(user)
algo = parse_algorithm(stored) # never inferred
if not verify_with(algo, password, stored):
return FAILURE # nothing is written on failure
if needs_rehash(stored, current_policy):
new = argon2id_hash(password) # plaintext exists only here
update users set password_hash = :new
where id = :id and password_hash = :stored
return SUCCESS
Two things in that sketch are load-bearing and routinely omitted.
The rehash condition is not algo != argon2id. It's "does this record match current policy," which also covers an Argon2id record hashed at m=12288 when policy has moved to m=19456. Build it that way from the start and the next parameter bump is a config change instead of a project. Most libraries hand you this: PHP's password_needs_rehash, argon2-cffi's check_needs_rehash, passlib's needs_update.
The write is conditional on the old value. Two concurrent logins from the same user — a phone and a laptop, or a retrying client — will otherwise both compute a fresh hash and race. The WHERE password_hash = :stored guard makes the loser a no-op rather than a lost update, and it means a rehash can never clobber a password change that landed in between. That's the real bug: the user changes their password in tab A while an in-flight login in tab B rehashes the old one back over it.
The rehash is also optional work on a successful login. If it throws, if the database is degraded, if you're shedding load — log it and return success anyway. A failed rehash means one account converts next time. A rehash allowed to fail the login means an outage caused by a migration nobody consented to.
And login isn't the only place plaintext appears. Registration, password change, and reset all have it and should all write the current-policy hash; reset in particular does your migration for free on every account that goes through support.
The record's real state machine
stateDiagram-v2
[*] --> Legacy: imported / historical
Legacy --> Native: successful login (rehash)
Legacy --> Wrapped: bulk sweep of dormant accounts
Wrapped --> Native: successful login (unwrap)
Legacy --> Native: password reset
Wrapped --> Native: password reset
Legacy --> Dead: forced expiry
Wrapped --> Dead: forced expiry
Native --> Native: parameter bump
Legacy → Native is the easy path and it converts your active population without anyone noticing. The two states worth the rest of this article are Wrapped and Dead.
Wrapped hashing, and the part everybody gets wrong
Three months in, the curve flattens. Your monthly-active population is converted; roughly 45% of all accounts still carry cost-8 bcrypt, and the daily conversion rate has fallen to a few hundred. Rehash-on-login will never reach them.
Wrapped hashing — layered or nested hashing — applies the new KDF on top of the old one's output, without plaintext:
wrapped = argon2id( bcrypt(password) )
You can compute this for every dormant row in a batch job tonight. Every wrapped record now has Argon2id's cracking cost, immediately, for accounts that have not logged in since 2019.
Here's the part that gets written down wrong almost everywhere: you must delete the inner digest and keep the inner salt.
If you store the full bcrypt string alongside the Argon2id wrapper "so you can reconstruct it," you have gained nothing. The attacker who steals your database ignores the Argon2id layer entirely and cracks the cost-8 digest sitting right next to it.
But verification must recompute the inner digest from the submitted password, and bcrypt is salted — so it needs that salt and cost. The record therefore keeps the inner parameters and destroys the inner output:
stored:
wrap = bcrypt$2a$08$LQv3c1yqBWVHxkd0LHAkC # variant, cost, salt — no digest
outer = $argon2id$v=19$m=19456,t=2,p=1$...$...
verify(password):
inner = bcrypt(password, salt="LQv3c1yqBWVHxkd0LHAkC", cost=8) # full 60-char string
return argon2id_verify(outer, inner)
on success:
store argon2id(password); drop wrap # unwrapped
That is the whole trick, and it has four costs you are committing to permanently.
You must pin the encoding of the intermediate, exactly. Do you feed Argon2id the full 60-character bcrypt string, or only the 31-character digest portion? Either works; only one is what you did. Write it in a comment above the function and in a design doc, because in four years someone will reimplement this in another language, and a one-character difference locks out every wrapped account with no diagnostic.
Verification cost stacks. A wrapped login pays bcrypt cost 8 plus Argon2id — roughly 20 ms plus 45 ms on a modern core, against 45 ms native. Wrapped records are your dormant tail, so it's a small share of live traffic, but it's permanent until each one unwraps.
You can never delete the bcrypt implementation. Not the library, not the code path, not the tests. That's a dependency you must patch forever for the sake of accounts unused in six years.
It is one-way. You destroyed the inner digests. If you later discover your Argon2id parameters were misconfigured, or that the batch job's encoding differed subtly from the verifier's, the affected accounts are unrecoverable except by reset. Run the batch on 0.1% of rows and verify a sample against known credentials in a staging copy before touching the rest.
Peppering is the option most teams should reach for first
Wrapped hashing solves "the stored value is too cheap to crack." An adjacent technique solves the same problem with better trade-offs for most teams.
A pepper is a secret living outside the database. The version worth your attention isn't the HMAC-before-KDF variant — that's also a one-way commitment, since rotating it requires plaintext. It's the encrypt-after-KDF variant:
stored = AES-GCM(key_in_kms, argon2id(password))
Both make a database-only breach useless. But the wrap is irreversible, while encrypt-after-KDF is a bulk transformation you can undo, rotate, or re-key at any time, in a batch job, with no plaintext and no user involvement. If your key policy changes, if you move clouds, if the scheme was a mistake — you decrypt and you're back where you were.
The costs differ in kind. You've added a key-management dependency to the login path — an HSM or KMS call per verification, meaning latency, availability, and per-operation cost — unless you cache the key in application memory, and a key in application memory is one heap dump away from the database it was protecting. Key loss is total: every peppered account becomes permanently unverifiable, which makes your key backup story the real security boundary of your password system.
The honest comparison: peppering defends against a database breach, the overwhelmingly common case — SQL injection, a leaked backup, an exposed replica. Wrapping additionally survives full compromise including your keys, and only by making offline work expensive rather than impossible. If your threat model is the common case, and it probably is, the reversible option is worth more than the irreversible one.
Work factor is a capacity decision, and you will feel it
The parameters are not just a security choice. Current OWASP guidance for Argon2id lands around m=19456 KiB, t=2, p=1 or m=12288, t=3, p=1; RFC 9106's second recommended option is 64 MiB with t=3, p=4. Pick from that range, then do the arithmetic — because the memory parameter has a property bcrypt's cost factor doesn't.
At m=19456, every concurrent verification holds 19 MiB. Two hundred concurrent logins is 3.8 GB resident, touched in a random-ish pattern. That isn't a CPU number you can average away; it's a hard concurrency ceiling per box, and it evicts the L3 cache of every co-tenant workload on the machine. Teams discover this as "our unrelated API's p99 got worse when we shipped Argon2id," and it's one of the better arguments for running authentication on its own fleet.
Illustrative single-core figures — re-measure on your own hardware, they vary threefold across CPU generations: bcrypt cost 8 ≈ 20 ms, cost 10 ≈ 80 ms, cost 12 ≈ 300 ms (each increment doubles); Argon2id at m=19456, t=2, p=1 ≈ 45 ms. At 400 logins/second peak, moving from cost 8 to Argon2id takes you from 8 core-seconds per second to 18 — roughly double the hashing fleet — and during the migration window you pay both, since a converting login runs the old verify and the new hash.
The attacker-side asymmetry is what you're buying. Public GPU benchmarks put a high-end consumer card at roughly 180,000 bcrypt hashes/second at cost 5, halving per increment: about 23,000/s at cost 8, two billion guesses a day from one card. A rules-based run against a billion-entry wordlist finishes overnight. Argon2id at 19 MiB is memory-bandwidth-bound rather than compute-bound, and the same card manages low thousands per second. Four orders of magnitude, and that gap is the entire justification for the project.
Failure modes worth rehearsing before you ship
The thundering herd on the flag flip. Turn rehash-on-login on globally at 09:00 Monday and every login in your busiest hour does verify-plus-rehash, with the Argon2id memory footprint arriving all at once. Ramp by a stable hash of the user ID — 1%, 5%, 25%, 100% — and bound concurrent rehashes with a semaphore that skips the rehash when saturated rather than queuing behind it.
Rollback that isn't. Once a record is Argon2id, a deploy of the previous release — which has no Argon2id verifier — locks that user out. Ship the reader first: release N verifies both algorithms and still writes bcrypt; release N+1 starts writing Argon2id. This is ordinary expand/contract discipline and people consistently forget it applies to password hashes, because hashes don't live in a migrations directory.
bcrypt's 72-byte truncation, and the trap in fixing it. bcrypt ignores everything past 72 bytes of input. If you allow passphrases, a user whose real password is 80 characters has been authenticating on the first 72 for years. Rehash them to Argon2id and you store the full 80 — they'll still log in fine, but a truncated variant saved in an old password manager entry stops working. Rare, real, and it arrives as an unreproducible support ticket.
If you're staying on bcrypt with pre-hashing to defeat the limit, the standard construction is bcrypt(base64(HMAC-SHA256(key, password))). Two details: base64 because a raw digest can contain a NUL byte and C implementations truncate there; and HMAC with a secret key rather than a bare sha256(), because an unkeyed fast pre-hash enables password shucking — if the same password appears in another site's leaked MD5 or SHA-1 corpus, the attacker tests your expensive bcrypt against those known digests directly instead of guessing passwords at all. The keyed version makes that impossible.
Changing two things at once. If you've never Unicode-normalized password input, do not add NFC normalization in the same release as the algorithm change. Normalization alters the bytes, which alters the effective password, and you won't be able to tell which change caused the failures. Separate deploys, weeks apart.
Timing disclosure during the window. Login latency now reveals which algorithm a record uses, and if you skip hashing entirely for unknown usernames it reveals account existence. Verify a dummy hash of your most expensive configured algorithm on the not-found path.
Deciding what to do with the accounts that never come back
Six months in, the curve has asymptoted. Perhaps 78% of accounts that authenticated at all this year are Native, and 30% of all accounts are still Legacy or Wrapped. That won't improve materially, and the decision has to be made rather than deferred — deferring is itself a choice, just one nobody signs.
The policy reflex is "force a reset after a deadline." Sometimes that's right. Reason about it on two axes instead.
What does the account control? An account holding a payment method, an address book, order history, or admin rights is worth an attacker's cracking budget. An account created once to download a whitepaper is not. Segment before deciding — a blanket policy is how you generate 200,000 support contacts to protect accounts nobody would attack.
What is the exposure if the database leaks tomorrow? Here the GPU numbers become a decision input rather than trivia. Cost-8 bcrypt over human-chosen passwords is not a meaningful defense against an attacker with a rig and a week; wrapped or peppered records are fine indefinitely. So the question collapses to: are the remaining Legacy records worth protecting, and if so, is cheap bulk protection enough, or does the password need to be gone entirely?
That gives three defensible positions. Wrap or pepper the tail and keep it — best where accounts hold residual value and reset friction is expensive; you carry a permanent legacy code path or a KMS dependency. Force reset on a deadline — best where accounts are valuable and you can absorb the support volume; note this reads to users as a security incident whether or not you say so, so expect the "were you breached?" thread. Expire the accounts — clean, and legitimately right for dormant free-tier accounts with no financial or social graph, but it's a product decision about deleting customer relationships, not an engineering one, and it needs a named owner.
The wrong answer is the fourth, which is what most organizations actually ship: leave the tail alone, leave the legacy verifier in place, and let the arrangement quietly become permanent. That's a defensible outcome — but only if someone chose it.
A timeline with exit criteria
- Weeks 0–2. Instrument first. Emit a counter per verification tagged with the stored algorithm and parameters, and build the histogram from the top of this article as a dashboard. Tag every untagged legacy row explicitly. Ship the reader-only release.
- Week 3. Enable rehash-on-login for 1% of user-ID buckets. Watch login p99 and hashing-fleet memory, not just CPU.
- Weeks 4–6. Ramp to 100%. Add the same rehash to registration, password change, and reset.
- Months 2–4. Watch the curve. A weekly-login product converts 70–80% of its active population in the first month; a product people open monthly takes a full quarter to reach the same point.
- Month 4. The tail decision, with product and support in the room. If wrapping or peppering, pilot on 0.1% and validate against known credentials before the full sweep.
- Months 5–9. Execute. Keep the legacy path behind a flag with a written kill date.
Measure percentage of daily successful logins served by a legacy verifier, not percentage of rows. The row count includes accounts that will never authenticate again and drags your number down forever; the login-served number reflects live exposure and goes to near-zero. And exclude accounts with no password at all — SSO and social-login users are not in the denominator, and leaving them in has convinced more than one team its migration stalled at 60% when it had finished.
"Done" is: legacy verifications below 0.1% of daily logins for thirty consecutive days, every remaining Legacy row covered by a written tail policy that has been executed, and a dated ticket to delete the old verifier. Not "zero bcrypt rows in the table." That number will never be zero, and waiting for it is how a nine-month project becomes a permanent one.