Why Security Questions Were Always a Bad Idea
The reset queue I was reading came from a mid-sized consumer product in 2019, but the flow it described had been shipped around 2008 and never touched since. Three questions at registration from a dropdown of eight, two required at reset, case-insensitive match, whitespace trimmed, five attempts before a soft lock.
The ticket that stopped me was from a woman who had held the account for eleven years. She knew her mother's maiden name. She could not reproduce what she had typed into the box in 2008. Was it capitalised? Was it the maiden name or the married name — her mother had remarried in 2004 and she wasn't sure which had been true in her own head at signup. She burned all five attempts, hit the lock, waited twenty-four hours, burned five more, and wrote a paragraph to support that ended the way these always end: I am the person. I have been the person for eleven years.
She was right, and the system had no way to agree with her.
Now hold that against a scene from the autumn the flow was written. In September 2008, someone got into Sarah Palin's Yahoo! account during a presidential campaign. There was no exploit. He opened the ordinary password-reset page, supplied her date of birth and ZIP code — both published — and answered the question standing between him and the mailbox: where did you meet your spouse? Wasilla High School. It was in her biography, and by contemporaneous accounts the whole thing took about forty-five minutes of searching. The control performed exactly as specified. It asked for a fact about her life, and it let in the person who had looked the fact up.
Those two scenes are the same control, and they are why this article exists.
A security question is a shared secret with five properties, each of which is individually disqualifying: an extremely small keyspace, a distribution that is publicly correlated rather than uniform, no rotation, no revocation, and a mandatory fuzzy-matching layer that you are required to build in order for the thing to work at all. Put together, they produce an authenticator that is harder for the legitimate user than for the attacker — a control operating with its sign inverted. And unlike almost everything else in this series, that was true on day one. Nothing changed to make security questions bad. What changed is that the excuse for shipping them anyway ran out.
This is the awkward entry in Security Advice We No Longer Believe, because the honest reconstruction doesn't end with "it was right for its era." It ends with "it was never right, and here is why intelligent, careful people shipped it in enormous volume regardless" — a more interesting question, with a transferable answer. So: the five properties and the arithmetic behind each; the inversion, with the published evidence and an honest account of its limits; the generous reconstruction of 1998; the narrow case where knowledge-based verification still earns its place; and then the part most articles skip, which is how to get them out of a live system without locking out the woman with the eleven-year-old account.
I'll stay out of covered ground. The general shape of the recovery trade-off is The Hardest Problem in Identity Isn't Authentication, It's Recovery; the modern version is Passkey Account Recovery Is the Whole Problem. Both dismiss security questions in a paragraph. This is the article behind the paragraph.
Five properties, and why each one is fatal on its own
There's a definition of a usable shared secret that the rest of your platform quietly relies on: a value that is (a) drawn from a space too large to search, (b) known to a bounded set of parties, (c) replaceable, (d) revocable, and (e) comparable exactly. A password satisfies all five — badly on (a), well enough on the rest. A private key satisfies all five superbly. A security question answer fails all five, and the failures compound in ways that aren't obvious from the list.
Readers who know Why Biometrics Aren't Secrets will recognise the structure, and the resemblance is the finding rather than a coincidence. A security question answer is an unrevocable, publicly-observable, fuzzily-matched fact about a person's life history. It is a biometric made of words, deployed without any of the architecture that makes a biometric safe.
1. The keyspace is tiny — and, worse, it isn't uniform
The naive defence of security questions is that a surname or a city name is drawn from a large set. There are hundreds of thousands of surnames in use in the United States and tens of thousands of populated places. Twenty bits, easily. Fine, surely?
No, and the reason is the single most important idea in guessing security: the relevant measure is not the size of the set, it's the shape of the distribution, and for a guessing attacker the shape that matters is the top of it.
Shannon entropy averages over the whole distribution and is close to useless here. An attacker with a small guess budget doesn't sample the distribution; they attack its mode. The right measures are min-entropy — -log₂ of the probability of the single most likely value — and, for budgets larger than one, the partial-guessing metrics Bonneau formalised in his work on password distributions (α-guesswork and friends). Every one of them says the same thing about human-fact distributions: the head is brutal.
Work the numbers on the classic question. In the 2010 US Census surname data, the most frequent surname covers a bit under 1% of the counted population and the top ten cover a few percent between them, so min-entropy for a US-population guess at mother's maiden name is around 7 bits — the attacker's single best guess wins about one time in 120. That is already worse than a four-digit PIN. But the population figure massively overstates the difficulty, because surname distributions are among the most strongly ethnically and geographically correlated variables that exist. Restrict to a Korean-American cohort and three surnames cover something close to half the population; restrict to a Vietnamese-American cohort and one surname covers roughly two in five. The attacker who knows nothing but your name and your city has, in many cases, reduced a "20-bit" secret to two or three bits.
That correlation property is the one people miss. Security question answers aren't merely low-entropy; they're low-entropy conditional on cheaply-observable attributes of the target, which is a strictly worse thing to be. Everything else on the standard list behaves the same way:
| Question | Naive keyspace | Realistic min-entropy | What collapses it |
|---|---|---|---|
| Mother's maiden name | ~17 bits (US surname set) | ~7 bits population-wide; 2–4 bits given ethnicity/region | Surname distributions are heavily correlated with demographics you can observe |
| First pet's name | ~12 bits | 5–7 bits | Naming fashion, correlated with the user's age cohort |
| City of birth | ~15 bits | 4–6 bits | Population-weighted; correlated with current location |
| High school | ~13 bits | 2–3 bits, often 0 | Determined by childhood address; usually published by the user |
| Favourite food | large | 2–3 bits | "Pizza" |
| Favourite colour | ~4 bits nominal | 1–2 bits | "Blue" |
The min-entropy figures in that table are my estimates from published distributional data, not measurements on your users; treat them as order-of-magnitude. The direction is not in dispute, and the direction is the whole argument: these are single-digit-bit secrets defended by a five-attempt budget. A four-digit ATM PIN is 13 bits and is protected by a hardware attempt counter and a captured card. Most of these are weaker than the PIN and protected by an HTTP endpoint.
2. It was never secret
Entropy would be the smaller problem if the attacker had to guess. Frequently they don't. Mother's maiden name is on marriage and birth records that are public in many jurisdictions and sold in bulk by people-search brokers for a few dollars; city of birth, high school, and first employer are on the professional profile the user maintains deliberately; first pet's name and childhood street are in the photo captions. And every couple of years a viral post circulates asking people to reply with exactly that list, framed as nostalgia. It doesn't need to be malicious — the replies are public either way. There is no comparable social ritual in which people post their passwords.
Ariel Rabkin made this concrete in a 2008 SOUPS paper that surveyed the questions actually in use at US banks and classified them by how findable the answers were. His conclusion — in the year the flow at the top of this article was designed — was that a large share of the questions in production were answerable from public records or from a social network profile. The critique was published, specific, and available before most of the deployments still running today were built.
The design consequence is the same one the biometrics article makes: any security argument containing the step "and the attacker does not know the user's mother's maiden name" is unsound at the first line, and no amount of care in your implementation repairs it.
3. It cannot be rotated
You can change the answer you have stored. You cannot change the fact.
This sounds like a quibble and it isn't, because rotation is how incident response terminates. "We rotated everything in scope" is the sentence that closes a postmortem. Consider what it means here. Your answers table leaks; you force every user to set new answers; every user now faces one of two options. Set a truthful answer to a different question — the new fact is drawn from the same distribution with the same properties and is discoverable in the same places, so you have re-issued the weakness rather than rotated away from it. Or set an untruthful answer — which introduces a memorised, arbitrary, unrehearsed string, the failure mode Stop Asking Users to Remember Secrets is entirely about, in its worst possible instance: retrieved once, years later, under stress, with no rehearsal history.
The asymmetry with passwords is the point. When a password leaks the user picks a genuinely new one from a genuinely large space, and the new one is exactly as good as the old one was. There is no equivalent move here. The pool of facts about a person's early life that are simultaneously stable, memorable, and unknown to the internet is close to empty, and it does not refill.
4. It cannot be revoked, and it is shared across every site that ever asked
This is the property with the largest blast radius, and it is what makes a breach of security questions structurally worse than a breach of passwords.
Password reuse is a behaviour. It's extremely common — that's the whole economics of credential stuffing — but it's in principle correctable, and a password manager corrects it completely, per site, for free. Security question answers are not reused by choice; they are reused by construction, because the question set is standardised across the industry and the answers are facts. Your mother has one maiden name. Every site that asked got the same string, and there is no manager that can generate a different mother for each relying party. When one site's answer table leaks, every other site that asked the same question is compromised, permanently, for that user, and no action by that user or by any of those sites undoes it.
The tables do leak, in exactly the worst form. Yahoo's disclosures around the 2013 and 2014 breaches stated explicitly that the compromised data included security questions and answers, some of them unencrypted. That's not an outlier implementation, and I'll come back to why plaintext storage is the norm rather than the exception — the reason is structural.
Then there's the industrial-scale version. Out-of-wallet KBA — the "which of these four addresses have you lived at" style used for identity verification — is backed by credit bureau data. In 2015 the IRS's "Get Transcript" service was drained by attackers who answered those questions correctly at scale; the initially reported figure of around 100,000 taxpayer accounts was revised upward twice, ultimately to more than 700,000. Two years later, Equifax — one of the very sources that populated those questions — disclosed a breach affecting around 147 million people.
Read those two events together and the conclusion is not "KBA was implemented badly." It's that the entire category has a single, shared, unrotatable key, and that key was published. The population's answers to knowledge-based questions are now, for practical purposes, a known plaintext. That is a state from which there is no recovery, because nobody can be issued a new childhood.
5. You are required to build a fuzzy matcher, and it is a security hole with a business case
This property gets the least attention and is the most instructive, because it's the one where the mechanism forces the defender to widen the attack surface.
Exact string comparison does not work on human-fact answers. Rex, rex, REX, Rex, Rex the dog and Rexy are one answer as far as the user is concerned, and rejecting four of them gives you a catastrophic false-reject rate that your support queue eats. So every production implementation normalises — case folding, whitespace trimming, punctuation stripping — and most go further: collapsing internal whitespace, reconciling St. with Street, sometimes accepting an edit distance of one or two, sometimes awarding partial credit across three questions. Every one of those relaxations was added for a good reason, in response to a real number on a real dashboard, by an engineer correctly trying to reduce lockouts.
The arithmetic is worth deriving, because "fuzzy matching weakens it" is usually just asserted. Let P be the distribution of true answers over strings. Your matcher doesn't compare strings; it compares equivalence classes — normalisation partitions the string space, and a guess succeeds if it lands in the same class as the stored answer. So the attacker's single-best-guess probability isn't max P(s) over strings, it's max Σ P(s) over classes. Relaxing the matcher merges classes, and merging classes can only increase the mass of the largest one.
Every relaxation strictly increases the attacker's single-guess success probability and never decreases it. That's a one-way ratchet, not a trade-off with an optimum.
The trade-off is real, though — the relaxation also reduces false rejects, which is why it's there. But look at who redeems it. The legitimate user benefits from the merging of their own class. The attacker chooses which class to redeem the benefit against, and picks the largest. Merging the variants of Rex helps the one user whose dog was Rex; it helps the attacker against every user whose dog was Rex, and Rex sits near the top of the distribution precisely because it's a popular name. The relaxation is worth more to the attacker than to the average user, because the attacker gets to pick where to spend it and the user doesn't.
Partial credit is the extreme case and deserves naming as a bug. "Two of three correct" against three questions of roughly 5 bits each does not give you 15 bits, or 10; it gives the attacker three independent lottery tickets and a tolerance for losing one. If you have it, you have a control substantially weaker than a single question with exact match — and it will have been added by someone reducing a support metric.
There's a second-order consequence that lands in the storage layer, which I'll take up in the migration section: a fuzzy matcher is very hard to reconcile with hashing at rest, and that tension is the direct cause of the plaintext-storage epidemic.
The inversion, stated precisely
Now put the five properties together and you get a result stronger than "weak credential."
The claim is: for a meaningful share of the population, a security question is easier for a motivated attacker to pass than for the account's actual owner. Not marginally. Measurably, in published studies, on real recovery attempts.
The two axes are actually one axis
Start with the structural reason, because it explains why no amount of question-set curation fixes this.
For a memorised item to be reliably retrievable years later without rehearsal, it must attach to structure the person already has — that's the chunking property, and it's why passphrases beat character soup. But an answer is memorable because it is a salient, stable, frequently-rehearsed fact about the person's life. And a salient, stable, frequently-rehearsed fact about a person's life is precisely the kind of fact that has been recorded, published, told to colleagues, entered on forms, and posted with a photograph.
Memorability and unguessability are not two dimensions you can optimise independently. They are both monotone functions of the same variable: how strongly the answer is determined by widely-shared facts. Turn one up and the other goes down. The quadrant you want — memorable and unguessable — is not underpopulated. It is empty by construction. This is the same tension Longer Beats Complex finds inside passphrases, except there it's a soft trade you can win with more words. Here there's no dial. The answer has to be one fact, and it either is or isn't the kind of fact the world knows.
What the measurements say
The best public data comes from Google. Bonneau, Bursztein, Miers, Mirian and Pullman published "Secrets, Lies, and Account Recovery" at WWW 2015, drawing on hundreds of millions of real recovery attempts across Google's actual user base — not a lab study, the production flow. I'll describe the shape rather than quote precisely, because the figures vary a lot by question and language cohort and you should read the paper:
- Single-guess attacks work far too often. For some common questions the modal answer covered close to a fifth of respondents in a given language cohort — the reported figure for English speakers' "favourite food" was in that region, and yes, the answer was pizza. One guess.
- Ten guesses reached something in the region of four in ten users for certain question-and-cohort combinations, including city-of-birth questions in cohorts with concentrated name and place distributions.
- Legitimate users failed their own questions at a startling rate. Recall for security answers came in around 60%, against roughly 80% for SMS reset codes and about three-quarters for email links, on the same population attempting the same task.
- Users who tried to defend themselves made it worse. A substantial fraction admitted entering deliberately false answers — and the falsified answers clustered, because people invent from the same small pool of jokes and defaults. Falsification hurt recall sharply and raised guessing difficulty far less than those users believed.
Microsoft Research reached compatible conclusions six years earlier. Schechter, Brush and Egelman's "It's No Secret" (IEEE S&P, 2009) ran the acquaintance threat model directly: participants nominated people who knew them, and those acquaintances guessed. Roughly a sixth of answers fell within five attempts, and around a fifth of participants had forgotten their own answers within six months.
Two research groups, six years apart, different methodologies and populations, one lab-based and one drawn from planetary-scale production telemetry. Both found the same shape: the attacker's pass rate is uncomfortably close to the owner's, and for some question-and-cohort combinations it is higher.
Be precise about what that establishes. It establishes the inversion for statistical attacks on common questions in large populations — the case that matters at consumer scale. It does not give you a per-account probability for your product, and I'd be sceptical of anyone quoting one. What it means is that the burden of proof has moved: the evidence has been against this control for a decade, and if you still run it you should be able to state your own numbers.
The escape hatch, and why it isn't one
The sophisticated user's response — I've made it myself, and it's wrong — is: generate a random string, store it in the password manager, and the question becomes a second strong credential.
You have just created an unmanaged password with no reset path. Lose the manager entry and there is no "forgot my mother's maiden name" flow, because this mechanism is the forgot-my-password flow; it's the terminal node. And the failure is correlated: the entry lives behind the account you're locked out of, or on the device stolen in the same event that made you need recovery. That's the correlated-failure question from the passkey recovery article, and the answer is no.
You have also put a high-entropy string behind a fuzzy matcher, so the normalisation now works against you — some implementations strip the very characters carrying your entropy, and you find out at reset time.
Most importantly, it doesn't generalise. A defence that works only for users who run a password manager and think about threat models isn't a defence of the control; it's a workaround by a population that could have used any better mechanism instead. A control is only as good as its behaviour on the median user, and the median user answers truthfully.
Why smart people shipped it anyway
Everything above was, in outline, knowable in 1998. So why did essentially the entire industry deploy this, for two decades, including organisations with excellent security teams?
The generous reconstruction has two parts, and only one of them is an excuse.
The constraints were genuinely different
Put yourself in a 1999 web team's position and enumerate the recovery channels available.
Email was not a reliable recovery channel. This is the fact modern engineers most consistently fail to internalise. In 1999 a very large share of consumer email addresses were bound to an ISP — @aol.com, @earthlink.net, @freeserve.co.uk — and changing ISP meant losing the address, which people did constantly. Free webmail was new and of unproven durability. Delivery was unreliable, spam filtering was primitive and aggressive, and a meaningful fraction of your users' addresses were dead within a year. Building recovery on email in 1999 meant building it on a channel that would silently evaporate for a large minority of your users. Email's status as a universal, portable, durable identity anchor is a recent property.
SMS was not available and not free. Mobile penetration was well short of universal, aggregator integration was a bespoke commercial contract per country, delivery was best-effort, and every message cost real money against a business model that often had no revenue at all. Why SMS OTP Won't Die covers how long the cost-and-coverage argument kept binding after the security argument was settled.
Everything else you'd reach for now did not exist — no smartphones, no push, no TOTP at consumer scale, no device recognition worth the name, no risk engines, no federation, no widely-deployed second factor of any kind. And the realistic alternative was frequently nothing at all. Not a weaker flow: lose your password, lose the account.
Against that, a mechanism costing zero — no infrastructure, no vendor, no per-use fee, works internationally, works for users with no phone and users whose email address died — is not stupid. It's people making a trade with the instruments available, the same trade that made 90-day rotation reasonable in 2003.
And it wasn't measurable. Nobody could evaluate this control because nobody had the data. There was no corpus of real answers to compute a distribution over, for the same reason there was no corpus of real passwords until RockYou in 2009 — you could not lawfully obtain one. And nobody instrumented the other half either: I have never seen a pre-2010 system that reported recovery success rate disaggregated by mechanism, which is the single number that would have exposed the inversion immediately. The control was silent, and as The Password Policy Arms Race argues at length, silent controls survive on narrative.
The part that isn't an excuse: it was a transplant
Here's the more interesting half, and the one I think is genuinely under-articulated.
Knowledge-based authentication did not originate on the web. It came from banking and telco call centres, where it had been in use for decades and where it worked acceptably well. Web teams imported it because it was the established practice for verifying a remote human, and it looked like a solved problem with a known implementation.
But look at what a call-centre verification actually was, as a control system:
flowchart TB
subgraph CC["Call centre KBA, c. 1995 — a control with six components"]
direction TB
Q1["The question<br/>(one weak signal)"]
Q2["A live human agent<br/>who hears hesitation"]
Q3["Voice + caller ID +<br/>callback to a number on file"]
Q4["Transaction history<br/>the caller must also know"]
Q5["A fraud team monitoring<br/>agent-level patterns"]
Q6["Financial reversibility:<br/>chargebacks, insured losses"]
end
subgraph WEB["Web self-service reset, c. 2005 — what got imported"]
direction TB
W1["The question"]
end
CC ==>|"transplanted"| WEB
style Q1 fill:#fde8e8,stroke:#c66
style W1 fill:#fde8e8,stroke:#c66
style WEB fill:#fff,stroke:#c66,stroke-width:2px
The question was never the control. It was one input to a control whose strength came from a human on a live line, an out-of-band callback, transaction data the caller had to corroborate, a fraud team watching for agents who approved too much, and — the piece that made the arrangement economically tolerable — the fact that if it failed, the money could be moved back.
Web self-service took the question, discarded the other five components, added unlimited parallelism and global reach, and ran the resulting fragment as a sufficient condition for account takeover.
That's the error, and it isn't specific to security questions. A control lifted out of its enforcement environment is not a weaker version of that control; it is a different control, and nobody has evaluated it. It's the same mistake as moving a biometric matcher out of the secure element and running it server-side: identical mechanism, no attempt budget, completely different security.
So the honest verdict on S9 is the sharp one: security questions were defensible on availability grounds in 1999 and never defensible on security grounds at any point. What changed isn't the security — that was always this bad — it's that every availability argument has expired. Email is durable and near-universal. SMS is cheap and reachable. Push, TOTP, passkeys and platform authenticators exist. Delay-plus-notification is free. The justification was "there is nothing else," and there is now a great deal else.
Which is why this one doesn't get the generous ending the rest of the series gets. The advice wasn't superseded. The excuse was.
Where knowledge-based verification still earns its place
I want to be careful not to overshoot, because there is a legitimate residual use and the distinction is precise.
Dynamic KBA derived from your own transaction data, in an attended channel, as one weak signal among several, is fine. "Which of these four amounts did you transfer last Tuesday?", asked by an agent on a live call with the answer set generated fresh from data your system produced, is a different animal from "what was your first pet's name" on a web form. The answer set is generated per attempt, so there's nothing stored to breach. It's derived from data you hold, not from a bureau that will eventually be breached — the IRS/Equifax lesson applies specifically to out-of-wallet questions sourced from third-party aggregators. It expires. It isn't a fact about the person's life, so it isn't unrotatable. And there's a human on the line, an attempt budget of one, and the rest of the control stack from the diagram above.
Even then, be honest about the weight: a dynamic KBA challenge contributes maybe a couple of bits of confidence. Treat it as an input to a score, in the sense of Adaptive MFA, never as a gate that opens on its own. And the corollary most call centres get wrong: if the agent can retry, or hint, or accept "close enough," the budget is gone and so is the signal.
Two things this does not license. Not static questions in an attended channel — an agent reading out "what's your mother's maiden name?" is exactly the flow the well-documented help-desk social engineering campaigns of recent years walked straight through. And not dynamic KBA in a self-service web flow, because self-service removes the human, the attempt budget, and the observability at once.
Where the standards landed is worth knowing, because it settles arguments with auditors. NIST SP 800-63B is explicit that verifiers shall not prompt subscribers to use knowledge-based authentication or security questions when choosing memorised secrets — a normative prohibition, not a preference. On the proofing side, SP 800-63A revision 3 permitted knowledge-based verification only under tight constraints (not free-form, sourced from data not in public records or marketing data, limited attempts, time limits, diversified sources), and revision 4 tightened it further. Read the current text rather than my paraphrase of the modal verbs; the direction across revisions has been consistently one way.
What replaces it
Briefly, because this ground is covered properly in the two recovery articles and I'd rather cross-link than re-derive.
| Mechanism | Where it fails |
|---|---|
| A second registered authenticator — the only one that removes the weak branch | Users won't enrol at signup. Prompt at the moment of proven multi-device use |
| Recovery codes | Lost, or stored behind the account you've lost. Fine for technical audiences, weak for consumer |
| Email / SMS to a verified channel | Your security becomes the provider's. State that rather than claiming otherwise |
| Delegated / social recovery | Coercion and relationship change; a trust graph to maintain |
| Identity document + liveness | Only works for accounts proofed to begin with; a few dollars per attempt; false rejects land hardest on those who most need help |
| Manual review with delay and out-of-band notification | Support headcount and product courage — but the universal fallback and the strongest cheap control you have |
Two points matter more than the table, and both are argued fully in the passkey recovery piece. Delay plus un-suppressible notification is nearly free and converts a silent instantaneous takeover into a slow, loud one — the legitimate user waits 48 hours, the attacker spends 48 hours being announced to every channel the real owner controls. And a recovered session should grant less authority than a login: enrol a credential and read the account, but no changing the recovery address, no export, no moving money, for a cooling-off window.
The most useful reframing, though, is the min(). Your account security is min(authentication, recovery). A security question in the recovery path caps the whole account at four to seven bits regardless of what you built at the front door — so every hour spent on the login ceremony while that flow exists is spent on the wrong side of a minimum.
Getting them out of production without locking anyone out
"Delete them" is correct and insufficient, because the woman with the eleven-year-old account has no other recovery method on file and you are about to remove the only one she has. This is the actual engineering problem, and it looks a lot like a credential migration — which is convenient, because we have a pattern for those.
Phase 0 — measure, for two weeks, before changing anything. Four numbers, and almost nobody has them: how many accounts have answers stored and no other verified recovery method (the size of your real problem, usually far smaller than the raw count); recovery attempts per month reaching the question step; the pass rate there and the attempt-count distribution behind it (a large share of successes landing on attempt three or later means either struggling users or an attacker enumerating, and you can't yet tell which — itself a finding); and the ratio of successful question-based recoveries to accounts later reported compromised. That last one ends the internal debate, and nobody computes it.
Then log every question attempt as a security event with account, source and outcome. You're about to argue for removing a control; do it with your own numbers rather than mine.
Phase 1 — stop the bleeding at registration. Remove the questions from signup and from account settings today. One-line change, zero effect on existing users, caps the problem permanently. Do it before you've finished designing the rest of the plan.
Phase 2 — demote from sole factor to one signal. Highest-value step, and cheap. Stop accepting a correct answer as sufficient: require it alongside something else — an email link, a known device, a familiar network — or make it one input to a risk score.
The moment you do, the standalone takeover primitive is gone. The attacker who researched the answer now needs the mailbox too, and the user who can't remember gets a second path rather than a wall. Note the asymmetry: this removes almost all of the attack value and almost none of the user value. If you ship one thing from this article, ship this.
At the same time, cut the attempt budget hard — three attempts, lifetime, long cooldown, notification on every failure. Don't implement it as an account lockout anyone with a username can trigger; that's a denial-of-service vector and these questions aren't worth acquiring one for. Throttle and notify; don't lock. And resist tightening the matching: it's the intuitive move and it's backwards, because tighter normalisation raises false rejects on legitimate users while the attacker, who researched the true answer, is unaffected. The budget is the lever, not the tolerance.
Phase 3 — drive re-enrolment at next login. An interstitial after successful authentication, never before — in front of a user whose identity you've just verified, not in the way of logging in. One screen, cheapest real option first, skippable twice, completion measured. Then take the lesson from Why SMS OTP Won't Die seriously, because it's this phase's trap: a hard, undismissable prompt drives some users to abandon entirely, and aggregate coverage can fall while average credential strength rises. Decide in advance which number you're optimising and plot both curves from week one.
Phase 4 — announce a sunset date and hold it. Twice, at least thirty days apart, to every channel on file. Then stop accepting them — for everyone, including support agents and the internal admin tool, which is where these survive longest and where nobody looks.
Phase 5 — the dormant tail, which has no clean answer. Your active population converts; a stubborn share never logs in again, never sees the interstitial, and holds no other method. This is structurally the dormant-account problem from Migrating Password Hashes Without Resetting Users, with the same three defensible positions: keep a manual, high-friction path (identity proofing or human review with a mandatory delay) for accounts that come back, which costs you a support process to staff and audit; force enrolment hard on next login, which works where you can absorb the friction and will read to users as an incident whether you say so or not; or expire the accounts on a published schedule, which is legitimately right for dormant free-tier accounts with no financial or social value and is a product decision needing a named owner outside engineering.
The wrong fourth option is the one most organisations ship: leave the questions accepted "just for the old accounts," let the sunset slip, and let the arrangement become permanent — which means the flow you were removing is still running, on exactly the population least able to notice it being abused.
Phase 6 — delete the data, properly. DROP COLUMN is the easy part. The answers are also in the support tool, in the CRM where an agent pasted one into a case note, in the warehouse that replicates the users table wholesale, in the request bodies your reset endpoint logged before someone added a redaction rule, and in backups. Backups usually age out on the normal retention cycle — but know the cycle and write down the date the last copy expires, because "we deleted them" and "no copy exists" are different claims and someone will eventually ask for the second.
How to store them until you do
Which brings us to the finding I promised, and the answer to a question that has bothered me for years: why are security question answers stored in plaintext so much more often than passwords?
It isn't carelessness. It's a genuine three-way constraint. You want all of: hashed at rest, with a per-user salt and a memory-hard KDF, like any other shared secret; fuzzy-tolerant matching, because exact match on human-fact answers has an unacceptable false-reject rate; and agent-readable, because the support process wants to read the stored answer and judge whether the caller's response is close enough.
You cannot have all three, for the same reason biometric templates can't be hashed. A cryptographic hash annihilates variation by design — one input bit flips half the digest — which is exactly what makes it safe to store and exactly what makes it useless for a comparison that must succeed on inputs that differ. Fuzzy matching fights hashing directly. And agent-readability doesn't merely fight it, it forbids it: if a human can read the answer, it is stored recoverably.
Faced with that, teams took the two requirements with visible stakeholders — support wanted readability, the dashboard wanted tolerance — and quietly dropped the one whose absence nobody would notice until the breach. Every plaintext answers table in the world is that decision, made once, usually by someone who never framed it as a decision.
The resolvable version drops agent-readability entirely and buys most of the tolerance with normalisation instead:
stored = argon2id( normalize(answer), salt_per_user )
normalize(s) = strip accents
→ casefold
→ collapse internal whitespace
→ trim
→ strip punctuation
Normalise identically at set time and verify time and pin that function exactly — write it down, comment it, test it, because in four years someone reimplements it in another language and a one-character difference locks out every account with no diagnostic. You lose edit-distance tolerance and partial credit, which is fine, since those were helping the attacker more than the user. You lose agent read-out, which is the point: an agent who can read the answer is a social-engineering target holding a permanent, cross-site, unrotatable fact about your customer.
Be clear-eyed that normalisation is itself the class-merging from section 5. You're choosing a specific, documented, minimal amount of widening instead of an undocumented, ever-growing amount added ticket by ticket. That's the whole improvement, and it's real.
If your answers are in plaintext today and you're a year from sunset, do this migration anyway. It's an afternoon's work, rehash-on-next-use if you want to be gentle, and it means that if the table leaks in month eight you're explaining a hashed column instead of the Yahoo sentence.
The line worth keeping
Most of Security Advice We No Longer Believe has the same shape: a control was a sound response to the constraints of its era, the constraints changed, and the control kept running because nothing about a control announces that its premise has expired.
S9 breaks the pattern, and I think it's worth being honest that it does. Security questions were never sound. The distribution of human-fact answers was as concentrated in 1998 as it is now; the answers were as unrotatable; the fuzzy matcher was as required. What made them shippable was that the alternatives genuinely didn't exist, that the control was free at the point of use, and — decisively — that it emitted no evidence about its own performance, so the inversion at its heart went unmeasured for a decade after it could have been measured.
The general lesson isn't "don't use security questions," which you already knew. It's two things.
A control extracted from its enforcement environment is a new control, and nobody has evaluated it. KBA worked in a call centre because of the agent, the callback, the transaction data, the fraud team, and the reversibility. The web took the question and left the control behind, and the resulting artefact inherited the reputation of the original without any of its properties. When you see a mechanism migrating between contexts — a matcher moving out of hardware, a fraud check moving from attended to self-service, a signal moving from one-of-six to sole-sufficient — that's the moment to re-derive its security from scratch rather than from its provenance.
And any control whose false-reject rate you have never compared against its false-accept rate is a control you have not evaluated at all. That's the sentence I'd carry out of this. The reason security questions survived twenty-five years is not that anyone believed they were strong. It's that the two failure modes landed on different dashboards, owned by different teams: the false rejects went to support as ticket volume, and the false accepts went to fraud as chargebacks, and no one ever put the two numbers on the same axes.
Put them on the same axes. If the attacker's pass rate is anywhere near the owner's, you don't have a weak control. You have one pointing the wrong way.