Why SMS OTP Won't Die

A payments company I talked to spent most of a year building a proper authenticator-app enrolment flow. TOTP, recovery codes, a nice QR screen, copy reviewed by a writer. The security team had wanted SMS gone since the last pentest, and now there was a real alternative to move people to.

They ran it as a staged migration. Existing SMS users got a banner, then an interstitial, then a hard prompt at login: set up your authenticator app. Two dismissals allowed, then the screen stopped being dismissible.

Eleven weeks in, someone finally plotted the two curves on the same axes. Authenticator adoption among prompted users: a bit over a third. Users who hit the hard prompt and then simply stopped logging in: enough that support noticed before the dashboard did. And the third curve, the one nobody had drawn in advance — users who dropped MFA entirely, because the path of least resistance out of an undismissable prompt was to turn off two-factor and go back to a password.

Net effect on the population: the share of accounts protected by any second factor went down. The share protected by a strong second factor went up. Which of those two numbers is the security outcome depends entirely on which threat you were worried about, and nobody in that org had written it down.

That is the whole SMS argument in one paragraph, and it is not the argument the industry keeps having.

The lecture, and why it keeps failing

The standard take is: SMS is insecure, SIM swapping is trivial, use an authenticator app. Every part of that sentence is defensible in isolation and the conclusion still doesn't survive contact with a real user base.

It fails because it compares SMS to TOTP, and the actual comparison a product leader faces is SMS versus nothing. Not because stronger factors don't exist, but because stronger factors have enrolment friction, and enrolment friction is not a rounding error. It is the dominant term.

A factor's contribution to your security posture is its strength multiplied by its adoption. Everyone can recite this. Almost nobody builds a rollout plan that respects it.

Here's the shape of the arithmetic. Take a consumer population of a million accounts and be honest about the enrolment rates you'd actually see, not the ones in the vendor deck:

Policy Enrols Strength vs. credential stuffing Strength vs. real-time phishing Accounts left on password only
SMS offered, low friction ~60–70% High Low ~300k
TOTP only, hard prompt ~30–40% High Low ~600k+
Passkey only, hard prompt Varies wildly by platform mix High High Large, and skewed toward older devices
SMS default, passkey opt-in ~70%+ on something High High for the opt-in cohort Small

The last row is the one that gets called "insecure by default" in a design review and is frequently the best aggregate outcome available. The population that never enrols in the strong factor is not a rounding error you can shame into compliance. It is a structural feature of consumer scale.

The public data supports the pessimistic read. Twitter's own 2021 transparency reporting put the share of accounts with any 2FA enabled at roughly 2.6%, and about three quarters of those were on SMS. That is a technically sophisticated company with a highly targeted user base. If your mental model is that users will migrate to a better factor when you explain the risk, that number is the correction.

What NIST actually said, and what it says now

This deserves accuracy because it gets cited badly in both directions.

In the 2016 public preview draft of SP 800-63B, NIST wrote that out-of-band authentication using SMS was deprecated and "may no longer be allowed in future releases." That sentence launched a thousand slide decks. It was also a draft.

The final publication (SP 800-63B, June 2017) removed the deprecation language. What replaced it was more interesting and much less quotable: SMS became a RESTRICTED authenticator. Restricted does not mean prohibited. It means permitted, at AAL2, with obligations attached — the verifier has to assess the risk of number porting and SIM change, has to offer at least one alternative authenticator that isn't restricted, has to give users meaningful notice of the risk, and has to maintain a migration plan for when the restricted method is eventually disallowed. Separately, and firmly: VoIP numbers and email are not acceptable out-of-band channels.

The restricted classification carried forward into Revision 4. SMS is still usable, still restricted, still carries the "offer something better and have a plan" obligations.

So the honest reading of the standard is neither "NIST banned SMS" nor "NIST blessed SMS." It is: SMS is a legitimate AAL2 factor whose known weaknesses you must acknowledge, disclose, and have an exit strategy for. That's a governance requirement dressed as a crypto requirement, and it's a considerably more useful instruction than the lecture.

Two threat models wearing the same coat

The reason the SMS debate goes in circles is that "SMS is insecure" bundles two attacks with almost opposite economics, and they point at different mitigations.

SIM swap is targeted and expensive per victim. It requires reconnaissance on a specific person, a carrier interaction — social engineering a support rep, a bribed insider, or a fraudulent port-out request — and a window of minutes to hours before the victim notices their phone has no signal. Real, damaging, well-documented, and it does not scale. Nobody SIM-swaps a hundred thousand accounts. It is deployed against people whose accounts are worth four or five figures minimum: crypto holders, executives, domain registrars, people with valuable handles.

OTP-relay phishing is cheap and scales infinitely. A phishing kit stands up a reverse proxy between the user and your real login page. The user types their password, the proxy replays it, your server sends a code, the proxy shows the user a code prompt, the user types the code, the proxy replays that too, and the attacker captures the resulting session cookie. Off-the-shelf kits do this. The marginal cost of the ten-thousandth victim is a domain name.

Here is the part that experienced engineers sometimes skip past:

sequenceDiagram
    participant U as User
    participant P as Attacker proxy
    participant S as Your login service
    U->>P: Password (on lookalike domain)
    P->>S: Password (relayed)
    S->>U: OTP delivered out-of-band
    Note over U,S: Channel is irrelevant here —<br/>SMS, TOTP app, or email
    U->>P: OTP code, typed by the user
    P->>S: OTP code (relayed, seconds old)
    S->>P: Session cookie
    Note over P: Attacker now holds the session.<br/>The factor did its job perfectly.

The relay attack does not care whether the code came from a carrier or from an authenticator app. TOTP is equally defeated. So is email OTP, so is a push approval without number matching, so is a hardware token that displays digits. Any factor whose output the user can read and retype is relayable, because relay exploits the user as the transport, not the channel.

Which means the popular remediation — "migrate from SMS to an authenticator app" — is a fix for the expensive, targeted, low-volume attack, and does approximately nothing about the cheap, scalable, high-volume one. That's usually backwards relative to actual loss.

The two attacks want different countermeasures:

Attack Real mitigation Migrating SMS → TOTP helps?
SIM swap Carrier port-out locks, SIM-change signals where the operator exposes them, re-authentication delays on high-value actions, device binding Yes, substantially
OTP relay Origin-bound credentials (WebAuthn/passkeys), session-binding, token-shape detection, blocking lookalike domains fast No
Bulk credential stuffing Any second factor at all Both work; adoption is the only variable that matters

If your loss data is dominated by credential stuffing — and for most consumer products it is — then the highest-value move is raising coverage, and SMS is the cheapest coverage per user that exists. If your loss data is dominated by relay phishing, the move is phishing-resistant credentials, and TOTP is a detour that costs you enrolment without buying resistance.

The reach argument, stated plainly

SMS has one property no competing factor has: it requires nothing of the user that they don't already have. No app install. No app-store account. No enrolment step to abandon halfway. No device that must be recent enough to hold a platform credential. It works on a five-year-old Android with 200MB free, on a feature phone, on a shared family handset, on a device where the user has forgotten their app-store password and cannot install anything at all.

That last population is larger than most architects believe, and it is not evenly distributed. It is concentrated among older users, lower-income users, users in markets where handset refresh cycles run seven years, and users who did not opt into the smartphone-with-a-password-manager world your team lives in. Removing SMS is a decision with a demographic incidence, and it will not show up in an aggregate conversion number — it shows up as a support queue.

There's a second-order effect worth naming. Every enrolment step is a place users drop out, and drop-out from an authentication enrolment doesn't leave the user unprotected in a neutral way. It leaves them and leaves your support desk holding a recovery flow, which is the softest surface you own.

The bill nobody budgets for

Two operational costs that don't appear in the security debate and dominate the actual TCO.

SMS pumping (artificially inflated traffic). An attacker controls, or is paid by, a party with revenue share on terminating traffic to certain number ranges — often in a handful of specific countries. They point a bot at your unauthenticated "send me a code" endpoint and pump millions of messages to those ranges. You pay per message. They collect the termination fee. Nobody gets breached; you just get invoiced. X/Twitter publicly claimed around $60 million a year in fraudulent 2FA SMS traffic when it moved SMS 2FA behind a paywall in 2023. Most teams discover this the way you discover a memory leak: on a bill.

The defences are mundane and must be built before you need them — per-number and per-IP rate limits on the send endpoint, geo-allowlisting the country codes you actually serve, a CAPTCHA or proof-of-work on send rather than only on verify, blocking known pumping ranges, and alerting on the ratio of codes sent to codes verified. That ratio is the single best detector: legitimate traffic converts; pumped traffic never verifies.

Delivery reliability and international variance. SMS is not a reliable channel. It is a best-effort one with no delivery guarantee, wildly varying latency, and per-country behaviour that will surprise you: sender-ID registration regimes, content filtering that silently drops messages containing URLs or the word "code," aggregator routes that degrade at 3 a.m. local time, carriers that deprioritise A2P traffic. A 97% delivery rate sounds fine until you notice it means 3% of your users cannot log in and the failure is invisible to them and to you. Instrument send-to-verify conversion per country and per route, or you are flying blind on an availability-critical path.

And note the coupling: SMS is a factor whose availability is outside your control, sitting on your login path. That's a different risk category than "it's phishable," and it's the one that pages you.

Where SMS actually does the most damage

Not as a login factor. As a recovery path.

If a user can recover their account with an SMS code, then every stronger factor sitting above it is decorative. Your passkey deployment is worth exactly what your recovery flow is worth, because an attacker will not attack the passkey. This is the min() problem, and it's argued properly in The Hardest Problem in Identity Isn't Authentication. It's Recovery. and Passkey Account Recovery Is the Whole Problem, so I won't relitigate it here.

The point specific to SMS is the silence of the downgrade. Teams deliberate hard over which factors to offer, then wire SMS into recovery as an implementation detail nobody reviewed, and ship a system whose real assurance level is set by a code sent to a phone number a carrier support rep can reassign. If you take one operational action from this article: go read your recovery flow and find out what assurance level it actually grants. It is almost certainly lower than the one on your login page, and it is the one that matters.

The decision rule

Stop asking "is SMS secure." Ask what is the marginal alternative for this specific population, and what is the loss if this specific account is taken over.

Situation Call
Consumer, low-value account, no second factor today Offer SMS. SMS beats nothing decisively, and nothing is the real alternative.
Consumer, mixed population, moderate value SMS as the default, passkey prominently offered, no forced migration. Let the strong factor win on merit; keep coverage high.
Consumer, high-value actions (payments, transfers, address change) SMS for login is fine; step up to a phishing-resistant factor for the action. See Step-Up Authentication: The Right Way to Handle Sensitive Operations — the factor should match the operation, not the session.
Any account holding crypto, domains, or direct money movement Remove SMS entirely, including from recovery. SIM swap is economically rational against these accounts and will be attempted.
Administrative and privileged access, any product Remove SMS entirely. The population is small, enrolment friction is affordable, you can mandate hardware. There is no reach argument for fifty admins.
Enterprise workforce, managed devices Remove SMS. The employer owns the device and can mandate the factor; the consumer constraints simply don't apply.
Recovery path, any tier Never SMS alone. SMS plus a delay plus notification plus a second signal, or a different mechanism entirely.

Two structural rules that make the rest of this maintainable.

First, authorize on assurance level, not on method. If services check "was this mfa" rather than "was this otp," then retiring SMS for a tier is a policy change in one place instead of a redeploy of everything. That's the acr versus amr distinction, and it's the difference between a migration you can execute and one you keep postponing.

Second, let risk decide when SMS is sufficient rather than deciding globally. A known device on a known network doing something ordinary needs less evidence than a new device in a new country moving money. SMS is adequate evidence for a lot of the first category and inadequate for most of the second — which is Adaptive MFA: Security Should Follow Risk, Not Rules applied to exactly one factor.

The uncomfortable conclusion

SMS OTP won't die because it is the only authentication factor that meets users where they already are, and reach beats strength in the aggregate more often than security architects like to admit.

That is not a defence of SMS. SMS is a bad factor with real, exploited weaknesses, an unbounded fraud bill, and no delivery guarantee. It should be removed from privileged access, from high-value accounts, and from every recovery path today, and none of those removals require a single consumer to install anything.

What it is, is an argument against the reflex. The move that improves your posture is almost never "ban SMS." It is: draw the line by population and by transaction value, make the strong factor genuinely easier than the weak one instead of merely mandatory, fix the recovery path first, and instrument the send-to-verify ratio before the invoice teaches you what it means.

Then write down which of the two attacks you are actually defending against. Most of the SMS debate is two people optimising against different threat models and neither one saying so.