MFA Fatigue: When More Security Makes Users Less Safe
The phone starts buzzing at 01:47.
It is the same notification the contractor has approved four hundred times: Are you trying to sign in? He denies it and goes back to sleep. Again at 01:49. Again at 01:52. Then a cluster — six in ten minutes, then eleven more. Around 02:30 a WhatsApp message arrives from someone claiming to be corporate IT: the alerts are a known glitch, approve one and they'll stop.
He approves one. They stop.
That is, in its essentials, how Uber was breached in September 2022. Uber's own incident update says the attacker had bought a contractor's corporate password — exposed by malware on a personal device — and then repeatedly triggered two-factor approval requests until one was accepted. The WhatsApp detail comes from the attacker's own account of it, given to researchers and journalists at the time.
Every control in that chain worked as specified. The password was correct. The second factor reached a registered device the real user was holding. The user affirmatively approved, and the system logged a clean multi-factor login. Nothing malfunctioned.
The usual reading of this is that the human failed and needs training. That reading is wrong, and it is wrong in a way that matters architecturally. The user was asked a question they had no means of answering correctly. A push approval prompt is an authorization decision presented to a party who cannot see any of the decision inputs — and no amount of awareness training fixes an under-specified interface.
The prompt is an under-specified authorization decision
Strip the product design away and look at what a push factor actually is.
Somewhere in the authentication service there is a pending-approval record, created when a session presented a valid password: an account identifier, a timestamp, an expiry, a state field. It is keyed, functionally, on the account — not on the session that caused it. When an approval arrives from the enrolled device, the service matches it to a pending record for that account and releases the session.
That last sentence is the whole vulnerability. The approval does not identify which session it authorizes; it identifies which account is approving. If there is exactly one pending request, the mapping is unambiguous by accident. The security property was never "this user approved this login." It was "this user approved a login, and we assume it's the one they were thinking of."
The out-of-band channel that makes push resistant to classical credential phishing is the same thing that makes it blind. The phone is deliberately not the browser: no view of the requesting origin, no shared cryptographic state with the requesting session, no way to tell the user's own laptop from a VPS in another country. It presents a claim it cannot verify to a person who cannot check it, and asks for a boolean.
Compare this to what a WebAuthn assertion carries — the challenge issued to that specific session, signed over the origin the browser reports, by a key scoped to the relying party ID. There, the answer is inseparable from the question. In push, the answer is a free-floating "yes" that any pending question can consume.
Once you see it in that shape, the attack writes itself. It isn't a clever exploit. It's a race that the attacker gets to run as many times as they like:
sequenceDiagram
participant A as Attacker (stolen password)
participant AS as Auth service
participant P as Victim's phone
participant V as Victim
loop until approved or attacker gives up
A->>AS: POST /login (valid password)
AS->>P: Push: "Approve sign-in?"
Note over P,V: Prompt carries no reference<br/>to the requesting session
V-->>AS: Deny (or ignore)
end
A->>V: Out-of-band: "IT here, known glitch,<br/>approve to stop the alerts"
A->>AS: POST /login
AS->>P: Push: "Approve sign-in?"
V->>AS: Approve
AS-->>A: Session established
Note over AS: Audit log: successful MFA login
The two mechanisms in that diagram do different jobs and are worth separating. The loop is push bombing — a volume attack on attention, which works on its own often enough. The message is the pretext, which converts a user who is denying correctly into one who approves deliberately. Uber was both. So was Cisco.
What the incidents actually show
Three cases are documented well enough to reason from, and their differences are more instructive than their similarities.
Uber, September 2022. Password from an infostealer log, an hour of repeated prompts, then a social pretext over WhatsApp. What followed matters as much: the attacker found a PowerShell script on an internal share with hard-coded admin credentials for the privileged access management system. The push approval was worth so much because everything behind it assumed the perimeter had held.
Cisco, May 2022. Per Cisco Talos's own writeup, the attacker compromised an employee's personal Google account, which was syncing Cisco credentials saved in Chrome, then ran voice phishing calls impersonating trusted support organizations alongside MFA push fatigue until an approval landed. Talos attributed the access to an initial access broker with ties to UNC2447, Lapsus$, and Yanluowang. Note the shape: the credential never came from a phishing page at all.
Lapsus$ / DEV-0537, 2021–2022. Microsoft's threat intelligence writeup describes repeated MFA prompting as routine for the group, alongside session token replay and calling help desks to have credentials reset. Their own Telegram channel, widely quoted at the time, put it more plainly than any vendor advisory: call the employee at 1am, repeatedly, and eventually they accept.
What all three demonstrate is not that users are careless. It's that the attacker only needs the mapping between "an approval" and "the right approval" to be ambiguous once, and the defender needs it to be unambiguous every time.
Number matching, and precisely what it buys
The standard fix is number matching: the requesting screen displays a number, and the authenticator asks the user to type it rather than tap a button. Microsoft made it the default for Authenticator push in May 2023; Duo's Verified Push and Okta's number challenge are variants of the same idea. It works and it is worth deploying — but the common implementations are not equally strong.
Number matching establishes a weak out-of-band channel from the requesting session to the approving device, routed through the user's eyes. The digits are a short one-time value that only someone looking at the requesting screen can produce. That is a real improvement on a boolean: the attacker on a VPS in another country cannot supply it.
But the strength depends entirely on whether the user transcribes or selects.
| Design | What the user does | Blind-approval odds | Survives flooding? |
|---|---|---|---|
| Tap-to-approve | Taps a button | 1 in 1 | No |
| Pick one of three shown numbers | Selects | ~1 in 3 | Barely — ~3 prompts on average |
| Type a 2-digit code | Transcribes | 1 in 100 | Yes, in practice |
| Type a 6-digit code | Transcribes | 1 in 10⁶ | Yes |
The middle row is the one people miss. A selection interface preserves a blind-tap path: a fatigued user who is guessing gets through in about three prompts, and the attacker has an unlimited supply of prompts. A transcription interface removes the path entirely — there is no gesture a distracted user can perform that produces the right digits by accident. The difference is not "more entropy." It is that one design still has a fast reflexive action available and the other does not.
Now the limit, which is the part that matters for threat modelling: number matching defeats blind approval, not a social-engineered relay.
The digits are not a secret. They are a value the attacker can read off their own screen. If the attacker has a live voice or chat channel to the victim — the Cisco pattern, the Scattered Spider help-desk pattern, every vishing campaign since — they simply read the number aloud. "This is IT, we're re-enrolling your device, you'll see a prompt, enter 47." The victim types 47. The channel binding is satisfied, because the human is the channel, and the human has been recruited.
So number matching moves the attacker's cost from "send 40 pushes" to "hold a two-minute phone call." That is a large increase — it eliminates the unattended attack, forces contact that carries attribution, and cuts viable targets per campaign by orders of magnitude. It does not change the category. The approval is still an assertion about a session the approver cannot see.
Rate limiting, and the denial you're throwing away
Capping prompts is the other standard mitigation: N pending approvals per account per window, then stop issuing them.
The obvious objection is the one made at length in Account Lockout Is a Denial-of-Service Vector: any threshold that disables an account on unauthenticated input is a denial-of-service primitive handed to anyone who knows a username. That argument holds here and I won't repeat it.
But push throttling has an asymmetry that raw login lockout doesn't, which changes the trade:
- The flood is triggered after a correct password. It is not anonymous input. An attacker who can drive your push rate limiter already holds a valid credential, which means the DoS is the least of your problems in that account.
- The throttle can degrade rather than lock. Stop issuing pushes and fall through to a factor that is not vulnerable to flooding — a TOTP code the user reads, or better, a security key. Availability is preserved; only the flood-able path closes.
Which points at the thing most deployments get wrong. A denied push is one of the highest-fidelity security signals you will ever collect, and almost every system treats it as a no-op.
A deny means someone presented the correct password for this account, and the account holder — the one person who knows whether they were logging in — said no. That is a confirmed credential-compromise report from the most authoritative source available, at the moment it happened. Most implementations increment a counter.
At minimum, a deny should invalidate the password that triggered it, kill existing sessions, and raise an event a human sees. Three denies inside a few minutes is not indecision; it is an attacker with a working password, iterating. The correct response is to make that password stop working, not to keep offering retries at a slower rate.
Rich context doesn't fix it, and the reason is arithmetic
The next thing everyone tries is putting more information in the prompt: application name, IP, city, map pin, browser, time. It feels like the right move — give the user the facts and let them judge.
It doesn't work, and the reason isn't that users are lazy.
Consider the base rate. Where a user gets ten prompts a week and a genuine attack arrives every three years, roughly one prompt in 1,500 is malicious. A user who approves everything without reading is correct 99.93% of the time. Reading carefully costs real time on essentially every prompt, for an expected benefit measured in events per decade. Reflexive approval is not user error. It is the rational strategy given the prior, and the interface is what set that prior.
And the context is largely uncheckable. A user in Munich sees "sign-in from Frankfurt." Is that wrong? Their VPN egresses somewhere, their carrier routes through a distant gateway, the corporate proxy moved twice this year. The one field that would be decisive — is this the browser tab I am currently looking at — is precisely the one the push channel structurally cannot carry.
This is the alarm-fatigue result operations teams already know from noisy pagers: a detector with a terrible prior and no way to verify its inputs converges on always answering the same way. Nobody fixes pager fatigue by telling on-call to read more carefully. You fix the signal.
Context in the prompt is still worth having — it is evidence at review time, and it occasionally catches someone. It is not a control.
The structural fix
A WebAuthn assertion cannot be approved for someone else's session, and this isn't a policy or a UX choice — there is no operation in the protocol that does it.
The mechanics are covered in Under the Hood of WebAuthn and The Magic Behind Passkeys; two properties kill this attack class:
The challenge belongs to the requesting session. The relying party generates a random challenge for this login attempt. The authenticator signs over it. An assertion produced for one session is not a valid answer to another session's challenge — there is no shared "the user said yes" state for a second session to consume.
The origin is inside the signed data. The browser, not the user, supplies the origin in clientDataJSON, and the credential is scoped to an RP ID. An attacker cannot get the victim's authenticator to produce a signature for the attacker's site, and cannot get a signature made for the real site to be usable from anywhere but the browser that requested it.
The consequence for MFA fatigue is total: there is nothing to flood. An attacker with a stolen password can hammer the login endpoint forever and the victim's device will never present a prompt they could mistakenly satisfy, because the only thing it can produce is a signature bound to the origin and challenge of the session in front of the user.
One honest caveat, because it is the same bug wearing a different hat: passkeys eliminate the flooding class, not every "approve something for someone else" primitive. The OAuth device grant has the same structural gap — the user authenticates perfectly, to the real authorization server, possibly with a passkey, and then approves a session that started on the attacker's machine. That's Approving a Login on Your Phone for a Smart TV Is an OAuth Flow, and it's why "we deployed passkeys" doesn't retire the whole family. Any flow where a human approves a session they cannot see inherits the same weakness.
The fallback path is your actual security posture
Here is where most rollouts quietly fail.
An organization deploys passkeys. Enrolment reaches 80%. The dashboard is green. Push stays enabled — for the 20% not yet enrolled, the shared frontline terminals with no per-user authenticator, the legacy VPN concentrator that speaks RADIUS and nothing else, the call-centre assisted channel, break-glass, and the executive whose security key is in a hotel safe two timezones away.
The security of a multi-factor deployment is the strength of its weakest available factor, not its strongest enrolled one. The attacker does not use the factor you prefer. They present the password, request the fallback, and the flood is back. Every step of the rollout was real; none of it closed the hole, because the hole is the alternative.
Two metrics distinguish a real rollout from a cosmetic one, and organizations usually publish the first and not the second:
- % of users enrolled in a phishing-resistant factor — the vanity number.
- % of successful authentications that actually used one — the real number. And its complement: for each user, can a push still satisfy this account today?
The remediation is unglamorous and per-user, not global. Once a user has a working phishing-resistant credential on at least two devices — the second device is what makes this survivable — remove push from that user's allowed factor set. Not deprioritize. Remove. A policy that says "prefer passkey, fall back to push" is a policy that says "push," because the attacker chooses.
For populations that genuinely cannot get there yet, the fallback should be narrow rather than universal:
| Situation | Fallback that keeps the hole | Fallback that mostly closes it |
|---|---|---|
| Not yet enrolled | Push, indefinitely | Push, with a hard enrolment deadline and admin-visible aging |
| Shared terminal | Push to personal phone | Per-shift hardware keys held at the location |
| Legacy RADIUS/IMAP | Push | Front it with a proxy that terminates modern auth; block basic auth |
| Lost device | Self-service re-enrol via push | Verified human recovery — see The Hardest Problem in Identity Isn't Authentication. It's Recovery. |
| Admin/privileged action | Whatever logged them in | Phishing-resistant required at the operation, via step-up |
That last row has the best ratio of effort to risk removed. You may not eliminate push from every login this year, but you can ensure it never satisfies the operations that matter — the argument in Step-Up Authentication: The Right Way to Handle Sensitive Operations, carried by an ACR or AMR claim the resource server actually checks rather than a policy set at the front door.
Where this leaves you
If you are running push today, in order of risk removed per unit of work:
- Number matching with transcription, not selection. Cheap, immediate, ends the unattended attack.
- Treat a deny as a compromise report. Invalidate the password, kill sessions, alert. This is often a small change and it is the highest-value one on the list.
- Cap pending prompts, and degrade to a non-floodable factor rather than locking the account.
- Require a phishing-resistant factor for privileged operations, independent of how the session began.
- Retire push per user as they enrol, and measure authentications rather than enrolments.
And the framing worth keeping after the tactics change: MFA fatigue is not a story about human weakness under pressure. It is a design in which the party asked to make an authorization decision has been given none of the information the decision requires, over a channel deliberately built to be separate from the one that would carry it. The 2am approval is the predictable output of that design, not a deviation from it.
You can make the question harder to answer wrongly, and number matching does exactly that. But the class does not close until the answer is cryptographically bound to the question — which is a property of the protocol, not of the person holding the phone.