Adaptive MFA: Security Should Follow Risk, Not Rules

There is a specific failure that no security architecture diagram contains.

There is a specific failure that no security architecture diagram contains.

An engineer is at her desk, on the corporate network, on the laptop IT issued her, at 10:40 on a Tuesday. She logs into the internal admin console. Her phone buzzes. She approves it without looking up from her screen — she has done this eleven times today. The approval takes 400 milliseconds of her attention and zero of her judgment.

At 22:15 the same phone buzzes again. She is watching television. She approves it. It was not her.

Nothing malfunctioned. The MFA worked exactly as designed: it presented a challenge, the user approved it, access was granted. The system's model of the world is that a second factor was verified. What actually happened is that a human being was trained, by roughly two thousand meaningless prompts, to treat the prompt as noise — and then the one prompt that carried information arrived and was processed identically.

This is the argument for adaptive MFA, and it is not primarily an argument about convenience. Uniform friction destroys the signal value of the challenge. A prompt that appears when nothing is wrong is not merely annoying; it is actively corrosive to the security control it implements, because it teaches the only component capable of detecting the attack — the user — that the prompt means nothing.

Adaptive authentication is the proposition that a challenge should be rare and therefore meaningful. Getting there is harder than the vendor slides suggest, and the hard parts are statistical rather than technical.

The two decisions, which people conflate

Before the signals, an important separation. A risk engine produces a score. What you do with it is a policy decision, and there are two distinct axes:

How much evidence do I require? No challenge, a soft challenge (push, TOTP), a hard challenge (phishing-resistant factor, WebAuthn), or a human verification.

What am I authorizing? Full session, restricted session, read-only, or refusal.

Most implementations only use the first axis, which is a waste, because the second is often the better answer. A moderately suspicious login doesn't have to be a binary "challenge or allow." It can be allowed, with the ability to read but not to change the payment details, and with the session flagged so that any high-value action triggers step-up at that moment. That is frequently a better trade than a challenge at login: the legitimate user proceeds unimpeded through everything they usually do, and the attacker hits a wall precisely when they attempt the thing they came for.

Adaptive MFA is not a smarter gate at the front door. It's the recognition that the front door is the wrong place to make most of the decision.

The signals, and what each one is actually worth

Everyone publishes the same list. What's missing is the honest assessment of each signal's strength and spoofability, so here it is.

Device recognition — the strongest signal you have. A durable device identifier, ideally a cryptographic one: a key in the TPM or Secure Enclave, a client certificate, or the presence of a registered passkey. Not a browser fingerprint. The distinction is enormous: a cryptographic device credential is unforgeable, whereas a fingerprint (user agent, canvas hash, fonts, screen size) is a guess that an attacker can copy from a session-hijacking payload and that legitimate users invalidate every time they update their browser.

If you build one signal, build this one. A first-party cookie or local-storage token issued at first successful authentication, cryptographically bound where the platform allows, gets you the highest-value input at the lowest cost. It also has a clean interpretation: this exact browser has previously completed a successful authentication for this account, which is a much stronger statement than anything else on this list.

Network reputation — moderate, asymmetric. The useful form isn't "is this IP good," it's ASN class and history. Traffic from a hosting provider or a known VPN exit for an account whose 400 previous logins all came from residential and corporate networks is a genuine anomaly. Conversely, a residential IP proves almost nothing — residential proxy networks are cheap and widely used by exactly the attackers you care about. The signal is much better at raising suspicion than at lowering it, which is a property you should encode deliberately: let network reputation add risk, but don't let it subtract much.

Geolocation — weak, and the source of your worst false positives. IP geolocation is accurate to the country perhaps 95% of the time and to the city considerably less. It breaks for mobile carriers that route through a distant gateway, corporate VPNs that egress in another continent, and CGNAT. "Impossible travel" is the classic rule and it produces a steady stream of false positives from people who use a VPN, from users whose carrier changed its egress, and from anyone who legitimately flies. Use country-level changes as a mild signal; treat city-level distance calculations with suspicion.

Behavioural history — underrated and cheap. Not keystroke dynamics; just the account's own patterns. This account has logged in 340 times, always between 07:00 and 19:00 local, always from two devices, always from one country. That's a distribution, and a new login is either inside it or outside it. The reason this signal is undervalued is that it looks like it needs machine learning, and it doesn't — a per-account histogram of hour-of-day, a set of known devices, and a set of known countries covers most of the value.

Velocity and population signals — the ones that catch actual attacks. Individual-login risk scoring misses the most common real attack, which is credential stuffing: thousands of accounts, one credential each, from a rotating IP pool. Every individual attempt looks unremarkable. What's anomalous is the population: a spike in failure rate, an unusual ratio of unknown-device attempts, many distinct accounts touched from one ASN, a sudden shift in the mix of user agents. These are tenant-level and global-level signals, and a risk engine that only looks at one login at a time is structurally blind to them.

Account state — high value, frequently omitted. Was the password changed an hour ago? Was the email address changed yesterday? Is this account newly created? Does it hold elevated permissions? Has it appeared in a recent breach corpus? Risk is not just about this login; it's about what this account has recently been through. An account whose recovery email changed twelve hours ago is a different risk proposition regardless of how familiar the device is.

The requested action — often the most informative input. "Log in" and "change the payout bank account" deserve different evidence. If you evaluate risk only at authentication time, you are making the decision at the moment you know the least about intent.

How to combine them without a labelled dataset

The honest problem: you want to train a model, and you have no labels. You do not know which of last month's logins were malicious. You have a handful of confirmed account takeovers, maybe dozens, against tens of millions of logins. That's not a training set; that's an anecdote.

So don't start with a model. Start with the thing you can reason about, and let it earn you the data.

Stage one: deterministic rules on the strongest signal. Known device + successful recent authentication + no account-state anomalies → no challenge. Everything else → challenge. This single rule, in most deployments, removes 80–95% of challenges while raising the security level, because now the challenge only appears in circumstances that genuinely differ from the account's norm.

That asymmetry is the key design move: use your strongest signal to build a high-confidence allowlist, and challenge everything else. You are not trying to detect attacks. You are trying to recognize the overwhelmingly common case of a user on their own device, and treat everything outside it conservatively. Recognizing the familiar is a much easier problem than detecting the malicious, and it captures most of the benefit.

Stage two: additive scoring, in shadow mode. Assign weights to the remaining signals and sum them. Yes, this is crude; the weights are your judgment rather than a fitted parameter. It's also transparent, explainable to an auditor, debuggable at 3am, and — for a first implementation — usually within a few percent of what a model would achieve, because the signal set is small and one signal dominates.

Run it in shadow mode first: compute the score, log it, act on the old policy. After two weeks you have the distribution, and now you can pick thresholds from data. This step is not optional. Teams that pick thresholds before seeing the distribution invariably choose numbers that challenge either 2% or 60% of logins, and discover which on the day they ship.

Stage three: let the challenges label your data. This is the elegant part, and it's the reason the stages are in this order. Every challenge you issue produces an outcome: approved-and-then-normal-session, approved-and-then-suspicious-activity, abandoned, failed, explicitly-denied-by-user. Those outcomes are labels — weak ones, but real. Six months of them gives you a dataset with genuine signal about which score bands are associated with bad outcomes, and then a model is worth building.

An adaptive system generates its own training data, provided you log the outcomes and can join them to the score that produced the challenge. Most implementations record the decision and discard the outcome, which forecloses the entire path.

The arithmetic that determines whether this works

Here's the calculation to do before promising anyone a detection rate, because it governs everything and it is routinely skipped.

Suppose 1 in 20,000 login attempts is a genuine account takeover. Your detector is good: it catches 90% of attacks and has a 1% false positive rate.

Per million logins: 50 attacks, of which you catch 45. And 999,950 legitimate logins, of which you flag 1% — 9,999.

Of the 10,044 logins you flagged, 45 are attacks. Your precision is 0.45%. For every genuine attack you challenge, you challenged 222 innocent users.

This is the base rate fallacy, and it is the central fact of risk-based authentication. It has three consequences that shape the entire design:

Never block on a score. At 0.45% precision, blocking means locking out 222 legitimate users per attack prevented. This is why the action must be a challenge — a control the legitimate user passes and the attacker (usually) doesn't. The challenge is what converts a low-precision detector into a usable control, because it costs the legitimate user seconds rather than access.

Your real metric is challenge rate, not accuracy. Nobody can act on "90% detection." The number that governs both user experience and the reflexive-approval problem is what fraction of logins get challenged, and you should set it as a budget — say, under 5% — and then maximize detection within that budget. This reframes tuning from an accuracy problem into a constrained optimization, which is both more tractable and more honest.

Precision improves fastest by improving the allowlist, not the detector. Every legitimate login you can confidently exclude via device recognition is removed from the false-positive pool before the detector runs. Going from "no device signal" to "strong device signal" typically improves precision by more than an order of magnitude — far more than any amount of model tuning on the remaining traffic. Which is the mathematical restatement of stage one: recognize the familiar, don't chase the malicious.

The failure modes to design for

Missing signals must fail toward challenge, not toward trust. If your geolocation provider is down, your device-recognition cookie was cleared, or your risk service times out, the score is computed on incomplete data — and a scoring function that sums positive risk contributions will output a low score when inputs are missing. Low score, no challenge. Your risk engine's outage becomes a silent security downgrade with no alert, and it looks like a good day on the dashboard because challenge rates dropped. Track signal availability as a first-class metric, and treat "insufficient signals" as its own decision branch that challenges rather than as a zero.

Attackers adapt to what you measure. Publish that you trust corporate IP ranges and attackers will find a way into one. Trust a device cookie and session-stealing malware will copy it — which is exactly why a cryptographically bound device key is worth the extra work over a bearer cookie. Assume every signal you rely on will eventually be spoofed, and prefer signals whose spoofing requires capabilities an attacker probably doesn't have.

Consistency of the unchallenged path matters more than of the challenged one. If a user is sometimes challenged and sometimes not with no discernible pattern, they conclude the system is broken and they stop reading the prompts, which returns you to the reflexive-approval problem you were solving. Explain it. "We didn't recognize this device" is a sentence users understand and it makes the prompt feel purposeful. Give them a "remember this device" that visibly works.

Privacy and regulation constrain the signal set. Device fingerprinting and location tracking carry GDPR implications; a security-purpose legal basis is generally defensible, but it needs to be documented, the retention needs to be bounded, and in some jurisdictions and sectors behavioural biometrics have specific consent requirements. Additionally, an adaptive decision that materially affects a user may attract a right to explanation — which is a strong practical argument for keeping the engine interpretable. "Your score was 0.83" satisfies nobody; "this login came from a device and country we've never seen for your account" satisfies a user, a support agent, and a regulator.

Adaptive MFA is not a substitute for phishing-resistant factors. It reduces the number of challenges, which mitigates fatigue behaviourally. Only WebAuthn removes the class of attack structurally, by making the approval unforgeable and origin-bound. The two are complements and the ordering is: deploy phishing-resistant factors, then use adaptive policy to make invoking them rare enough to stay meaningful.

Where to start

Concretely, in order, for a team that has static MFA today:

  1. Issue a durable device credential on first successful authentication, cryptographically bound if the platform allows. This is the whole foundation.
  2. Suppress the challenge for known devices with clean account state. Ship this alone and you'll cut challenge volume by an order of magnitude.
  3. Log every decision with its inputs and the resulting outcome, joined by an ID. You are building the dataset that lets you do anything sophisticated later.
  4. Run additive scoring in shadow mode for two weeks before it makes any decision, then pick thresholds against a challenge-rate budget.
  5. Add population-level anomaly detection at the tenant level, separately from per-login scoring — it's the only thing that sees credential stuffing.
  6. Move the high-assurance requirement to the action, not the login, so a moderate-risk session is allowed but constrained.
  7. Alert on challenge rate going down, not just up. A drop is either a genuine improvement or a broken signal pipeline, and you want to know which.

The goal is not to challenge intelligently. It's for the challenge to be rare enough that when the engineer's phone buzzes at 22:15, she notices — because it hasn't buzzed unnecessarily in three weeks, and that makes the buzz itself information.

That's the entire value proposition. Not fewer prompts for their own sake: prompts that still mean something.