Step-Up Authentication: The Right Way to Handle Sensitive Operations
A bank I worked with had a rule: transfers over £10,000 required a second factor.
A bank I worked with had a rule: transfers over £10,000 required a second factor.
The implementation was a checkbox on the transfer page — I confirm this transfer — plus an SMS code, verified by the payments service, which then set a boolean in the session. session.stepped_up = true. The boolean had no expiry, because nobody had thought to give it one, and it applied to the session rather than to the transfer, because that's what a session boolean does.
So you authenticated once with a second factor, and every subsequent transfer that day passed the check. Worse, the flag was set by a service that had no idea which transfer the user had confirmed. The amount, the beneficiary, the reference — none of it was bound to the verification. The user confirmed something. The system recorded that a confirmation had occurred.
That gap between a confirmation occurred and this specific action was confirmed is where nearly all step-up implementations live, and it's the thing worth getting right before any of the protocol details matter.
Why login-time MFA isn't the answer
The reflex is to require MFA at login and consider the problem solved. It isn't, for two reasons that pull in opposite directions.
Time. A session that begins with MFA at 9am is, at 4pm, protected by whatever happened seven hours ago. The laptop has been to a coffee shop and a meeting room in between. Authentication is an event, and its evidentiary value decays; what you know at 4pm is that someone authenticated this morning and the session cookie is still present.
Proportion. Mandating strong authentication at login for every user, every time, in order to protect an operation that 3% of users perform monthly is a tax paid by everyone for the benefit of a few. It also generates the exact behaviour you don't want: users who tap "approve" reflexively, dozens of times a week, until the prompt carries no information and MFA fatigue attacks become trivially effective.
Step-up inverts the trade. Keep login proportionate to ordinary risk, and demand stronger evidence at the moment of consequence — when the evidence is fresh, the user understands what they're approving, and the prompt is rare enough to still mean something.
Step-up is not re-authentication, and the difference is not pedantic
These get used interchangeably and they're distinct operations.
Re-authentication asks: are you still there? Same factor, prove presence again. Password, or a passkey tap. It answers freshness — nobody sat down at your unlocked laptop.
Step-up asks: can you prove more than you did before? A stronger factor, or an additional one. It answers assurance — you have moved up a level, not merely repeated a level.
You need both, for different things, and conflating them produces the common bug where a "step-up" prompt accepts the same password the user typed at login. That's re-authentication with a security label on it. It confirms presence, adds no assurance, and is precisely useless against the attack you're worried about — an attacker who has the password and the session.
The corresponding protocol machinery is separate too. max_age handles freshness. acr_values handles assurance. Reaching for the wrong one is how you end up with a prompt that the attacker can satisfy.
The claims that carry this
OIDC gives you three pieces, and most teams use one of them.
auth_time — when the user actually authenticated, as a timestamp. Not when the token was issued; when the human proved something. It's the foundation, and it's the claim most implementations quietly drop.
acr — Authentication Context Class Reference. A string naming the assurance level reached. The spec deliberately doesn't define values, which is either flexibility or a gap depending on your mood. Some deployments use the ISO/NIST-style ladder, some use URNs, most invent their own. What matters is that they're ordered and few. I'd argue for three:
urn:acme:acr:pwd password only
urn:acme:acr:mfa password plus a second factor
urn:acme:acr:phishres phishing-resistant (passkey, hardware key, mTLS)
Three is enough to express real policy. Once you're at eight you have a taxonomy nobody can map to an actual decision, and services will start doing string equality against whichever value they saw in testing.
amr — Authentication Methods References. An array of what was actually used: ["pwd", "otp"], ["hwk"], ["swk", "user"]. Where acr is the policy conclusion, amr is the evidence.
The rule I'd hold to: services authorize on acr, and log amr. Authorizing on amr means every service embeds knowledge of which method combinations are acceptable, and when you deprecate SMS you get to redeploy all of them. Authorizing on acr means that decision lives in one place — the authorization server — and deprecating SMS is a policy change that removes otp from the set that yields mfa.
Then the request-side parameters. acr_values is a voluntary hint: the IdP may honour it, and if it doesn't, you get a token back at whatever level it felt like. To make it binding, use the claims parameter with essential: true:
{
"id_token": {
"acr": { "essential": true, "values": ["urn:acme:acr:phishres"] }
}
}
Now the IdP must either satisfy it or return an error. This distinction — hint versus requirement — is responsible for a whole category of implementations that appear to work and enforce nothing, because the happy path in testing had the user doing MFA anyway.
And max_age=0 combined with prompt=login is the freshness hammer: authenticate again now, regardless of session state. Note that prompt=login alone doesn't guarantee the factor, only that a login occurred — pair it with an essential acr when you need both properties.
The part almost nobody implements: the resource server can ask
Here's the design most teams end up with. The client knows that transfers over £10,000 need step-up, so it checks the amount, and if it's high it sends the user through a step-up flow before calling the API.
The policy now lives in the client. Every client. The web app, the mobile app, the partner integration, and the internal tool someone built in a weekend. Change the threshold to £5,000 and you're chasing four codebases, one of which you don't control. And any client that simply doesn't implement the check calls the API successfully, because the API was never the thing enforcing it.
RFC 9470 fixes this, and it's the piece of this subject most likely to be new even to people who've built step-up before. It defines a challenge the resource server returns when the token it received is real, valid, and insufficiently strong:
HTTP/1.1 401 Unauthorized
WWW-Authenticate: Bearer error="insufficient_user_authentication",
error_description="A phishing-resistant credential is required for this transfer",
acr_values="urn:acme:acr:phishres",
max_age="300"
Read what that does. The API — the only component that actually knows what's being requested and how much money is involved — states the requirement. The client doesn't decide, it reacts: it takes those parameters, sends the user back to the authorization server with them, gets a token at the required level, and retries. The client needs no knowledge of the policy at all. It needs to know how to handle a 401 of this shape.
That's the same inversion as WWW-Authenticate in ordinary HTTP auth, and it lands the policy where it belongs: next to the resource. A new client written by a team you've never met gets the enforcement for free, which is the actual test of whether a security control is architectural or conventional.
sequenceDiagram
participant U as User
participant C as Client
participant AS as Authorization server
participant RS as Payments API
C->>RS: POST /transfers (token: acr=mfa)
RS-->>C: 401 insufficient_user_authentication<br/>acr_values=phishres, max_age=300
C->>AS: authorize(claims: acr essential=phishres, max_age=300)
AS->>U: passkey prompt — "Transfer £40,000 to ACME Ltd"
U-->>AS: ✔
AS-->>C: token: acr=phishres, auth_time=now
C->>RS: POST /transfers (retry)
RS-->>C: 201 Created
The max_age in the challenge is doing quiet work in that diagram: it forces the resource server's freshness requirement into the flow, so a token that has acr=phishres from two hours ago won't satisfy it. Assurance and freshness, both enforced by the party that cares.
Binding to the transaction, not the session
Everything so far raises assurance for a session. The bank's bug was that it raised assurance and then let it apply to every subsequent transfer.
For genuinely consequential operations, the authentication has to be bound to the specific action. This is the difference between "the user did MFA recently" and "the user approved this transfer of £40,000 to this account."
Two mechanisms exist, and both are underused.
Rich Authorization Requests (RFC 9396) replace scope strings with structured detail on the authorization request:
{
"type": "payment_initiation",
"actions": ["initiate"],
"instructedAmount": { "currency": "GBP", "amount": "40000.00" },
"creditorAccount": { "iban": "GB29NWBK60161331926819" }
}
The authorization server can display those details in the prompt — the user approves a transfer, not "an operation" — and the issued token carries the same structure, so the resource server can verify that the token authorizes this payment and not merely payments in general. Change any field and the token no longer matches.
CIBA (Client Initiated Backchannel Authentication) decouples the approval channel from the request channel: the transaction is initiated on one device and approved on another, out of band. It's the right shape for call-centre and point-of-sale flows, and it's the standardized version of the thing everyone builds badly with push notifications.
You don't need either for "let me see my payslip." You want at least the first for anything irreversible, because without transaction binding, a step-up is a general-purpose elevation that any concurrent request can ride on — including, in an agentic system, a request the user knows nothing about.
That last point deserves stating directly. When there's a process acting on the user's behalf, session-scoped elevation is worth very little: the elevated session is exactly what the agent is using. Transaction-bound approval is the only version of step-up that survives an autonomous caller, because it requires the human to confirm particulars that only exist for one specific action.
Making it not miserable
The most common failure of step-up isn't a security hole. It's a design that users route around.
Don't make it a second full login. Sending the user back through username, password, and a second factor when you only needed the second factor is the pattern that gets step-up removed from the roadmap after the first round of complaints. The user has a session; they're authenticated. You need one additional proof. prompt=login with a hint of who's logging in — and better, a passkey prompt that's a single touch.
Preserve context ruthlessly. The user was mid-form. If they come back to an empty form, you've taught them to avoid the operation, and they'll do it from a path that doesn't trigger the check if one exists. That means state that survives the redirect round-trip, and it means testing the back button, which nobody does.
Show what's being approved. A prompt saying "confirm your identity" is a prompt the user approves reflexively, which is the MFA-fatigue failure mode in miniature. "Approve transfer of £40,000 to ACME Ltd" is a prompt they read. This is a security property, not a UX nicety: it's the only defence against approving something you didn't initiate.
Set an explicit elevation window, and make it short. Five to fifteen minutes, and it must be an absolute expiry evaluated from auth_time, not a sliding one. A sliding window means an active session never de-elevates.
Instrument abandonment. If 40% of users hit the step-up and don't complete the operation, you have a broken flow, not a security success. That metric is the one that tells you whether the control is working or is simply being avoided.
A short checklist
Worth reviewing against your own implementation — most of these fail at least twice:
- Does your IdP issue
auth_timeandacr, and does anything downstream actually read them? - Is your
acrlist ordered, comparable, and shorter than five entries? - Do services compare
acrlevels ordinally, or do they do string equality against one expected value? (String equality means a higher level fails the check.) - Is the requirement expressed as an essential claim, or as an
acr_valueshint the IdP is free to ignore? - Can a resource server issue a step-up challenge, or does every client hold the policy?
- Is the elevation window absolute, and does anything log when it's used?
- For irreversible operations, is the approval bound to the transaction's parameters — or to the session?
The last one is the bank's bug, and it's the one that survives every code review, because the code is correct. It just answers a slightly easier question than the one the business asked.