The Identity Latency Budget: Where 200 Milliseconds Actually Goes
A user clicks "Sign in." Somewhere between that click and the moment they see their dashboard, they've formed an opinion about your product.
A user clicks "Sign in." Somewhere between that click and the moment they see their dashboard, they've formed an opinion about your product.
The research on this is boring and consistent: under about 100ms feels instant, around 300ms feels responsive, and past a second people start to notice they're waiting. Login is the very first interaction anyone has with your software, which means it's doing a disproportionate amount of work in setting expectations. A slow login doesn't read as "the identity service is under load." It reads as "this product is slow."
So you have a budget. Call it 200 milliseconds for the authentication round trip — ambitious but achievable. Let's spend it, line by line, and see where it actually goes.
The thing that makes this hard
Login isn't one request. That's the first thing to internalize, and it's why latency budgets that look fine on paper feel sluggish in practice.
A single "log in" is a redirect chain. The browser goes to the application, gets bounced to the identity provider, possibly bounces again to an upstream IdP, comes back with a code, exchanges it for tokens, and lands back at the application. Each hop is a full request — DNS, TLS, routing, processing — and each one is an opportunity to spend budget you don't have.
sequenceDiagram
participant B as Browser
participant App as Application
participant IdP as Identity Runtime
B->>App: GET /dashboard
App->>B: 302 → IdP /authorize
B->>IdP: GET /authorize
IdP->>B: Login page
B->>IdP: POST credentials
IdP->>B: 302 → App /callback?code=...
B->>App: GET /callback
App->>IdP: POST /token (back channel)
IdP->>App: Tokens
App->>B: 302 → /dashboard
That's five browser round trips and one server-to-server call, minimum, before the user sees anything. If each one costs 100ms, you've spent half a second and you haven't validated a password yet.
This is why the single most effective latency optimization in identity isn't making any individual operation faster — it's not doing a redirect you didn't need.
Spending the budget: the credential submission
Let's zoom in on the one request that does the real work: the user has typed their password and hit submit. Here's where the 200ms goes.
DNS: 0–100ms. Zero if cached, which it usually is by this point in the flow. But the first hop to a new identity domain isn't cached, and a cold DNS lookup can genuinely cost 50–100ms. This is a real argument for keeping your identity endpoints on a domain the browser has already resolved, and for reasonable TTLs — aggressive short TTLs on an auth domain buy you flexibility you rarely use and cost you latency on every cold client.
TLS handshake: 0–150ms. This one hurts. A full handshake is two round trips; TLS 1.3 gets it to one; session resumption gets it to zero. On a mobile connection with 80ms RTT, the difference between a resumed session and a full handshake is over 150ms — most of your budget, spent before your server has seen a single byte of the request.
The takeaway: connection reuse across the redirect chain matters enormously, and it's mostly determined by infrastructure decisions (HTTP/2, keep-alives, TLS config) rather than anything in your application code.
Gateway and routing: 1–10ms. Load balancer, ingress, WAF. Usually small, occasionally not. A WAF doing deep inspection on a POST body, or an ingress controller doing per-request cert lookups, can quietly add tens of milliseconds. Worth measuring rather than assuming.
Tenant resolution: 0–5ms. In a multi-tenant platform, before you can do anything you have to know which tenant this request belongs to, and load its configuration: password policy, MFA requirements, branding, token settings, federation config.
This should be a cache hit essentially always. Tenant config changes rarely and is read constantly — it's the single most cacheable thing in the system. If tenant resolution is hitting your database on every login, that's usually the first big win available, and it's often worth 20–50ms.
Password verification: 50–150ms. Deliberately.
Here's the line item that breaks people's mental model. This is, by design, the most expensive operation in the entire flow, and making it faster makes your system less secure.
bcrypt, scrypt, and Argon2 are intentionally slow. That's the whole point — a fast hash means an attacker with a stolen database can test billions of candidate passwords per second. The work factor is a dial that trades user-facing latency for offline attack cost, and the standard guidance lands somewhere around 100ms of work.
So roughly half your budget goes to one deliberately slow function. Everything else in the login path has to be fast because this part can't be.
Two consequences worth internalizing. First, password hashing is CPU-bound, not I/O-bound, which means it doesn't parallelize away — it consumes actual cores, and login capacity planning is really CPU capacity planning. Second, a login flood is a CPU exhaustion attack whether or not anyone intended it that way, which is why rate limiting in front of the credential endpoint is a capacity control, not just a security control.
Session creation: 1–5ms. Write a session record, typically to Redis. Fast, but it's a network hop and therefore a dependency. Under Redis failover, this line item becomes "however long failover takes," which is worth knowing before it happens.
Token signing: 1–10ms. Signing a JWT is fast if the key is in memory. It is dramatically not fast if you're calling out to an HSM or a cloud KMS per token — that's a network round trip plus the HSM's own processing, and it can be 20–50ms per signature.
Teams that use KMS-backed signing usually end up caching a key handle or using an envelope approach specifically because per-token KMS calls don't survive contact with production traffic. If you're on RSA-2048, signing is more expensive than verification; ECDSA flips that trade-off. At high volume this is a real decision, not a detail.
Audit write: 0ms — if you did it right. The login event needs to be recorded. If that's a synchronous database insert, you've just added 5–20ms to every login and coupled your login availability to your audit store.
Publish to a queue and move on. The audit record will be there in a moment, and the user won't wait for it. This is the clearest example of a general principle: anything that doesn't affect the response should not be in the response path.
Response and redirect: 1–5ms. Set cookies, build the redirect, send it. Small — but note that cookie size affects every subsequent request. A fat token in a cookie is a latency tax you pay forever, on every single request the user makes for the rest of their session.
Adding it up
| Component | Realistic | Bad day |
|---|---|---|
| DNS | 0 | 100 |
| TLS | 0 (resumed) | 150 |
| Gateway | 5 | 30 |
| Tenant config | 1 (cached) | 40 (uncached) |
| Password hash | 100 | 100 |
| Session write | 2 | 50 (failover) |
| Token signing | 3 | 50 (KMS call) |
| Audit | 0 (async) | 20 (sync) |
| Response | 3 | 5 |
| Total | ~115ms | ~545ms |
Both columns describe the same system. The difference is entirely cache state, connection reuse, and whether the non-essential work was moved off the critical path.
That's the honest summary of identity performance: the fast case and the slow case are not different architectures. They're the same architecture on a good day and a bad day, and engineering effort mostly goes toward making bad days rare.
What this means for how you build
The password hash sets your floor. If it's 100ms, you cannot be faster than 100ms, and every other component is competing for what's left. This reframes the whole optimization problem: you're not making login fast, you're preventing it from becoming slow.
Cache misses are the main variable. Tenant config, signing keys, policy, client metadata — every one of these is read constantly and written almost never. Uncached, each adds tens of milliseconds. This is why identity platforms are, in practice, mostly caching problems.
Async everything that doesn't gate the response. Audit, analytics, notifications, webhooks. If a component's failure shouldn't fail the login, it shouldn't be in the login path — those two properties are the same property.
Redirects are the hidden cost. Six round trips at 80ms of network latency each is 480ms of pure network time, and no amount of server optimization touches it. Reducing hops is worth more than micro-optimizing any single hop.
Measure the flow, not the endpoint. Your /token endpoint's p50 can be 8ms while users experience a 900ms login, because the endpoint metric doesn't include DNS, TLS, redirects, or the upstream IdP that took 600ms to respond. The number that matters is click-to-dashboard, and it's the one that's hardest to instrument. Instrument it anyway.
The part nobody budgets for
Everything above assumes local authentication. Add federation and the arithmetic changes completely.
When a user authenticates against an enterprise customer's Azure AD or on-prem ADFS, there are two more redirects and a call to a system you don't operate, don't monitor, and can't optimize. If that IdP takes 800ms, your login takes 800ms plus your own budget, and no amount of caching helps.
This is worth communicating clearly, both internally and to customers: in a federated flow, most of the latency belongs to someone else's infrastructure. What you control is the part you control — don't let it be the reason the flow feels slow, because when a federated login is sluggish, the vendor who gets blamed is the one whose login page the user was looking at.