Seven Identity Bugs Every Team Eventually Ships

Identity bugs have a distinctive personality, and it's an unpleasant one.

Identity bugs have a distinctive personality, and it's an unpleasant one.

Most bugs are cooperative. They reproduce. You attach a debugger, narrow the input, find the bad line. Identity bugs don't. They happen to 3% of users, or only in Frankfurt, and by the time you've been paged and found the request ID the token has expired, the cache entry has been evicted, and the offending pod has been recycled by an autoscaler that considered itself helpful. The evidence expires before the investigation starts.

They're also asymmetric. A rendering bug shows you a broken page. An identity bug either locks out paying customers or, much worse, lets someone in — and the second kind doesn't page anyone. It sits there working correctly, from the attacker's point of view.

And they're largely invisible in tests, because tests run on one machine, with one clock, one tenant, one signing key, and fixtures built by someone who was trying to test something else.

Seven that most teams ship eventually. None are exotic.

1. The one pod whose clock is wrong

"Roughly 3% of logins are failing with invalid_token. Users say if they just try again it works. Sometimes twice."

Every debugging instinct is now pointed the wrong way. Three percent smells like a flaky dependency or a partial deploy. It does not smell like time.

Whether you hit it depends on which pod issued your token and which validated it. If the issuing pod's clock is 90 seconds fast, its tokens carry an iat and nbf in the future from the perspective of every other pod, and validators reject them as not-yet-valid. Hit the same pod twice and everything works, because it agrees with itself. That's the "try again and it works" pattern, and it's why it never reproduces in staging, where you run two replicas and one of them is you.

issuer=pod-7f2a  iat=1754204461  nbf=1754204461
validator=pod-91c now=1754204372  -> token used before nbf (89s)

The fix has two halves and teams ship only the first: validate with a deliberate skew tolerance, a small number of seconds chosen on purpose rather than the library default nobody looked at. The half that matters is the second — treat clock drift as a first-class monitored metric, per host, with an alert. chrony and ntpd both expose the offset. Scrape it. A fleet where one node is silently 90 seconds off has a latent failure rate in every time-sensitive system you own, not just auth.

2. The session that remembers who you used to be

"I made Priya an admin twenty minutes ago and she still gets 403 on the settings page. I've checked the database three times. She's an admin."

Or the version that ends up in an incident review instead of a Slack thread: you revoked someone's admin rights during an offboarding, and they kept using them.

You can't reproduce it, because you test by logging in fresh — and a fresh login mints a fresh token with fresh claims, which works perfectly. The bug exists only for sessions that predate the change, exactly the population you never test with.

A token or session is a snapshot. It was minted at 14:02 with the claims true at 14:02 and is cheerfully valid until 15:02, because that's what "valid until" means. A write to a roles table does not reach backward into a signed artifact already sitting in someone's browser.

The fix is not "shorten lifetimes until the problem is small enough to ignore," though everyone tries that first. It's to treat a privilege change as an authentication-level event: force re-issuance of the session and everything derived from it, rotating the session identifier while you're there. That rotation is the same mechanic as session fixation defence, and the rule generalizes — any change to who the user is, or what they may do, must produce a new session identifier. Login, step-up MFA, grant, revocation, tenant switch, impersonation start and stop. If the identifier survives the transition, so does the old authority.

3. The redirect URI that matched a bit too well

"Hi — I'm a security researcher. I was able to obtain an authorization code for arbitrary users on your platform. Do you have a disclosure process?"

There's no reproduction problem here. There's a "why did nobody notice for two years" problem.

Redirect URI validation is one of the few places where the spec says exact string comparison and means it, and the place where product pressure pushes hardest the other way. A customer has forty subdomains, so someone adds prefix matching. Someone else, being tidy, normalizes both sides before comparing — lowercasing, resolving .., stripping default ports. Individually reasonable, collectively fatal.

registered:  https://app.example.com
attacker:    https://app.example.com.evil.com/cb    # prefix match says yes
attacker:    https://[email protected]/cb    # userinfo, parsers disagree
attacker:    https://app.example.com/logout?next=https://evil.com/cb   # exact match, still yours

The third is the interesting one, because it survives exact matching. The target is registered and genuinely yours — it just happens to contain an open redirect. Your authorization code lands on your own domain and is immediately forwarded, query string and all, somewhere else.

Fix: exact string comparison, no normalization, no wildcards, no "the host has to match." And accept the uncomfortable corollary — an open redirect anywhere on a registered redirect host is an authentication vulnerability in your identity system, even though it lives in someone else's codebase and their tracker will rate it "low."

4. The cache TTL nobody chose

"We disabled the account at 14:31 during the call. Their API requests kept succeeding until 14:35. The customer's security team is asking us to explain the four minutes."

It's hard to investigate because it heals itself. By the time anyone looks, the cache has expired, requests are being rejected correctly, and the system is indistinguishable from one that never had the bug.

The cause is almost always a local, per-instance cache of user or tenant state with a TTL someone set to 300 in a config file in 2022 because five minutes seemed like a nice round number. It was a performance decision. It became a security parameter the moment the cached value could answer "is this person allowed."

The fix is rarely "remove the cache" — read amplification makes that unaffordable. It's that every cached datum feeding an access decision needs a stated maximum staleness someone can say out loud without opening the code. Sixty seconds is defensible. Four minutes is probably defensible. Not knowing is what ends the deal, because the customer's next question is "what else don't you know?"

If you can't state the window in seconds, you don't have a revocation mechanism. You have a tendency.

5. Rotation day, no kid

"Auth is down. Everything. All tenants. We did the key rotation like the runbook said."

Routine rotations are the dangerous ones, because the runbook has worked four times and nobody reads it closely anymore.

Two variants, identical from the outside. In the first, tokens are issued without a kid header. With one signing key that's fine — validators try the only key they have. With two keys published during an overlap window a validator has to guess, and library behavior on ambiguity ranges from "try all keys" to "take the first" to "throw." You learn which you have at the worst possible moment.

In the second, kid is present and correct, but the validator's JWKS cache has a TTL measured in hours and hasn't seen the new key. Every token signed with it fails unknown key id, and the validator declines to refetch, because from its point of view a cache is a cache.

jwt.header = {"alg":"RS256"}                # no kid — validator is guessing
jwt.header = {"alg":"RS256","kid":"2026-08"} # kid present, JWKS cached at 2026-07

Fix: emit kid on every token, starting on day one when you have exactly one key and it seems pointless. Publish new keys well before signing with them, waiting longer than the longest plausible consumer TTL — including consumers you don't operate. And on unknown key id, refetch JWKS once before failing, rate-limited so it can't be turned into a DoS. That single conditional refetch is the difference between a rotation that's a non-event and one that's an outage.

6. Milliseconds, seconds, and the local timezone

"Tokens issued from the EU cluster expire instantly. Same code. Same build. I've diffed the deploy twice."

Or the mirror image, which is worse because nobody files a ticket: tokens that never expire. Nobody reports working software.

Two causes, both boring, both shipped constantly. The first is unit confusion: exp is UTC epoch seconds, and most modern runtimes hand you milliseconds by default. Set it in milliseconds and you've issued a token valid until the year 57,000. Set it in seconds where the validator expects milliseconds and it expired during the Nixon administration.

The second is building expiry from a local-time datetime — a naive datetime, a LocalDateTime, a Date built from parts — then converting as though it were UTC. On a container running UTC this is invisible. Where TZ is Europe/Berlin, your one-hour token is born two hours old in summer. Same code, same build, different environment variable, set eighteen months ago by someone fixing log timestamps.

exp = 1754208061          # seconds — correct
exp = 1754208061000       # ms — expires in year 57649
exp = 1754200861          # local time treated as UTC — already expired

Fix: one function constructs time claims, it takes a duration and nothing else, and it goes through an explicit UTC epoch-seconds conversion. Then assert it — not that expiry "works," but that the integer lands in a plausible band, and that the suite still passes under TZ=Australia/Adelaide, which is UTC+9:30 and has caught more bugs than any timezone deserves credit for.

7. The tenant filter that a refactor ate

"I'm in the Acme account and I can see three invoices belonging to Globex. Is this a demo data thing?"

It is not a demo data thing.

The worst bug on the list, and structurally the simplest. Someone consolidated four nearly-identical repository methods into one helper. Three had WHERE tenant_id = ?. The consolidated version kept the shape of the one that didn't — an admin query, or a background job that legitimately ran across tenants. Nothing threw. No test failed. The review looked like a net deletion of duplication, the kind of diff reviewers approve quickly and warmly.

The tests passed because the fixtures contained one tenant. With one tenant in the database, a query missing its tenant filter returns exactly the same rows as one that has it. The suite is structurally incapable of detecting this class of bug and stays that way until someone changes the fixtures.

The fix that holds isn't vigilance; vigilance loses to refactoring over a long enough horizon. Make isolation structural — a query object that cannot be constructed without a tenant context, row-level security in the database, a repository where the tenant is a constructor argument rather than a parameter someone might forget. The goal is that the broken version doesn't compile, not that someone catches it in review.

And change the fixtures. Every test database should contain at least two tenants, the second holding data no test is ever allowed to see. It costs an afternoon and turns a category of silent data-leak bugs into ordinary test failures.

The common thread

Read the seven together and two patterns account for all of them.

The first is state that was true once and is no longer. A session minted before a role changed. A cached account status from before the account was disabled. A JWKS response from before the key existed. Each was correct when captured; each is a piece of the past that outlived its accuracy and kept being consulted anyway. Distributed systems are made of snapshots, and identity is the domain where a stale one has legal consequences.

The second is two systems disagreeing about time, identity, or scope. Two pods on what time it is. An issuer and a validator on whether exp is seconds or milliseconds. Your matcher and an attacker's parser on where a hostname ends. A repository and its caller on whether tenant scoping was already applied. Both parties internally consistent, both locally correct, the bug living in the gap between them — which is precisely where no unit test looks.

That tells you where to point your attention. Not at the crypto; the crypto is library calls. Not at the protocol; the protocol is written down. Point it at the joints — where a fact crosses from one system to another and gets re-interpreted, and where you're acting on something you learned a while ago and haven't rechecked.

Write down the staleness window. Emit the kid. Put a second tenant in the fixtures. Alert on clock drift.

None of it is clever, and that's rather the point. These don't get shipped because the team wasn't smart enough. They get shipped because each one is invisible from inside the component that causes it.