Why Every Identity Platform Eventually Looks the Same

If you open up the architecture of Auth0, Okta, Keycloak, Ping, and any serious in-house identity platform built in the last decade, something slightly uncomfortable becomes obvious: they all look alike. Not similar in the way that all web apps look alike. Similar in the specific, structural way tha

If you open up the architecture of Auth0, Okta, Keycloak, Ping, and any serious in-house identity platform built in the last decade, something slightly uncomfortable becomes obvious: they all look alike. Not similar in the way that all web apps look alike. Similar in the specific, structural way that suggests nobody had much of a choice.

There's a control plane and a runtime. There's a cache in front of the database. There's a key management story with rotation and a JWKS endpoint. There's an event bus feeding audit, provisioning, and notifications. There's a federation layer that speaks SAML and OIDC. There's a session store that isn't quite the same thing as the token store.

The interesting question isn't whether they converge. It's why. Because the answer isn't that everyone copied Okta. Most of these systems were designed independently, by teams that thought they were making original choices, and they arrived at the same shape anyway.

That's not imitation. That's convergent evolution.

The B-tree comparison

Almost every serious database, built by different companies across different decades with different goals, ends up storing indexes as some variant of a B-tree. Postgres, MySQL, Oracle, SQL Server, SQLite. Not because their architects all read the same paper and agreed. Because the physical constraints are identical: you're storing more data than fits in memory, on a device where sequential access is dramatically cheaper than random access, and you need range queries to be fast. Given those constraints, the solution space narrows brutally. B-trees aren't a choice so much as a discovery.

Identity platforms have their own version of this. The constraints are so specific, so non-negotiable, and so universally shared that the architecture stops being a design decision and starts being a consequence.

Let me walk through the constraints and show how each one forces a piece of the architecture into existence.

Constraint 1: Reads dominate writes by orders of magnitude

Every request to every protected resource in your entire system triggers an identity operation. Token validation, session lookup, key fetch, policy check. Meanwhile, writes — a user changing a password, an admin creating a client, a tenant updating its SAML config — happen comparatively never.

The ratio isn't 10:1. In a mature deployment it's closer to 10,000:1.

Any system with that read/write asymmetry gets pushed, hard, toward the same answer: separate the path that serves reads from the path that accepts writes, and put aggressive caching in front of the read path. You cannot have configuration writes contending with token validation reads on the same hot tables and expect either to behave well.

This single constraint is responsible for more of the architecture than anything else. It's what produces the control plane / runtime split, which every mature identity platform has, whether or not they use those words for it.

flowchart LR
    Admin["Admin writes\n(rare, consistency-critical)"] --> CP["Control Plane"]
    CP -.->|"projected config"| RT["Runtime"]
    Users["Auth traffic\n(constant, latency-critical)"] --> RT
    RT --> Cache["Cache"]

Constraint 2: Configuration correctness and protocol latency are incompatible goals

A control plane wants serialized writes, strong consistency, full audit of every change, and validation that refuses anything ambiguous. It's fine if an admin API call takes 400ms.

A runtime wants to be stateless, horizontally scalable, and finish a token validation in single-digit milliseconds. It cannot afford a synchronous lookup against a consistency-critical store on every request.

These two sets of requirements don't just differ — they actively fight. Consistency mechanisms that make the control plane trustworthy are exactly the mechanisms that make the runtime slow. So every platform that scales eventually separates them, and then faces the follow-on question: how does configuration get from one to the other?

There are only really three answers. Synchronous calls (rejected, because it reintroduces the latency and couples availability). Shared database tables (rejected eventually, because it recreates the write contention and couples deployments). Or asynchronous projection — the control plane publishes configuration changes, the runtime consumes and caches them locally, and you accept bounded eventual consistency as the price.

Everyone lands on the third option. Everyone then discovers the same consequence: a config change takes a moment to propagate, and you have to decide what the runtime does when its projected state is stale or missing. The mature answer is always some version of "fail secure, never fall back to a synchronous call" — because a fallback path that only activates under load is a fallback path that will take you down under load.

Constraint 3: Cryptographic keys must rotate, but tokens outlive rotations

You sign tokens with a private key. Security practice and compliance both say that key must rotate periodically. But tokens signed with the old key are still in flight, still valid, still being presented by clients who have no idea a rotation happened.

This forces overlapping key validity, which forces publishing multiple active public keys, which forces a discovery mechanism for consumers to fetch them, which is exactly what JWKS is. And since consumers now fetch keys over the network on a path that must be fast, JWKS responses get cached, which forces cache invalidation semantics and key IDs (kid) so a consumer can tell which key a given token needs.

Nobody designs JWKS because it's elegant. You design it because you tried to rotate a key once, broke every downstream service simultaneously, and worked backward from "that must never happen again."

Constraint 4: Everything must be auditable, and audit can't be in the request path

Identity systems are the first thing a security auditor, an incident responder, or a support engineer asks about. Who logged in, from where, when, what changed, who changed it. This isn't optional and it isn't just compliance theater — when something goes wrong at 2am, the audit log is the only thing that can tell you what actually happened.

But writing an audit record synchronously in the login path means your login latency is now coupled to your audit store's availability. That's an unacceptable trade: you've made the most critical user-facing flow in your system depend on a component whose failure should be invisible to users.

So audit goes asynchronous. Which means you need a durable, ordered, replayable channel between "the thing that happened" and "the thing that records it."

You just invented the event bus.

Constraint 5: More than one thing needs to react to identity events

Once the event bus exists for audit, the second consumer arrives almost immediately. SCIM provisioning needs to know when a user is created. Notification services need to know when a password changed. Analytics needs login patterns. Session revocation needs to know when a user is deactivated. Downstream applications want webhooks.

Each of these, implemented as a direct synchronous call from the login path, would be a new availability dependency and a new source of latency. Implemented as an event consumer, each is independently deployable and independently failable.

This is why every identity team eventually runs Kafka, or something shaped like it — not because they wanted distributed streaming infrastructure, but because they had five consumers of the same events and no other sane way to fan them out.

flowchart LR
    RT["Runtime"] --> Bus["Event Bus"]
    Bus --> Audit["Audit"]
    Bus --> SCIM["Provisioning"]
    Bus --> Notify["Notifications"]
    Bus --> Metrics["Analytics"]
    Bus --> Hooks["Webhooks"]

Constraint 6: Multi-tenancy is either structural or it's a future incident

If one deployment serves multiple customers, isolation has to be enforced somewhere. You can do it with a WHERE tenant_id = ? on every query and a code review culture that catches the ones you forget. Or you can make tenant context a required input that resolves before any lookup happens, so that a query without tenant scope is structurally impossible rather than merely discouraged.

Every platform that's been around long enough moves toward the second, usually after an incident or a near-miss that makes the first approach feel unacceptably fragile. And the strongest version of it — separate issuers, separate signing keys per tenant, tokens that simply cannot validate across tenant boundaries — shows up repeatedly across vendors because it's the version where cross-tenant leakage requires a cryptographic failure rather than a forgotten SQL clause.

Constraint 7: Enterprise customers arrive with protocols you didn't choose

You don't get to pick which identity protocol your customers use. They arrive with Azure AD, or Okta, or a Ping deployment from 2014, or an on-prem ADFS that someone is afraid to touch. Some speak SAML. Some speak OIDC. Some want you to be their IdP; some want you to consume theirs.

So every platform grows a federation layer that abstracts "an external identity source" into an internal model, with claim mapping to translate whatever the external system calls things into whatever you call things. The shape of that layer — connection config, certificate handling, attribute mapping, a test/debug mode because federation always breaks in ways that are hard to see — is remarkably consistent across products, because the problem is remarkably consistent.

What convergence actually tells you

There's a temptation to read all this as depressing. If everyone ends up in the same place, what's the point of thinking hard about architecture?

I'd read it the opposite way. Convergence is evidence that the problem is well-understood and the solution space has been thoroughly explored. When independent teams keep discovering the same structure, that structure is probably load-bearing. It means the architecture isn't arbitrary — it's the residue of a lot of people hitting the same walls.

The practical implication is worth stating plainly: if you're building an identity system and your architecture doesn't have these pieces, you haven't avoided them. You've deferred them. The control plane / runtime split will show up the first time an admin operation degrades login latency. The cache layer will show up the first time your database is the bottleneck for token validation. The event bus will show up when your third consumer of login events makes the synchronous-call approach untenable. Key rotation will show up the day you need to rotate a key and realize you can't without downtime.

Where the real differentiation lives isn't in the boxes on the diagram — everyone has the same boxes. It's in the quality of the answers to the questions those boxes force you to ask. How stale can projected config be before the runtime refuses to serve? What's the cache invalidation story when a tenant changes a policy? Does the audit log tell you enough to reconstruct an incident, or just enough to satisfy a checkbox? Can you rotate a signing key on a Tuesday afternoon without a maintenance window?

Those answers are where identity platforms actually differ from each other. The architecture diagram just tells you they're all solving the same problem — and that they've all figured out the same thing about what that problem demands.