Building Identity Like Kubernetes

The first time I explained our identity architecture to a platform engineer, I was three diagrams in and losing them. Then I said "it's basically Kubernetes," and they got the entire thing in about fifteen seconds.

The first time I explained our identity architecture to a platform engineer, I was three diagrams in and losing them. Then I said "it's basically Kubernetes," and they got the entire thing in about fifteen seconds.

That shortcut turned out to be more than a convenient analogy. The structural correspondence is close enough that Kubernetes' hard-won lessons about desired state, controllers, and reconciliation transfer almost directly — and the places where identity teams get into trouble are usually the places where they've violated a principle Kubernetes made explicit years ago.

The correspondence

Kubernetes splits the world in two. There's a control plane where you declare what you want — deployments, services, config — which is validated, versioned, and stored in etcd. And there's a data plane, the kubelets and running pods, which does the actual work of serving traffic.

The control plane never serves your users' HTTP requests. The data plane never decides what should be running. Between them sit controllers that watch for differences between declared state and actual state and work to close the gap.

An identity platform has exactly this structure, whether or not the team building it uses these words:

Kubernetes Identity Platform
API server (declare desired state) Control plane / admin API
etcd (authoritative store) Configuration database
Controllers (reconcile) Projection / sync workers
kubelet + pods (do the work) Runtime — token issuance, login flows
Node-local cache Runtime's projected config cache
Events Audit and event stream
flowchart LR
    Admin["Administrator"] --> CP["Control Plane\n(desired state)"]
    CP --> Store["Config store\n(authoritative)"]
    Store --> Proj["Projection\n(reconciliation)"]
    Proj --> RT["Runtime\n(protocol execution)"]
    Users["Auth traffic"] --> RT
    RT --> Cache["Local projected state"]

An admin declaring "this tenant requires MFA and its access tokens live 900 seconds" is doing exactly what a kubectl apply does: writing desired state to an authority, and trusting that reconciliation will make reality match.

Rule one: the data plane never calls the control plane

This is the principle Kubernetes is most rigid about, and it's the one identity teams most often break.

A kubelet doesn't ask the API server what to do for each incoming packet. It's told what should be running, it caches that, and it keeps running it — even if the API server is completely down. You can lose the entire Kubernetes control plane and your running workloads keep serving traffic. That's not an accident; it's the central design goal.

The identity equivalent: the runtime must never make a synchronous call to the control plane during request execution.

The temptation to break this is enormous, because the alternative feels harder. You need a tenant's password policy during login — the control plane owns it, so just fetch it. It works in development. It works in staging. It works in production, right until the control plane has a slow query, and now every login on the platform is slow, because you made your most latency-sensitive path depend on your least latency-optimized service.

Worse is the fallback pattern: cache the config, but on a cache miss, call the control plane. This looks like defensive engineering and is actually a loaded gun. That fallback path is exercised approximately never during normal operation — which means it's untested, unoptimized, and completely un-load-tested. It activates for the first time during an incident, when cache state is bad and traffic is high, and it turns a degraded cache into a full outage by stampeding the control plane.

Kubernetes' answer is that a kubelet with a stale view keeps doing what it was last told. The identity equivalent is that a runtime with stale projected config keeps enforcing the last configuration it successfully received, and if it has no config, it fails secure — refuses the request — rather than falling back to a synchronous fetch or, catastrophically, to a permissive default.

Rule two: reconciliation, not orchestration

The naive way to propagate a config change: the admin API updates the database, then calls each runtime instance to tell it about the change.

That's orchestration, and it's brittle for the same reasons Kubernetes rejected it. What if an instance is mid-restart? What about an instance that scales up thirty seconds later — who tells it? What if one call fails? Now you have some instances on the new config and some on the old, with no mechanism that will ever converge them.

The reconciliation model instead: the control plane records desired state and emits that a change occurred. Runtime instances converge toward that state on their own — by consuming events, or polling a version, or both. A new instance starting up doesn't need anyone to catch it up; it reconciles from current desired state as part of coming online.

The property this gives you is the valuable one: the system is self-healing rather than dependent on every notification succeeding. A dropped event is a delay, not a permanent divergence, because the next reconciliation cycle catches it.

This is also why identity platforms end up with a version or generation number on projected config. It's the same thing as a Kubernetes resource version — it lets a runtime answer "am I current?" without diffing the entire configuration, and it lets you build the one dashboard you'll actually want at 3am: which instances are lagging, and by how much?

Rule three: eventual consistency is a feature you must bound

Kubernetes doesn't pretend to be immediately consistent. You apply a deployment; pods come up over some seconds. Everyone accepts this because it's explicit, and because the convergence window is bounded and observable.

Identity teams often struggle here, because "eventually consistent security config" sounds alarming. An admin disables a user — surely that must be instant?

The honest answer: it isn't, and pretending otherwise leads to worse architecture. Even a synchronous update has propagation delay through caches, replicas, and in-flight requests. The difference between systems isn't whether there's a window — it's whether the team knows how big the window is.

So bound it and publish it. "Configuration changes take effect within N seconds across all runtime instances" is a real engineering commitment you can measure, alert on, and tell a customer's security team. "Changes are instant" is a claim that's false and will be discovered to be false during an audit.

And where a specific operation genuinely can't tolerate the window — emergency session revocation, a compromised credential — build a targeted fast path for that specific operation rather than making the entire configuration system synchronous. Kubernetes does the same thing: most changes reconcile, but pod deletion has a direct path, because "stop this now" has different requirements than "make this the new normal."

Rule four: separate the failure domains

The strongest argument for this architecture isn't performance. It's blast radius.

Control plane down: admins can't change configuration. Logins continue normally, running on last-known-good projected state. Users notice nothing. This is a work-hours incident.

Runtime down: nobody can log in. This is a 3am, everyone-wakes-up incident.

Those two events have completely different severities, and the architecture is what keeps them separate. If they share a database, share a deployment, or share a process, then every control plane incident is also a runtime incident, and you've taken your annoying-but-survivable failure mode and merged it into your catastrophic one.

The practical test is a deployment question: can you deploy the admin API on a Wednesday afternoon without a login outage risk? If the answer is no, you don't have a separation. You have two names for one system.

Where the analogy stops

Three places it breaks down, worth naming so nobody pushes it too far.

Identity's data plane is stateful in ways pods aren't. Sessions, refresh token families, replay-protection state — the runtime genuinely owns mutable state that can't be treated as disposable. You can't kill an identity runtime instance as casually as a stateless pod, unless that state lives entirely in a shared store.

The security consequences of staleness are asymmetric. A stale deployment config means old code runs a bit longer. A stale authorization config means someone has access they shouldn't. Identity's tolerance for reconciliation lag is genuinely lower in one direction than the other, and that asymmetry should show up in your design: propagate revocations faster than you propagate grants.

Multi-tenancy is much sharper. A Kubernetes namespace is a soft boundary. A tenant boundary in identity is a hard security boundary — cross-tenant leakage is a breach, not a noisy-neighbor problem. Identity platforms therefore push isolation deeper than Kubernetes does: separate issuers, separate signing keys, tenant context resolved before any lookup rather than filtered afterward.

Why the analogy is worth keeping anyway

Because it front-loads the right questions.

When someone proposes that the login path fetch a value from the admin database, "would a kubelet call the API server for that?" ends the discussion in one sentence. When someone worries that config propagation isn't instant, "how does Kubernetes handle this?" reframes it from a bug into a bounded, measurable window. When someone asks why there are two services instead of one, the failure-domain argument is already familiar.

Kubernetes made these patterns legible to an entire industry. Identity platforms arrived at the same structure independently, driven by the same constraints — a control loop, a fast data plane, declarative config, bounded reconciliation. Borrowing the vocabulary doesn't just make the architecture easier to explain. It makes the mistakes easier to name before you've shipped them.