The 10 Architectural Decisions That Define Every Identity Platform

Every identity platform is the sum of about ten decisions, most of which get made in the first six months and none of which get revisited cheaply. They're not implementation details. They're the load-bearing choices that determine what the system can do five years later — and more importantly, what

Every identity platform is the sum of about ten decisions, most of which get made in the first six months and none of which get revisited cheaply. They're not implementation details. They're the load-bearing choices that determine what the system can do five years later — and more importantly, what it can never do without a rewrite.

What follows is the list, along with what each choice actually costs you. There are no universally correct answers here. There are answers with known consequences, and the mistake is almost always making the choice implicitly instead of deliberately.

1. Is tenant isolation structural or filtered?

The fork: Do tenants share tables with a tenant_id column, get their own schemas, or get their own databases entirely?

Shared tables with a filter is the cheapest to build and the easiest to operate — one migration, one connection pool, one backup. It's also the option where a single forgotten WHERE clause is a cross-tenant data breach. The mitigation isn't "be careful"; it's making tenant context a required parameter that resolves before any query is constructed, so an unscoped lookup isn't a bug you catch in review — it's code that doesn't compile.

Database-per-tenant gives you the strongest isolation and the most painful operations: migrations across thousands of databases, connection pool exhaustion, and per-tenant backup complexity that grows linearly with sales.

What actually decides it: how much you'd pay to make cross-tenant leakage structurally impossible rather than procedurally unlikely. Regulated industries push toward isolation. Volume pushes toward sharing. Most platforms end up with shared storage plus isolation enforced hard at a layer above the query.

The related decision people miss: should each tenant be its own token issuer? If yes — separate issuer URL, separate signing keys — then a token from Tenant A cannot validate against Tenant B even if every application-layer check fails, because the cryptography won't allow it. That's a meaningfully different security posture than isolation that depends entirely on correct code.

2. Control plane and runtime: separate or unified?

The fork: Does the system that configures identity share code, database, and deployment with the system that executes protocol flows?

Unified is faster to build and simpler to reason about, right up until the two workloads start interfering. Admin operations want strong consistency and full validation. Protocol flows want statelessness and single-digit-millisecond responses. Sharing a database means schema migrations for admin features block runtime traffic, and a control plane incident becomes a login outage.

Separating them forces a harder question immediately: how does configuration cross the boundary? Synchronous calls from runtime to control plane reintroduce exactly the coupling you separated to avoid — now your login path fails when your admin API is degraded. Shared tables reintroduce the write contention. The remaining answer is asynchronous projection: the control plane publishes, the runtime consumes and caches locally, and you accept bounded eventual consistency.

The consequence you must design for: if the runtime consumes projected state, what does it do when that state is stale or absent? The only safe answer is fail-secure, with no synchronous fallback. A fallback path that only activates when projection is broken is a fallback path that activates during exactly the incident where it will make things worse.

3. JWT or opaque access tokens?

The fork: does the token carry its own claims, or is it a reference that must be introspected?

JWTs are self-validating. A resource server checks a signature locally and never calls you. That's fantastic for latency and availability, and it means your token endpoint isn't in the critical path of every downstream API call.

It's also why revocation is hard. A JWT is valid until it expires, full stop. "Revoke this token now" isn't something the format supports — you can only shorten lifetimes and accept a revocation window, or maintain a denylist that every validator must check, at which point you've reintroduced the network call you chose JWTs to avoid.

Opaque tokens invert everything: revocation is instant because the token means nothing without a lookup, but now every validation is a network round trip to your introspection endpoint, which is now on the critical path of your customers' entire API surface.

The usual resolution: short-lived JWTs (minutes, not hours) plus refresh tokens that carry the actual revocable state. You get local validation for the common case and a real revocation point at refresh time. The revocation window becomes a tunable number rather than an architectural dead end.

4. Where does session state live?

The fork: Redis, the database, or the cookie itself?

Cookie-only sessions are stateless and scale beautifully, and they make server-side termination effectively impossible. "Log this user out everywhere, right now" is a requirement that arrives eventually, and it arrives from a security team during an incident.

Redis is the common answer — fast, TTL-native, purpose-built for this. It also becomes an availability dependency for every authenticated request, which means your identity platform's uptime is now bounded by your Redis cluster's uptime, and you need a real answer for what happens during a failover.

Database-backed sessions survive restarts and give you queryability ("show me all active sessions for this user"), at meaningfully higher latency per request.

The thing to decide early: whether sessions are recoverable across a cache flush. If a Redis restart logs out every user on the platform simultaneously, that's a decision — make sure it's one you made on purpose.

5. What's the caching strategy, and what invalidates it?

The fork: local in-process cache, distributed cache, or both.

Local caches are the fastest thing available and the hardest to invalidate — every instance holds its own copy, and a config change has to reach all of them. Distributed caches are invalidated in one place and cost a network hop.

Most mature platforms run both: a distributed cache as the shared tier, plus short-TTL local caches for the hottest, most stable data (signing keys, tenant config, policy). The local tier is what keeps the p99 honest.

The question that matters more than the topology: what is your invalidation story when a tenant changes something security-relevant? If an admin disables a user, how long until every runtime instance stops honoring that user's sessions? "Whenever the TTL expires" is an answer, but it should be a number you've chosen and can defend, not a number you discovered during an incident.

6. Event-driven or direct calls?

The fork: when a login happens, does the runtime call audit/provisioning/notifications directly, or publish an event?

Direct calls are simpler and traceable in a stack trace. They also mean the login path's latency and availability are the sum of every consumer's latency and availability. Add a fifth consumer and you've added a fifth way for login to fail.

Events decouple all of that, at the cost of a real infrastructure dependency, eventual consistency, and the operational burden of consumer lag, replay, and ordering guarantees.

The tell that you need events: the moment there's a second consumer of the same event. One is a function call. Two is a design smell. Three is a bus you haven't admitted to building yet.

7. Where does authorization live?

The fork: does the identity platform decide what a user can do, or only who they are?

Put authorization in the token — roles, permissions, scopes as claims — and you get fast, local decisions with no callback. You also get stale permissions until the token expires, and tokens that grow uncomfortably large for users with many roles. There's a real limit here: HTTP header size caps are not theoretical once a token carries a few hundred permissions.

Put authorization in a policy service and decisions are always current and arbitrarily expressive, but now every request has a policy lookup, and that service is in everyone's critical path.

The pattern that works: coarse-grained, slow-changing claims in the token (tenant, role, group). Fine-grained, resource-specific decisions at the application layer, where the context actually lives. The identity platform shouldn't be deciding whether Alice can edit this specific document — it doesn't know what a document is, and building it so it does is how identity platforms become unmaintainable.

8. How do keys rotate?

The fork: manual rotation with a maintenance window, or automatic rotation with overlapping validity.

This decision is made once and regretted for years if you get it wrong, because the wrong answer is invisible until the day you need to rotate under pressure.

Rotation requires: multiple simultaneously-valid keys, a kid on every token so validators know which key to use, a JWKS endpoint publishing the current set, and a caching strategy for consumers that refetch. Miss any one of these and rotation means downtime.

The test: can you rotate a signing key right now, on a weekday afternoon, without notifying anyone? If not, you don't have key rotation — you have a key rotation project, and you'll be running it under emergency conditions the day a key is compromised.

9. Single region or multi-region?

The fork: and specifically, is identity data replicated, partitioned, or pinned?

Identity is uniquely hostile to multi-region because it's simultaneously latency-sensitive (every request touches it), consistency-sensitive (a revoked session must be revoked everywhere), and residency-sensitive (that customer's user data legally cannot leave the EU).

Full replication gives low latency everywhere and forces you to confront replication lag on security-critical state. Regional partitioning respects residency and means users can't roam. Active-passive is simplest and wastes half your capacity while making failover a rehearsed event rather than an automatic one.

Decide early whether tenants are region-pinned, because retrofitting residency onto a globally-replicated identity store is one of the genuinely miserable migrations in this space.

10. What does audit actually capture, and is it immutable?

The fork: log lines, or an append-only event stream that is itself a first-class data model.

Audit-as-logging is easy and fails the moment someone asks a question logs can't answer: what was this tenant's configuration on the day of the incident? Who changed it? What did the previous value look like?

Append-only, event-sourced audit answers those questions and costs you storage, a retention policy, and the discipline to never mutate a record.

The framing that gets this prioritized correctly: audit isn't a compliance feature. It's the debugging tool for the one system where you can't reproduce the bug locally. When a customer says "our users got logged out at 3pm yesterday and we don't know why," the audit trail is the entire investigation. Teams that treat it as a checkbox for auditors build something that satisfies auditors and helps nobody at 2am.

The meta-decision

Read that list again and notice what these ten have in common: each one is cheap to decide on day one and expensive to change on day one thousand. They're not the decisions that make a product good — features do that. They're the decisions that determine whether adding features stays possible.

The most common failure I've seen isn't picking the wrong option. It's not noticing that a decision was being made at all — shipping a synchronous audit call because it was the obvious way to write the code, and discovering three years later that it's load-bearing in the login path and can't be removed without touching everything.

Make these ten explicitly. Write down what you chose and why. The next architect will need it, and there's a decent chance the next architect is you, several years from now, with no memory of why any of this looked like a good idea.