The Identity Platform I Wanted to Exist
Every identity platform I've worked with has asked me to make the same choice, and I've never been convinced the choice is real.
Every identity platform I've worked with has asked me to make the same choice, and I've never been convinced the choice is real.
On one side: something simple. A library, a hosted service with a good quickstart, a working login in twenty minutes. Excellent until an enterprise customer asks for SAML with their own certificate rotation, SCIM deprovisioning, per-tenant session policy, and an audit export their SIEM can ingest. Then you discover the simplicity was purchased by not modeling any of that, and the additions arrive as bolt-ons that don't share concepts with the thing you already built.
On the other side: something enterprise-grade. It does all of the above. It also takes a quarter to deploy, requires a specialist to configure, and greets a two-person startup with a screen containing forty fields, thirty-six of which are irrelevant to them and none of which are labeled in terms they'd recognize.
The industry treats this as a natural law — that identity products segment by customer size because the requirements genuinely diverge. I think that's mostly wrong. The requirements diverge. The concepts don't. A startup and a Fortune 500 both need to know who a user is, which organization they belong to, what they're allowed to do, and what happened. The Fortune 500 needs more configuration of those same concepts, not different ones.
What I wanted was a platform where growing up meant configuring more, not learning a different system.
How products end up on one side or the other
It's worth being fair about why the fork exists, because it isn't laziness.
Products that start simple make an entirely reasonable early decision: model the common case cleanly, skip the enterprise machinery nobody's asking for yet. Then enterprise deals arrive and each requirement gets added as a feature adjacent to the existing model rather than integrated into it — because integrating it would mean reworking the core, and there's a deal closing this quarter. Five years later there's a coherent original model with a ring of enterprise features orbiting it, each with its own vocabulary.
Products that start enterprise make the opposite reasonable decision: model the full generality up front, because that's what the buyers demand. The result is correct and expresses every requirement — in the vocabulary of the protocols rather than the vocabulary of the person configuring it. Then someone tries to make it approachable, and what gets built is a quickstart wizard that generates configuration the user can't subsequently understand or modify.
Both paths produce a product where the simple experience and the powerful experience are different systems sharing a login page. That's the thing I wanted to avoid, and avoiding it turns out to require a handful of decisions made early, each with a real cost.
Decision one: standards at the core, not a translation layer
Plenty of platforms treat OAuth, OIDC, and SAML as protocols they export — there's a proprietary internal model, and adapters translate it outward. This is pleasant for the vendor. Everything you learn is vendor-specific, which is also pleasant for the vendor.
I wanted the standards to be the actual model. Not because standards are elegant — SAML is not elegant — but because the accumulated design knowledge in those specifications is worth more than anything I'd invent, and because it means the knowledge a team builds is transferable. If someone spends two years learning how the platform handles token exchange, they've mostly learned RFC 8693, which will still be true if they leave.
The cost is real: you inherit the protocols' complexity and their awkward edges, and you can't paper over them with a nicer abstraction of your own design. Some things are harder to express than they'd be in a bespoke model.
The compensating move is to keep the standards at the core and put intent at the surface. When someone registers an autonomous agent, they shouldn't be choosing grant types and token endpoint auth methods — they should be declaring what kind of thing it is and what it's for, and the platform should derive the protocol configuration from a policy. Underneath, it's an ordinary OAuth client with ordinary grants. Nothing proprietary. But the person configuring it never had to think in protocol vocabulary to get a correct result.
That's the pattern I keep returning to: standards underneath, intent on top. Not a proprietary model with standards bolted to the side, and not raw protocol surface handed to the operator.
Decision two: configuration, and specifically not customization
This is the decision that costs the most in the short term and matters most in the long term, and I've written about it at length elsewhere.
The short version: a platform that lets customers inject code into the authentication path trades away determinism, latency bounds, availability, and the ability to refactor its own internals. Those are precisely the properties an identity platform is selling. Every "can I just run a small Java class during login" that gets a yes is a piece of the product's core value quietly spent to close one deal.
So: no customer code in the runtime. Declarative claim mapping, a policy engine, async events, and pre-synchronized attributes instead.
The cost is not hypothetical. There are requests we cannot satisfy as directly as a plugin API would, and there will be deals where a competitor's plugin system is the deciding factor. The bet is that "not that way, but here's how" holds for the large majority of real needs, and that the platform being fast, predictable, and upgradeable is worth more over five years than the flexibility would have been. That's a bet, not a certainty.
Decision three: isolation that's structural rather than remembered
Multi-tenancy is easy to get 95% right and the last 5% is where breaches live. A WHERE tenant_id = ? on every query works until the one query where someone forgets, and "we're careful in code review" is not a security control.
What I wanted was isolation that doesn't depend on anyone remembering. Tenant context resolved before any lookup happens, so an unscoped query isn't a bug you catch — it's code that can't be written. And the strongest available version at the token layer: every tenant its own issuer, with its own signing keys, so that a token from one tenant failing against another isn't a matter of application logic being correct. It's a matter of the cryptography not permitting it.
The cost is rigidity. Cross-tenant features — a user who legitimately belongs to two organizations, platform-wide administrative views — are harder to build when isolation is structural rather than filtered. You end up building explicit, audited mechanisms for the cases that genuinely need to cross the boundary, rather than just... querying across it. That's more work, and it's the right kind of more work.
Decision four: separating what configures from what executes
Configuration wants strong consistency, validation, and a full audit trail; a 400ms admin API call is fine. Protocol execution wants statelessness, horizontal scale, and single-digit-millisecond responses. These requirements don't merely differ — the mechanisms that make one trustworthy are what make the other slow.
So the control plane and the runtime are separated, and the runtime never synchronously calls the control plane during a request. Configuration reaches the runtime by projection: published, consumed, cached locally. If the control plane is down, admins can't change things and users keep logging in — which is the correct blast radius, because those two incidents deserve different severities.
The cost is eventual consistency, and it's a cost I'd rather state plainly than paper over. A configuration change takes a moment to propagate. That window is bounded and measurable, but it is not zero, and any platform claiming its config changes are instantaneous is either wrong or has hidden a synchronous dependency somewhere it will hurt later. Where an operation genuinely can't tolerate the window — emergency revocation — that gets a targeted fast path, rather than making the entire configuration system synchronous to serve the rare case.
Decision five: machine identity as a concept, not a workaround
Most identity platforms model humans, then accommodate machines by handing them a client credential and calling it a service account. That worked when machines were a handful of backend integrations. It works considerably less well now that a meaningful share of authentications come from automated callers — CI systems, scheduled jobs, and increasingly agents acting on a specific person's behalf.
An agent doesn't fit either existing shape. It isn't a human at a browser, so the interactive grants don't apply. But it usually is acting for a specific person who simply isn't present at the moment of the call — so a pure machine identity loses the thread of who the work is ultimately for.
I wanted that to be a first-class concept rather than a convention teams invent locally: an agent as its own registered object with an owner, a purpose, and a lifecycle, governed by policy rather than hand-configured, with delegation expressed through standard token exchange so a resource server several hops downstream can still see the whole chain — the original user, and every actor standing in for them.
The cost here is different in kind: it's a bet on where things are going. If agentic access stays niche, this is a concept that didn't need to be first-class. I don't think that's how it goes, but I'd rather name it as a bet than pretend it's obvious.
What this makes the platform bad at
An honest version of this essay has to include the cases where these decisions are the wrong ones.
If you want a platform you can extend with arbitrary code, this is the wrong product, deliberately and permanently. That's not a roadmap gap.
If you need instantaneous global configuration propagation with no bounded window, the architecture doesn't offer it, and I'd be suspicious of anyone who claims to.
If you're a small consumer app with no multi-tenancy, no federation, and no compliance obligations, a good auth library is probably a better fit than any identity platform, including this one. The value here shows up when identity becomes a governance problem, and if it hasn't yet, you'd be paying for machinery you don't need.
And if your evaluation weights ecosystem breadth and years of production hardening above architectural fit, the incumbents have a decade of advantage that no amount of design coherence offsets. That's a legitimate way to choose.
The thing I actually wanted
Not a product with a startup tier and an enterprise tier. One set of concepts — tenants, clients, resource servers, policies, agents — that a two-person team can use on day one with almost nothing configured, and that a large enterprise can configure heavily without switching to a different mental model or a different product.
Enterprise capability without enterprise complexity isn't a slogan; it's a claim about where complexity should live. It should live in configuration, which can be defaulted, templated, and left alone. It should not live in concepts, which every user has to learn regardless of size.
Whether that's been achieved is genuinely not for me to assess. But it's what the decisions above were for, and each one has a cost I'd rather have stated than discovered.
Disclosure: I work on ClavionX, which is the platform these decisions describe. I've tried to write this as an argument about design trade-offs rather than a case for the product — partly because that's more useful, and partly because the trade-offs apply whether you end up here, on Keycloak, on a commercial platform, or on something you build.