Why Enterprise Capabilities Shouldn't Require Enterprise Complexity
The call was scheduled for thirty minutes and ran ninety.
On one side, a customer's IdP administrator, who had done this before and was competent. On the other, an integration engineer whose entire job was connecting enterprise customers to a SaaS product. Between them, a SAML connection screen with — I counted, later, from a screenshot — forty-one fields.
They spent the ninety minutes on things like this. Is the ACS binding POST or Redirect? (POST. It is always POST. Redirect has a URL length limit that a signed assertion will exceed.) Should WantAssertionsSigned be checked, or WantResponseSigned, or both? (At least one, and if you don't know which, the answer is both.) What NameID format? (This one actually mattered, and nobody in the meeting knew it mattered.) Signature algorithm? Digest algorithm? Should we enable IdP-initiated? What's an Attribute Consuming Service Index?
At the end of the ninety minutes, SSO worked. And here is the thing worth sitting with: of the forty-one fields, three had been genuine decisions. Three. Everything else was either copied verbatim out of an XML metadata document that both systems already had, or set to the only value that has ever been correct, or left at a default that nobody in the room could have defended if asked.
The three real decisions were: which IdP attribute carries identity, what happens to a user who authenticates successfully but has no local account, and whether the customer's security team requires assertion encryption on top of TLS. Those are interesting questions. They have consequences. A reasonable person could answer them differently at two different companies.
The other thirty-eight fields were the SAML specification's table of contents rendered as a form.
That is the argument of this article, and I want to make it precisely rather than as a complaint about UI. SAML, SCIM, and audit are not inherently complex capabilities. They are made complex by platforms that take the protocol's vocabulary — its grant types, binding modes, assertion consumer URLs, PATCH path filters, event type enums — and expose it directly as configuration surface, on the reasonable-sounding theory that exposing everything is the safest way to support everyone.
It isn't. It transfers the design work the platform declined to do onto every customer, once per customer, forever.
Fields are not decisions
The useful distinction is between the number of fields a configuration surface has and the number of free variables it actually contains. They are wildly different numbers, and almost nobody measures the second one.
A field is a free variable only if both of these hold:
- Two real deployments legitimately differ on it. Not hypothetically — actually. If every working connection you have ever shipped has the same value, it is a constant wearing a text input.
- Its value cannot be derived from something the system already has. If the answer is sitting in a metadata document, a discovery endpoint, or another field on the same form, it is an output, not an input.
Run a SAML connection form through those two tests and it collapses. Here's the honest accounting from that forty-one-field screen:
| Category | Roughly how many | Example | Why it isn't a decision |
|---|---|---|---|
| Present in the IdP's metadata XML | ~16 | Entity ID, SSO URL, SLO URL, signing certificate, supported bindings, NameID formats offered | The IdP already published it, in a machine-readable document, signed |
| Protocol constants with one correct answer | ~11 | ACS binding, WantAssertionsSigned, signature algorithm, digest algorithm |
One value is correct; the others exist for 2005 |
| Determined by another field | ~7 | Encryption certificate, SLO binding, ACS index | Derivable once you know one upstream choice |
| Cosmetic or organizational | ~4 | Display name, description, icon | Real, but not protocol |
| Genuine decisions | 3 | Identity attribute, JIT provisioning behavior, encryption requirement | Two customers reasonably differ; nothing derives them |
The second row is worth dwelling on, because it's where the "we expose it for flexibility" defense is weakest. The HTTP-Redirect binding for an ACS endpoint is not a supported option that some customers prefer. It's a spec artifact that will break the moment an assertion carries a few group memberships, because the entire response has to fit in a URL. SHA-1 as a digest algorithm is not a preference. Offering these as equal-weight choices in a dropdown is not flexibility; it is a platform declining to state what it knows.
And the first row is the part that genuinely surprises engineers who haven't worked on federation directly: SAML already ships an intent-to-protocol compiler, and the industry mostly ignores it. SAML 2.0 metadata is a signed XML document — EntityDescriptor, containing IDPSSODescriptor with its KeyDescriptor use="signing" entries, its SingleSignOnService elements with their binding URIs, its advertised NameIDFormat list. It exists specifically so that two parties can exchange one artifact and derive everything mechanical from it. OIDC did the same thing more cleanly with /.well-known/openid-configuration and a JWKS URL.
The protocols solved this. The configuration surfaces did not adopt the solution, because a metadata upload field is one row in a table and forty-one fields look like a feature list.
The real cost isn't typing. It's the invariants nobody can express.
If the problem were only that people type sixteen values they could have uploaded, it would be a tedium problem, not an architecture problem. The architecture problem is sharper, and it's this:
Protocol fields are not independent, and a flat field surface cannot express the dependencies.
Every one of those forty-one inputs renders as its own control, validated on its own, saved on its own. But the correctness of a federation connection is a property of combinations. A few that have caused real incidents:
NameIDFormat: transientplus SCIM deprovisioning that matches on NameID. Each is defensible alone. Together, deprovisioning silently matches nothing, forever, and nobody notices until an audit asks why terminated employees still have active accounts. The form validated both fields successfully.WantAssertionsSigned: falseplusWantResponseSigned: false. Individually each is a checkbox. Together they describe a system that will accept an assertion authored by anyone who can reach the endpoint. (This one has its own article: Why Unsigned SAML Is Dangerous.)- IdP-initiated SSO enabled plus an unrestricted
RelayState. An open redirect with an authenticated session attached to it. Both fields are legal; the pair is a vulnerability. The trade-offs live in IdP-Initiated vs SP-Initiated SAML. - JIT provisioning on, with the identity attribute set to an email address, at a company that recycles email addresses after departure. Correct-looking, and it merges two humans into one account.
Forty-one fields, mostly binary or small enumerations, is a space with something on the order of 10¹⁰ distinct configurations in it. The number of correct configurations — ones that authenticate reliably, deprovision correctly, and don't accept forged assertions — is perhaps a few dozen. Any surface that hands you the raw product space is asking every customer to independently rediscover which few dozen points are safe, using the only feedback signal available, which is whether login appears to work.
Login appearing to work is not the property you cared about. Three of the four examples above pass that test.
I've made an adjacent argument about code branches in Configuration Beats Customization: n independent flags means 2ⁿ behavioral states nobody tests. This is the same arithmetic arriving at the configuration surface instead of the codebase, and it's worse in one specific way. A flag combination that breaks is at least the vendor's bug. A field combination that breaks is, contractually, the customer's configuration — which is why these incidents have such a distinctive flavor of nobody being wrong. See Who Owns the Incident When Every Layer Worked?.
Modeling intent means naming the valid points
The alternative isn't fewer capabilities. It is choosing a different set of primitives: instead of exposing the axes, name the points on them that are actually valid, and let people select among those.
Concretely, the three capabilities in the roadmap of every B2B product:
| Capability | What protocol vocabulary exposes | What the intent actually is |
|---|---|---|
| SAML / OIDC federation | Bindings, NameID formats, signature and digest algorithms, ACS index, AuthnContextClassRef, clock skew tolerance |
This organization's users sign in with their own IdP. This attribute is their identity. Unknown users are created / rejected. |
| SCIM provisioning | PATCH path grammar, externalId vs id, filter expressions, active: false vs DELETE, bulk semantics |
The customer's directory is authoritative for who exists and for these attributes. Departure means this. |
| Audit | Event type enums, retention tiers, export formats, log levels, sink configuration | Reconstruct who did what, on whose authority, in an order a regulator will accept. |
The right-hand column is short, and every entry in it is a sentence a CISO would actually say. That's the test I use for whether something is an intent: can the person who has the requirement state it, in their own words, without learning your protocol? If yes, that sentence is your configuration primitive. The protocol vocabulary is the compiler's intermediate representation, not the user's input language.
The SCIM row deserves an aside, because SCIM is the case where the "the protocol is just complex" defense is least available. SCIM is a small spec. Resources, a filter grammar, PATCH, /Users and /Groups. An engineer can hold it in their head in an afternoon. And yet SCIM integrations are reliably miserable, because the actual hard question is one the protocol never asks: for each attribute, which side is authoritative, and what does deprovisioning mean?
active: false and DELETE are both legal SCIM. One means "suspended, sessions should die, data retained." The other means "gone." Enterprises differ on which one their offboarding process sends, and products differ on how they interpret each. That mismatch — not PATCH path syntax — is where offboarding silently fails. A protocol-shaped config surface has a checkbox for "support DELETE." An intent-shaped one asks: when this user leaves, should their sessions terminate immediately, should their data be retained, and for how long? Those answers compile down to how you treat both verbs, plus your session revocation behavior, plus your retention policy, which the protocol surface would have scattered across three unrelated screens.
Audit is the same shape and I've written about its operational cost in The Hidden Cost of Audit Logs. The complexity there is never the number of event types. It is whether the record supports the one question auditors ask — who was this done on behalf of — which requires a stable subject identifier and an ordered chain of actors, a property you either designed in from the beginning or cannot retrofit at all.
The compile step, and the thing that breaks it
So the shape of the answer is a compiler: intent in, protocol configuration out.
flowchart LR
I["Intent<br/>policy · connection · lifecycle contract"] --> C["Compile<br/>validate cross-field invariants"]
C --> P["Protocol configuration<br/>grants · bindings · algorithms · TTLs"]
P --> R["Runtime"]
M["IdP metadata / OIDC discovery"] --> C
P -.->|"read-only, for debugging"| D["Effective config view"]
Two details in that diagram carry most of the weight.
The first is that the invariant checking lives in the compile step, where it can see all the inputs at once. This is the entire reason the indirection is worth anything. A flat field surface validates fields; a compiler validates configurations. "Transient NameID with directory-matched deprovisioning" is a rejectable program, not a runtime surprise.
The second is the dotted arrow, which I'll come back to in a moment, because it's where the honest counterargument gets its due.
And the failure mode to name explicitly: the compiled output must not be independently editable. This is the same law that governs every generator anyone has ever shipped — generated protobuf stubs, Terraform state, compiled CSS. The moment humans hand-edit the artifact, the source of truth forks and the generator becomes a liability rather than a tool. Identity platforms hit this constantly in a specific form: a quickstart wizard that produces a configuration, and a raw field editor that operates on the same object. Now every connection is in one of two states — generated, or diverged — and no one can tell which without reading it. That's not intent modeling with an escape hatch. That's two configuration systems sharing a database table.
If a platform offers both, the intent layer has to own the object, and the raw view has to be either read-only or a deliberate, marked, one-way conversion.
It's worth being clear that this is a different argument from the one against plugin systems in The Extensibility Trap. That one is about execution: customer code in the identity runtime destroys determinism and availability. This one is about expression: a configuration surface can be entirely declarative, entirely safe to execute, and still be badly modeled. The two problems have a shared root, though — a configuration language that can't express intent is precisely what drives customers to ask for a hook, because the hook is the only way left to say what they mean.
One illustration, disclosed
I work on ClavionX, so treat this as a worked example of the pattern rather than a recommendation, and note that the machine-identity pieces I'm describing are in design, not shipped — I'm describing a model, not production experience with it.
The case that made the abstraction earn its keep was agents. An autonomous agent needs an OAuth client. Registering one, protocol-first, means answering: which grant types, what token endpoint auth method, what access token TTL, what refresh token TTL, whether refresh rotation is on, which scopes, which audiences. Seven-plus fields, most of which an engineer registering a nightly reconciliation job has no principled basis for answering — and one of which, grant types, contains a genuinely dangerous option.
An agent is a machine. It has no browser and no user at the keyboard. AUTHORIZATION_CODE is never correct for it. But a generic client form offers it, because a generic client form is the OAuth spec rendered as a form, and someone will eventually tick it to make a stubborn integration work.
So the model puts an Agent in the data model as its own object — with an owner, a purpose, and a lifecycle (REGISTERED → ACTIVE → SUSPENDED → RETIRED) — and realizes it 1:1 as an ordinary OAuth2 client underneath. Nothing proprietary on the wire. You don't configure the agent's grants or TTLs; you assign an Agent Policy, a reusable rule set, and the policy compiles onto the client. Agent policies are constrained to machine-to-machine grants — CLIENT_CREDENTIALS, TOKEN_EXCHANGE, REFRESH_TOKEN — and there is no policy that can produce AUTHORIZATION_CODE. The dangerous point in the product space is not defaulted away from. It is not in the space.
MCP servers get the same treatment: an MCP Server is a Resource Server with kind=MCP, governed by an MCP Policy whose knobs are stated as intent rather than protocol. requiresDelegatedUser defaults to true, meaning the server refuses tokens that carry no human principal — the intent being this tool acts for people, never on its own account. discoveryVisibility chooses public, authenticated-only, or hidden. Two questions an architect can answer without reading a spec.
Underneath, delegation is plain RFC 8693 token exchange, with no agent-specific or MCP-specific casing in the runtime at all: sub remains the original user through every hop, and each intermediary accumulates in the act chain. That's the audit property from two sections ago, obtained structurally — the chain exists because the standard produces it, not because a logging layer remembered to write it down. The general version of this is in Auditing Agent Actions, and the framing of agents as clients and MCP servers as resource servers in Agent Authentication: The Client Analogy and MCP: The Resource Server Analogy.
Two costs, stated plainly because they're the interesting part:
Compilation is asynchronous, and that's visible. Changing a policy recompiles every governed object, and there is a window during which some are updated and some aren't. This falls out of a hard architectural rule — the runtime never synchronously calls the control plane, consuming projected state only and failing secure when it's missing — which buys the availability separation described in Identity Is Mostly Read Traffic and pays for it with a bounded staleness window. Any platform claiming instant global config propagation has hidden a synchronous dependency somewhere it will hurt later.
Policies are a shared blast radius. The upside of a reusable rule set is that one edit fixes a hundred agents. The downside is identical. Per-object configuration has the property that mistakes are contained, and abstractions like this trade that away for consistency. That's a real trade, not a free win, and it means policy edits deserve the review rigor of a code change rather than the rigor of a settings toggle.
When exposing the protocol is the right call
An argument that only says "hide the protocol" is wrong, and I'd rather draw the line than pretend it isn't there. There are four cases where raw protocol surface is correct, and three of them are common.
Debugging. Always debugging. When federation breaks — and it will, because the failure is usually in someone else's IdP — the person diagnosing it needs to see exactly what the system is doing at the wire level. InResponseTo doesn't match a request you issued. The assertion's NotOnOrAfter is in the past because a domain controller drifted four minutes. The IdP rotated a signing key and published metadata but didn't tell anyone. None of these are diagnosable from an intent view that says "SSO: enabled." Anyone who has lived through debugging SAML knows the intent abstraction is exactly the wrong altitude at 2am.
This is what the dotted arrow in the diagram is for, and it's the reconciliation I'd argue for generally: the protocol view should be a derived, read-only output, not an input. An effective-configuration view that shows every compiled value, where it came from, and which policy produced it gives integration engineers everything they need to diagnose, without turning the surface back into a write path where drift can start. Read surface and write surface are different problems and there is no reason to make them the same screen. Most platforms conflate them and lose both.
Genuinely non-conformant peers. Some percentage of enterprise IdPs — my experience says under 5%, but not zero — need something no sane policy would offer. An IdP that requires SHA-1 because a fifteen-year-old appliance can't do anything else. One that sends a malformed AudienceRestriction. One that needs NameQualifier populated in a way the spec permits but nothing else does. If the platform has no answer for these, the answer becomes a support escalation and a database patch, which is strictly worse than a marked override.
The design constraint on that escape hatch matters more than its existence. An override should be a first-class object with a reason, an owner, and a review date; it should be visible in the connection's diff and in the security review; and it should never be reachable by accident. The mechanism is the one from Configuration Beats Customization: capture the justification at the moment of creation, because that is the only moment the information required to eventually delete it exists. An override without a reason field is a permanent one.
Migration. Moving off an incumbent means reproducing its behavior exactly, quirks included, before you can normalize it. During that window you want the raw knobs, because "compile a correct configuration" is not the goal — "reproduce this specific one" is. That's a phase with an end date, not a permanent mode. Migrating Off an Identity Provider covers the rest of that.
When the protocol vocabulary is the user's vocabulary. This is the one people forget. If your users are identity engineers building on top of your platform — not application developers, not IT admins, but people whose job title contains the word "identity" — then grant types and binding modes are their native language, and translating into a friendlier intent vocabulary makes their work harder, not easier. Abstractions have an audience. Choosing the wrong one is a real failure mode in the other direction, and the reason a certain kind of low-level tool is beloved by exactly the people it was built for. The mistake is assuming that audience is everyone, which is roughly the argument in Enterprise Identity Is Not Consumer Authentication run in reverse.
Complexity should live in configuration, not in concepts
The pattern underneath all of this is a specific claim about where complexity is allowed to accumulate.
Every platform has irreducible complexity — federation genuinely is hard, deprovisioning genuinely has ambiguous semantics, audit genuinely requires care. That complexity doesn't disappear because you built a nicer form over it. The question is only who holds it: the platform, once, in a compiler and a set of policies it tests; or every customer, independently, in a ninety-minute call they'll repeat with the next vendor.
The version I'd defend is this. Concepts are what every user must learn regardless of size; configuration is what they select once they've learned them. A two-person startup and a bank should encounter the same objects — tenants, clients, resource servers, connections, policies — and differ only in how much of the configuration they touch. When a capability arrives as a new concept rather than as configuration of an existing one, the product has forked into a small version and an enterprise version that happen to share a login page. That fork is what "enterprise complexity" actually names. It isn't the number of features; it's the number of mental models. I've described the design decisions that follow from taking that seriously in The Identity Platform I Wanted to Exist, and the related discipline of refusing capabilities that belong to other layers entirely in Can We Make Identity Systems Simpler Again? — the surface stays small for two independent reasons, and it's worth not confusing them.
The forty-one-field screen is the visible symptom, and it's tempting to treat it as a UI problem. It isn't. It's a modeling decision leaking upward: the platform never decided which combinations were valid, so it exposed all of them and called it flexibility. Forty-one fields with three decisions in them isn't a form that needs redesigning. It's three decisions and thirty-eight admissions.
The work is deciding which three.