Every Extension Point Is a Permanent Public API

The ticket said: "Enrichment hook stopped populating department for tenant 4471 after 3.9.0."

It had been open for nine days when it reached me, which tells you something about how long it takes to believe a report like that. The hook interface hadn't changed. The signature was byte-identical. The customer's JAR was the same artifact, same checksum, same 2021 build date, sitting in the same directory. We had run their extension against a 3.9.0 build in a scratch environment and it worked.

The difference turned out to be four lines in our own code that nobody had thought worth mentioning in the release notes.

In 3.8, when a mapping expression referenced an attribute that didn't exist, the resolver dereferenced a null and threw NullPointerException. In 3.9 we did the obvious hygiene thing and introduced a typed AttributeResolutionException with a useful message, because the NPE stack traces were making support tickets unreadable. Strictly better. No public signature changed.

The customer's hook — written by a contractor who had left the company three years earlier, against an HR system that only two people understood — was structured like this:

try {
    dept = ctx.get("hr.department");
} catch (NullPointerException e) {
    dept = ctx.get("legacy.dept_code");   // pre-2019 users
}

They had discovered our NPE by accident, decided it was a reliable signal for "this user predates the HR migration," and built their fallback on it. It worked for three years. Then we fixed our bug, the NullPointerException stopped arriving, the AttributeResolutionException propagated out of their hook, and — because their tenant was configured fail-closed, which is the configuration we recommend — every pre-2019 user in that tenant could no longer log in.

We broke a customer by removing a defect. We never promised the NPE. We never documented it. It was never in the interface. It was, nonetheless, the contract.

The law, stated properly

Hyrum Wright's formulation is worth quoting exactly, because it is usually softened into something less useful:

With a sufficient number of users of an API, it does not matter what you promise in the contract: all observable behaviors of your system will be depended on by somebody.

Two words carry the weight. Observable — not documented, not declared, not intended. And all — not "the important ones," not "the ones a reasonable engineer would rely on." If it can be detected from outside, someone has already built on it.

The Extensibility Trap argues that identity platforms should not run customer code at all, and I agree with it; one of its section headings is the title of this article. This piece starts one step later, at the point where that argument has already lost — or has been deliberately overridden, because you decided that offering something was correct. You are going to ship an extension point. The question is what shape it should have so that shipping it doesn't pin your architecture for the next decade.

Why identity is the worst case for Hyrum's Law

Every API has this problem. Identity has it in a uniquely bad configuration, for three reasons that compound.

The dependent population cannot redeploy on your schedule. The standard mitigation for Hyrum's Law is a corpus you can test against: Google has a global build graph, Rust has crater runs against crates.io, a library maintainer can at least grep GitHub. You have none of that. Your customers' extensions live in private repositories, inside their build systems, often as binary artifacts whose source has been lost. You cannot enumerate your dependents, let alone compile them. The population is invisible by construction.

The code was written by someone who is gone. This is not a rhetorical flourish; it is the modal case. Identity integrations are written once, during an implementation project, frequently by a systems integrator whose engagement ended at go-live. There is no maintainer. There is a JAR and a wiki page that is four versions out of date. "Just update the plugin" is a request to an organization that has no one who can read it.

The extension runs in the authentication path, so breaking it is not a degradation. This is the one that changes the math. If you break a reporting plugin, a report is wrong. If you break an authentication hook, the failure mode is binary and it is the worst one available: fail closed and nobody in the tenant can log in, including the administrators who would fix it; fail open and you issue tokens missing an attribute that downstream systems use for access control, which is a silent privilege bug that nobody will notice for a quarter. The Extensibility Trap calls this the tell that the architecture has left you with no correct behavior. It is. But if you're shipping the extension point anyway, you need an answer better than picking one of the two wrong ones, and there is one — see the failure-semantics rule below.

Add these up and you get a specific consequence: your extension interface has a longer support horizon than your architecture. The plugin API you designed in year one, when you understood the problem least, will still be load-bearing in year ten, after you've replaced the storage layer, the token format, the deployment topology, and most of the team.

The observable surface is much bigger than the signature

Here is the part that experienced engineers still get wrong, because it isn't a knowledge gap so much as a habit of attention. When we talk about "the extension API," we look at the type signature. The type signature is a small and unrepresentative fraction of what customers can observe, and everything they can observe is what you've actually shipped.

Take the most boring possible hook:

Map<String, String> enrich(AuthContext ctx);

One method, two types. Now enumerate what a customer's code can detect about it at runtime:

Observable dimension What a customer can come to depend on The change that breaks them
Cardinality Called exactly once per login Adding a refresh path, a retry, or a second call for step-up doubles their ERP writes
Iteration order ctx iterates in registration or insertion order Swapping LinkedHashMap for HashMap; adding parallel resolution
Execution order Their hook runs after the group resolver, so groups are populated Reordering the pipeline; running resolvers concurrently
Mutability & aliasing Mutating the map they were handed persists into the token Passing a defensive copy — which is the correct thing to do
Object identity Same ctx instance across two calls, so they cached in a field Per-call allocation; object pooling; a different instance per node
Nullability ctx.get("x") returns null, never throws Introducing a typed exception, as above
Error type A specific exception class, message text, or error code string Rewording a message; wrapping in a cause chain
Absence vs. emptiness Missing attribute arrives as "", not absent Normalizing empties to null, or vice versa
Timing The hook has ~2s before anything upstream times out Tightening the login latency budget
Ambient context Runs on the request thread, so their ThreadLocal, MDC, or transaction is visible Moving to a worker pool or a virtual thread
Encoding A numeric claim serializes as 12345, not "12345" Fixing a JSON serializer inconsistency
Side effects of failure Throwing aborts the login, so they use it as a policy mechanism Making the hook non-fatal

Not one of those is in the signature. Every one of them is in the contract, whether or not you agreed to it, from the first day a customer ships against it.

The useful mental reframe: an extension point's contract is not what you declared, it is what a sufficiently determined observer can measure. Design review should ask "what can they measure?" not "what did we document?" — because the answers to those two questions diverge immediately and never reconverge.

The NPE story above is one cell of that table. I've since collected the others from real incidents, mine and other people's. My favorite is a platform that ran enrichment hooks sequentially, then parallelized them for latency; a customer had two hooks where the second read a key the first wrote into the shared context. Nothing in the interface said hooks ran sequentially. Nothing needed to. It was observable, so it was load-bearing.

The evolution story, in full

Let me walk one interface through four years, because the pattern is more instructive than the individual failures. Each change below is one a competent engineer would make without hesitation.

Year 1. Ship Map<String,String> enrich(AuthContext ctx). AuthContext is the live internal object — it's already there, wrapping it feels like ceremony. Hooks run sequentially in configuration order. Exceptions abort the login.

Year 2, change: make attribute resolution concurrent. Login p99 drops 80ms. Three tenants break because their hooks depended on ordering. You add a dependsOn field to the hook configuration to let them restore ordering. You have now made execution order an explicitly declared part of the contract, forever, in order to fix a problem caused by it having been an implicitly declared part of the contract.

Year 3, change: pass a defensive copy of the context. A support incident showed one tenant's hook mutating shared state that leaked across requests under load. Copying is unambiguously the right fix. It breaks every customer who was writing their output by mutating the context instead of returning a map — a usage you never intended and, it turns out, roughly a third of them adopted, because it was possible and it was convenient.

Year 3, change: add a riskScore field to AuthContext's serialized form. Two self-hosted customers had written strict deserializers. Unknown field, parse error, no logins. You did not add a method, break a signature, or change a type; you added information.

Year 4, change: invoke hooks on token refresh, not just on interactive login. Correct — attributes should be re-evaluated when a token is renewed. It quietly triples call volume for every hook that writes to an external system, and one customer's hook was posting an audit row per invocation into a database sized for daily logins.

Year 4, change: fix an NPE into a typed exception. You know how that one ends.

Six changes. Zero signature breaks. Four production incidents in customer environments, at least two of which locked users out. Note what all six have in common: each one was the removal of an accident. Hyrum's Law is not mostly about people depending on your features. It is about people depending on your accidents, and the accidents are exactly what you eventually want to clean up.

Manufacture the corpus you don't have

If the standard defense against Hyrum's Law is a global test corpus, and you can't have one, the move is to build a private one out of the traffic you already see.

Record every extension invocation: the serialized input payload, the returned value, the exception if any, the wall-clock duration, and the resulting decision. Sample it — you do not need all of it, you need coverage of shapes. Store it per-tenant with a retention window, and treat it as sensitive, because it is: these payloads contain identity attributes and belong under the same handling rules as your audit trail.

Then, before any change that touches the extension pipeline, replay the corpus against a candidate build and diff the results.

flowchart LR
    RT["Runtime<br/>hook invocations"] -->|"sampled: input, output,<br/>error, duration"| Corpus[("Invocation corpus<br/>per tenant, TTL'd")]
    Corpus --> Replay["Shadow replay<br/>candidate build"]
    Replay --> Diff["Diff: output, error class,<br/>call count, ordering"]
    Diff -->|"differences by tenant"| Gate["Release gate:<br/>named blast radius"]

What this buys you is not proof of compatibility — you're replaying against your own harness, not their JAR, and their code can branch on things your corpus never captured. What it buys you is a named blast radius. Instead of "we changed the exception type, hopefully nobody cared," you get "this changes observable behavior for 3 tenants, here they are, here's what changes." That converts an unbounded risk into a list of phone calls, which is a completely different kind of problem and one an organization can actually act on.

Two refinements make it much more valuable. First, diff the error class and message, not just the success path — that's where the breakage lives, and it's the field everyone forgets to capture. Second, record the call count per login, because cardinality changes are invisible in any per-call diff.

This is the same instinct as instrumenting customer-specific branches so you can eventually delete them, argued in Configuration Beats Customization: confidence, not capability, is the scarce resource. You are not trying to make change safe. You are trying to make change legible.

Designing the surface so it doesn't pin you

Everything above is diagnosis. Here is the constructive half — the rules I'd apply to an extension point I had to ship. Each one is a deliberate reduction of the observable surface, and each has a cost worth stating.

Pass immutable copies; accept return values; never hand out a live object. If they can mutate what you gave them, mutation is your contract and you can never introduce copying, pooling, sharing, or concurrency. Give them a frozen snapshot and require the result as a returned value — ideally as an explicit delta (set, remove) rather than a whole replacement map, because a delta lets you distinguish "didn't touch it" from "set it to the same value," which is the distinction you'll want in year three when you add a second extension to the same pipeline. Cost: allocation on a hot path, and an interface that feels more ceremonious than the obvious one.

Version the payload, and prove the ignore-unknown-fields behavior on day one. Structured payload with an explicit version field, additive-only within a major version, and a documented rule that consumers must ignore unknown fields. Documenting it is not enough — nobody's parser obeys documentation. So ship an unknown field immediately, in v1: a junk key whose name and value vary per release. Anything with a strict deserializer breaks on day one, in the customer's integration testing, when it costs an afternoon. Otherwise you discover it in year four, when it costs an outage. Deliberately exercising your own tolerance rules is the only way they stay true.

Never promise ordering you don't need — and actively destroy it. Declaring order unspecified achieves nothing; observable order is order. If hook execution order is genuinely not part of your contract, shuffle it, the way Go randomizes map iteration and gRPC clients jitter. Nobody can depend on what varies. The obvious objection is that nondeterminism makes support harder, and it does — so seed the shuffle from the request ID and log the seed, which makes any single execution exactly reproducible while making the aggregate undependable. You get the property you want without giving up debuggability. If some ordering is genuinely required, make it explicit and declared, and pay for it knowingly; what you must not do is have real ordering that you merely decline to mention.

Specify cardinality, then violate it in test. Decide up front whether the point is at-most-once, exactly-once, or at-least-once per authentication, write it down, and pass an idempotency key so a customer whose hook has side effects can deduplicate. Then make your staging environment invoke twice, sometimes, deliberately. At-least-once that is never actually more-than-once is at-least-once in the documentation and exactly-once in the contract.

Decide failure semantics at design time and make the degraded result visible. This is the answer to the fail-open/fail-closed dilemma, and it's better than choosing a side. Make the failure mode an explicit per-extension configuration with a small enumerated set — fail closed, omit and continue, or use last-known-good — chosen by the customer rather than defaulted by you. Then, critically, make the degraded outcome distinguishable in the output: an assertion issued without the enriched attribute because the extension failed must not be byte-identical to one where the attribute is legitimately absent. Carry a marker — a claim, an ACR-style indicator, an audit correlation — so a relying party's policy can decide for itself whether to trust a token built on partial enrichment. Fail-open is only a silent privilege bug while it is silent. Downstream authorization is a separate concern from authentication regardless; the seam is argued in Authentication Is Not Authorization.

Prefer declarative extension to imperative, for a mechanical reason. The usual case for declarative is safety and latency. The Hyrum's Law case is sharper: a declaration constrains the observable surface to what the evaluator deliberately exposes. A customer who writes a mapping expression cannot observe your object identities, your thread, your iteration order, your exception types, or your call count — not because they've promised not to, but because the vocabulary has no way to name those things. That's what lets you rewrite the evaluator, memoize it, reorder it, run it on a different node, or replace the whole engine. The constraint isn't protecting you from bad input; it's protecting your ability to keep building, which is the argument Freedom Is Conserved makes across both LLM tool schemas and identity.

But constrain the language itself, or you've just moved the problem. Expression languages grow their own Hyrum surface, and it is not small: implicit type coercion (the moment "1" == 1 evaluates truthy, that is frozen forever), short-circuit evaluation order when subexpressions have observable cost, null propagation semantics, regex engine backtracking behavior, numeric precision, collation in string comparison, and error message text that someone will inevitably parse. Design a declarative surface the way you'd design a type system: total functions, no implicit coercion, explicit null handling, no reflection-shaped escape into the host runtime. A small language with stated semantics is a genuine contract. A small language with emergent semantics is a plugin API with worse ergonomics.

The platform I work on, ClavionX, takes this route deliberately — disclosure: I'm affiliated, and the machine-identity work described here is in design, not shipped, so treat it as a description of intent rather than a maturity claim. Agents and MCP servers are configured through reusable Policy objects that compile asynchronously onto the underlying OAuth2 client and resource server, rather than customers configuring grant types and TTLs directly. That indirection is exactly the surface-reduction argument: the customer's contract is the policy vocabulary, and the compilation target underneath stays ours to change. The related structural rule — the runtime never synchronously calls the control plane, only consuming projected state and failing secure when it's missing — has a Hyrum's Law consequence people miss: there is no synchronous callback surface to accidentally expose, because there is no synchronous path at all. Architectural constraints and interface constraints are frequently the same constraint viewed from two directions.

The counterpoint: offering nothing has a bill too

An argument that only ever says no is a temperament, not an architecture, and the strongest objection to this whole article is that a sufficiently locked-down platform simply loses.

It's worse than losing deals. Refusing to offer an extension point does not remove the extension point; it relocates it to somewhere you cannot see. The customer who can't run a hook will put a reverse proxy in front of your token endpoint and rewrite claims in flight. They'll stand up a shim IdP that federates to you and decorates the assertion. They'll poll your admin API every thirty seconds and mutate user attributes out of band. Every one of those is now an unobservable, unsupported, un-versioned component in the authentication path — and when it breaks, the ticket still arrives at your support desk, minus any ability to diagnose it. A sanctioned extension point with a recorded invocation corpus is, on that comparison, a considerable improvement over a forbidden one.

There's a second cost that's easy to underweight: the platform that offers nothing learns nothing. Extension usage is the highest-signal product research available — it tells you precisely which needs your declarative surface failed to express. Teams that ship a constrained hook and instrument it discover the four recurring shapes that should become first-class features. Teams that ship nothing get told "we went with someone else" and never find out why.

So: escape hatches, deliberately, with a sunset. Which is a sentence everyone agrees with and almost nobody executes, so it's worth being concrete about what makes one actually terminate:

  • The migration target must exist when you grant the hatch. "We'll build declarative support for this later" is how a temporary hatch becomes permanent. If the replacement doesn't exist yet, you're not granting an escape hatch, you're granting an API.
  • A date, in the contract, with a named owner on both sides. Not a comment. Not a wiki page. The organizational mechanism has to outlive the individuals, because it will need to.
  • Telemetry from day one, per tenant. You cannot sunset what you cannot count, and "who's still using this" is unanswerable retroactively.
  • Don't document it. This is the uncomfortable one. A per-tenant flag granted to three named customers is retirable. The same capability in the public documentation acquires dependents you never negotiated with, including customers who adopted it after the sunset date was set. Discoverability and retirability are directly opposed, and you have to choose.

Even with all four, expect most sunsets to slip. Price the hatch as a recurring cost, the way The Hidden Cost of Customer-Specific Features argues for customer-specific code, and make sure whoever trades revenue against velocity is doing it with the number in front of them. And before granting one at all, it's worth running the request through the five questions in The Composability Test — particularly the last, "can we ever remove it?", which for extension points has a default answer of no.

What to take away

The version I'd want a principal engineer to carry into a design review:

The interface you froze is the one you shipped, not the one you documented. Everything observable is shipped. So the design question is not "what should this method take and return" but "what can a determined observer measure, and which of those measurements am I willing to still be honoring in ten years?"

Everything else follows from that. Copies rather than live objects, because aliasing is observable. Return values rather than mutation, because mutation is observable. Versioned payloads with tolerance you exercise yourself, because parser strictness is observable. Shuffled order where order is unspecified, because order is observable whether or not you name it. Failure semantics chosen at design time and marked in the output, because failure behavior is the most load-bearing observable of all in an auth path. Declarative over imperative, because a vocabulary is the only mechanism that makes a behavior genuinely unobservable rather than merely undocumented.

And then record what people actually do with it, because the one thing you can be certain of is that the contract in production is not the contract in the README, and the only version of it that will ever matter is theirs.