The Hidden Cost of Customer-Specific Features

The flag was called acme_nameid_legacy_upn.

It existed because in March 2019, four days before quarter close, a customer's IdP emitted a SAML NameID in a format that was not what their profile said it was — unspecified, containing a UPN, where the metadata promised emailAddress. Their identity team could not change it; the ADFS instance was owned by a subsidiary being divested and nobody was permitted to touch it. Our options were to lose the deal or add fourteen lines to the assertion parser.

We added the fourteen lines. It was the right call. The deal was worth more than the team's annual budget, the code was trivially reviewable, and the flag defaulted to off so it could not possibly affect anyone else.

Five years later there were 402 flags in the tenant configuration table, a release took six weeks, and nobody could delete acme_nameid_legacy_upn — including after ACME churned, in 2022, to a competitor.

That last part is the interesting part. Not the accumulation, which is well understood and which I've argued about in general terms in Configuration Beats Customization. The interesting part is that the flag survived the death of its only reason to exist by three years, and that no engineer on the team was being negligent. They tried. Twice. Both times they backed out, correctly, because they could not answer a question that turns out to be much harder than it looks: does anything still depend on this?

This article is about that question. Why it's hard, why the obvious instrumentation answers it wrong, and what a retirement discipline looks like that actually terminates.

How you get to 402

Nobody approves 402 flags. They approve one flag, 402 times, in a context where each approval is correct.

The context matters, so it's worth being specific about it. A customer-specific flag is almost never proposed by an engineer who wants one. It arrives through a channel that looks like this: an account executive has a signature blocked on a technical objection, a solutions architect has already told the customer "I'm sure we can handle that," and the ask lands in engineering with a date attached. The engineering answer that costs the least today is a narrow, defaulted-off branch. It is also, genuinely, the answer that carries the least risk today. An engineer who fights it is fighting for a diffuse future benefit against a concrete near-term loss, in a forum where they will be characterised as the obstacle.

So the flags accumulate in shapes. In identity work they're remarkably consistent, and I'd guess four families cover 90% of them:

Family Example Why it looked cheap
Protocol tolerance Accept a NameID format the metadata doesn't advertise; tolerate an unsigned <Assertion> inside a signed <Response>; skip InResponseTo correlation for an IdP that omits it One conditional in a parser, defaulted off
Timing and lifetime This tenant gets a 12-hour session; this one gets a 90-second clock-skew tolerance; this one wants refresh tokens that don't rotate One number, read from config
Token shape Add a non-standard claim; emit groups as a comma-joined string instead of an array; keep an old claim name alongside the new one A few lines in the token builder
Flow suppression Skip email verification because the customer provisions via SCIM; skip step-up for one client; allow IdP-initiated for one connection An early return

Every one of those is defensible in isolation, and several are things a reasonable platform ought to support. The problem is not the individual flag; it's what the population does that no member of it does.

The combinatorial part, and why nobody notices for years

The arithmetic is the obvious part and I won't belabour it: n independent booleans give you 2ⁿ behavioural states, and you are testing a few dozen of them.

The non-obvious part is why the pain arrives suddenly, years after the arithmetic went bad. The mechanism is that your effective test matrix is not the flag space — it's the set of flag combinations your live tenants actually occupy. Call that the occupied set. With 402 flags you might have 180 distinct occupied combinations, which is enormous but finite, and crucially it grows roughly linearly with customers rather than exponentially with flags. For a long time, a team can keep up: someone maintains a set of "customer-shaped" integration environments, the big accounts get a pre-release soak, and the escaped defects stay in single digits per release.

What changes is not the number of flags. It's the arrival of the first change that cuts across the occupied set.

You rewrite session storage. You upgrade the XML parser. You move token issuance behind a new signing service. Suddenly the unit of testing is no longer "one flag's behaviour" but "does this structural change hold under 180 combinations," and the 180 were never enumerated anywhere — they exist as rows in a production table owned, in practice, by whoever last ran an onboarding script. So the release process grows a phase that consists of finding out. That phase is what people call "hardening," and it is really discovery: exporting tenant configs, clustering them, building environments per cluster, running them, triaging what falls out, and then arguing about whether a given behavioural difference is a regression or a customer-specific feature nobody wrote down.

Six weeks is not the cost of testing 402 flags. It's the cost of rediscovering the occupied set every release because nobody maintains it as an artifact. That's a fixable problem, and it's much cheaper to fix than the flags themselves: a nightly job that computes the distinct configuration combinations across live tenants and writes them somewhere versioned turns a six-week discovery phase into a diff. Do that before you do anything else in this article — it costs a day and it is the only piece of this that pays back immediately.

The deletion problem

Now the hard part.

The reason 402 flags persists is not that removal is expensive. Deleting fourteen lines is free. It's that removal requires a proof of non-dependence, and the system has been quietly engineered — accidentally — to make that proof unobtainable. There are at least six distinct failure modes, and they compose.

1. There is no telemetry on flag reads. The most common state of the world. The flag is a column, or a JSON key, or a row in tenant_settings. Reading it produces no event. You can query which tenants have it set, which is not the same as which tenants depend on it, and the difference is where the incidents live. Everyone assumes the config table is the record. The config table is a record of what someone typed during onboarding, some of it in 2019, some of it copy-pasted.

2. Flags are read at config load, not at the decision. This one is the trap, and it's the reason the first instrumentation attempt usually makes things worse. Config is loaded once per tenant into an object — a TenantConfig, a projection, a cached struct — and the flag is a field on it. If you instrument the read, you instrument the load: every tenant that has a config looks like a user of every flag on it, because deserialization touched the field. Teams add the counter, look at the dashboard, see all 900 tenants "using" the flag, and conclude the flag is load-bearing everywhere. It isn't. It was deserialized.

flowchart TB
    subgraph L["Instrumenting the load — measures nothing"]
      C1[("tenant_settings")] --> D1["Deserialize TenantConfig"] --> M1["counter: flag read<br/>fires for every tenant, every reload"]
    end
    subgraph U["Instrumenting the decision — measures dependence"]
      D2["TenantConfig in memory"] --> B["NameID resolution<br/>compare flag path vs default path"]
      B -->|"same result"| N["no signal — not a dependency"]
      B -->|"different result"| S["counter: divergence<br/>tenant, flag, call site"]
    end

3. The signal you want is divergence, not usage. This is the piece I most wish someone had told me earlier, because it changes the shape of the whole problem. A flag being read is not evidence that anything depends on it. A flag is a dependency only when it caused a different outcome than the default would have produced. So the thing to emit is not "flag X was read by tenant T" but "flag X changed the answer for tenant T at call site S."

Concretely: at the branch, compute what the default path would have decided, compare, and emit only on divergence. In the NameID case that's cheap — you already have the assertion; resolving it both ways costs microseconds. Where the default path has side effects or real cost, you can't run it, but you can usually characterise it statically enough to record the comparison (flag=on, format=unspecified, default_would_have=reject).

The payoff is large. Divergence counters are sparse — most flags turn out to change nothing for most of the tenants that have them set, because the flag guards a condition that no longer occurs. acme_nameid_legacy_upn was set on eleven tenants. It diverged on zero, because the other ten had it copied in from a template and none of their IdPs emitted unspecified. Usage telemetry said eleven. Divergence telemetry said zero. Only one of those numbers is a deletion decision.

4. The default drifted. The subtle one. A flag written in 2019 as "opt in to the lenient behaviour" is often, by 2023, inverted in practice: the lenient behaviour got promoted to the default because eight more customers needed it, and the flag now means something narrower, or nothing, or — worst — the opposite of what its name says. Now removing the flag doesn't restore 2019 behaviour; it changes behaviour for every tenant that never set it, because deleting the branch also deletes the fallback the branch was written against. Any flag older than two refactors of the code around it needs its semantics re-derived from the code, not from its name or its ticket.

5. The owning customer is gone, and the flag isn't. Churn does not clean up. ACME left in 2022; the flag stayed because (a) their tenant row was soft-deleted rather than removed, so a query for "tenants with this flag" still returned it, (b) the onboarding template that a solutions architect had cloned from ACME's config in 2020 carried the flag into every subsequent SAML tenant, and (c) nobody connected the churn event to an engineering cleanup task, because those systems have never spoken to each other anywhere I've worked. The general form: the lifecycle of a customer-specific feature is not coupled to the lifecycle of the customer. If you build one coupling from this article, build that one — a churn or downgrade event should open a ticket enumerating the flags that customer was the stated owner of.

6. Flags entangle. Flag A is only correct when flag B is set, because the engineer who added A in 2021 tested it in an environment that happened to have B on. Nobody recorded the dependency; it lives in the interaction, and it means the safe-deletion question is not per-flag but per-subset. This is the one that turns a tractable cleanup into an intractable one, and it's the strongest argument for divergence telemetry recording the call site as well as the flag — entanglement shows up as two flags diverging at the same site for the same tenants.

Notice what all six have in common: they're not knowledge problems that harder thinking solves. They're observability problems. Which is good news, because observability problems have engineering answers, and the answer is not a spreadsheet audit.

A retirement discipline that terminates

Four practices, in the order that gives you leverage soonest. None of them require permission from anyone in sales, which is the point.

Record the deletion criterion at creation time. Not "review by Q3." A falsifiable condition: "Delete when no tenant diverges on this flag for 90 consecutive days," or "Delete when ACME's ADFS migration completes (their ticket IDP-4471)," or "Delete when we ship configurable NameID mapping." Alongside it: the owning tenant, the requesting account, the engineer, and one sentence on what breaks without it. This is the same argument as the reversibility point in Configuration Beats Customization, but sharpened — a review date only schedules the moment when someone will fail to answer the question. A deletion criterion specifies what evidence would answer it. Put it in a structured field next to the flag definition, not in a comment, so a job can evaluate it.

Instrument divergence in production, at the decision site. Covered above. Two operational notes. First, the observation window must exceed the flag's natural period — and identity flags have long periods. A flag on certificate-rotation handling might diverge once a year; a flag on a quarterly access-review import diverges four times a year; a flag on the SCIM deprovisioning path diverges when someone leaves the customer's company, which for a 40-person tenant might be twice. Ninety days of silence is meaningless for a flag whose event is annual. Set the window from the event frequency, not from a policy. Second, emit tenant, flag, call site, and the values compared — not just a counter — because the first thing you'll want when a divergence appears after 200 quiet days is to know exactly which condition fired.

Deprecate by darkening. When telemetry says a flag is dead, you still have a residual risk: the divergence might be real but rare, or the instrumentation might have a gap. Darkening resolves it empirically, and it's the inverse of a dark launch — instead of running new code without effect, you remove old behaviour for one tenant and watch.

The protocol matters more than the idea:

  • Pick the lowest-value tenant that has the flag set, not the loudest one.
  • Turn the flag off for that tenant only, during a window when their traffic is live but their business hours are not peak, and where you can flip it back in under a minute — a config change, not a deploy.
  • Watch for the absence of success, not the presence of errors. This is the part teams get wrong. A flag suppressing email verification, turned off, produces no errors at all: it produces users stuck at a verification screen, which surfaces as a drop in completed logins, not as a 500. Define the success metric for the affected path before you darken, and alert on its rate rather than on exceptions.
  • Then widen: one tenant, three tenants, all but the largest, all. Each step long enough to cover the event period.
  • Only after a full period at zero divergence, delete the code. Removing the config key comes last, in a separate change, so that the rollback during the risky phase is a config flip rather than a deploy.

Darkening is slower than deleting and enormously faster than the alternative, which is not deleting. It also converts the political problem into a technical one: you are no longer asking anyone to approve a deletion on faith, you are producing evidence.

Promote what has become product. A meaningful fraction of the 402 will turn out to be flags that many tenants diverge on. That flag is not a customization any more; it's an undocumented feature with a bad name and no tests. The correct response is not deletion and not preservation — it's promotion: give it a name in the configuration vocabulary, define its valid values, pick and document a default, write the tests, cover it in the docs, and migrate every tenant onto the real setting. You have converted an unenumerated dimension into an enumerated one, which is exactly the move Configuration Beats Customization argues for at design time, applied retroactively. In practice this is where most of the value is: on the cleanups I've watched, the split lands around 60% deletable, 25% promotable, 15% genuinely still customer-specific. It's the 25% that makes the exercise worth funding, because promoting one flag often lets you delete six others that were approximating the same thing badly.

One thing to be careful about: promotion is not free, and a promoted flag is a permanent public commitment. Promote when several tenants diverge for the same reason. If five tenants diverge for five different reasons, you have five leftovers, not one axis.

When a customer-specific path is the right call

An article that only says "don't" isn't an argument. Sometimes the branch is correct, and pretending otherwise is how architects stop being invited to the deal calls.

The honest cases:

The requirement is genuinely singular and genuinely revenue-critical. One customer, one weird protocol behaviour, a large contract, and no prospect of the requirement generalizing. Take the money. The mistake is not taking it — it's taking it without containment.

The divergence lives at the boundary, not in the core. This is the containment move worth internalizing, and it's the identity-specific version of the general "adapter at the edge" pattern. acme_nameid_legacy_upn did not have to be a branch in the assertion parser. It could have been a per-connection normalization step that runs before the parser and rewrites the NameID into the format the metadata promised — a declared mapping on the connection, not a conditional in shared code. The core stays single-path; the weirdness is data attached to one connection object, visible in one place, deletable by deleting the connection. Same for token shape: a per-client claim mapping is an adapter; an if (client == X) in the token builder is a branch. Ask, every time: can this be expressed as data on the edge object, rather than a condition in the middle? Often it can, and then it retires itself when the customer does.

The requirement contradicts a design premise. Not extends — contradicts. Then it isn't a flag, it's a different product, and the containment is a separate deployment with its own release train, priced accordingly. This is also the honest answer when a customer wants behaviour you would not want to defend in a security review, and it's adjacent to why running customer code inside the runtime is a different and worse trade — see The Extensibility Trap.

A paid fork with a sunset date. Rare and legitimate: a fixed-scope contract, a maintained branch, a stated end date, and a price that reflects a second release train. The failure mode is a fork with no sunset date, which is just a second product you didn't decide to build. If you take this route, put the sunset in the contract, not in a wiki page.

Across all four, the containment question is the same: when this customer leaves, what deletes the code? If the answer is "a person who remembers," you haven't contained it. If the answer is "the connection object," or "the separate deployment," or "the contract expiry," you have.

The line

Everything above reduces to one asymmetry, which is worth stating in the form you can use in a design review:

A customer-specific feature is created with a customer's name on it and lives without one. The name is in the ticket, the ticket gets archived, the customer churns, the template propagates the flag to eleven tenants that never needed it, and what remains is a branch that nobody can prove is dead — not because it's load-bearing, but because you never built the instrument that could tell you.

You cannot prevent the flags. The economics that produce them are real, and mostly the deals are worth it. What you can do is refuse to create one that has no deletion criterion, measure divergence rather than usage, darken before you delete, and promote the ones that stopped being customer-specific years ago.

Four hundred and two flags is not a discipline failure. It's four hundred and two decisions made without the evidence that would have allowed the sixth one to be reversed.