Keep Business Logic Out of Identity

The ticket was titled "remove cost_center from the access token." One line of claim mapping. The estimate was two points.

It took eleven months.

Not because the mapping was hard to delete — because nobody could say what would break. The claim had been added in 2019 for a finance dashboard that had since been decommissioned. But cost_center was in every token the platform issued, which meant it was visible to every application in the estate, which meant that at some point somebody had started using it. Two teams admitted to it immediately. A third denied it, then found a switch statement on the value in a reporting service. A fourth was a vendor SaaS product with a JSONPath expression in a customer-managed configuration screen, and the person who wrote it had left.

The last one was the interesting one. A procurement tool used cost_center to decide which approval chain a purchase request entered. Not "displayed it" — routed on it. That was a real access decision, made in a real system, on a string identity had been passing along for six years as a convenience. Deleting the claim would have silently routed every request to the default chain, which had no approver in it.

Somewhere in month seven, an architect on that project said the thing this article is about: we never decided this attribute belonged to us. We just had a place to put it.

That's the failure mode. Attributes aren't adopted into an identity platform by a design decision. They're adopted because there was a table with a row for every person and a token that reached every application, and the marginal cost of one more field was approximately zero — right up until it was eleven months.

The Extensibility Trap ends by arguing that the resolution for most of these requests is to pre-compute the attribute instead of fetching it during login. That's correct and it's about three sentences of the actual work. This article is the sequel: how to tell which attributes belong to you in the first place, and the mechanical migration for moving one from fetched-at-login to synced-ahead-of-time without losing the capability that justified it.

"Is it about the user" is not a test

Everything in an enterprise is about a user. Salary, laptop asset tag, Slack handle, desk number, parking permit, dietary preference for the offsite — all about a user, all arriving in the SCIM payload, all landing in the identity store because SCIM is an identity protocol and that is where SCIM writes. Identity Isn't Your User Database makes the case that most of these fields shouldn't be there, with the useful sorting question: would this field ever participate in an access decision?

That question is right, and it under-determines the answer. Plenty of fields do participate in an access decision somewhere and still don't belong in identity — the procurement routing above is exactly that case. So here is a longer test: three questions, applied to a specific named attribute rather than to a category.

1. Does an access decision depend on it? Not "could it inform one" — does a permit/deny, or a token audience, or an authentication requirement, change value when this attribute changes value? If nothing gates on it, you are storing and distributing a fact for someone else's convenience, and the identity platform is functioning as a CDN for HR data.

2. Does identity have authority to change it? Can an administrator in the identity platform set this value and have that be correct — not overwritten on the next sync, not divergent from a system that considers itself the owner? If the answer is no, identity holds a replica, not a master, and a replica has obligations a master doesn't.

3. Does it change on a timescale slower than a session? Compare the attribute's natural rate of change against your session and token lifetimes. If it changes faster than a token lives, a claim carrying it is a photograph of a value that has already moved, and every consumer of that claim is deciding on stale data by construction.

Each question routes differently, which is what makes this usable. It isn't a vote where two out of three wins.

Answer What it means Where the attribute goes
No to Q1 Nothing gates on it Application store, keyed on sub. Not identity, not the token — regardless of how convenient it is
No to Q2 Someone else owns the value Identity may hold it as a read-only replica with a named owner and a staleness budget. Never editable on both sides
No to Q3 It moves faster than a token Not a claim. Checked at the point of use, by the system that owns the decision
Yes to all three Identity fact Store it, project it, claim it

This is not the test in Identity Isn't Your Authorization Engine, which asks where a decision should be evaluated. That one places logic; this one places facts. You can pass one and fail the other: "which regions may this user operate in" is a fine identity fact and a terrible identity decision if the rule needs to know what the resource is. The two compose; neither substitutes for the other.

Working the borderline cases

The test is only worth anything if it produces non-obvious answers, so here are the cases that actually show up in enterprise deals, worked honestly.

Cost center. Q1 is usually no, which surprises people, because cost center feels load-bearing. Chargeback is accounting, not access. When it does gate something — the procurement routing — the gate lives in a domain system that has its own copy of the org structure and should read it from there, not from a token minted by the login service. Q2 is emphatically no: SAP owns it and wins any disagreement. So: not an identity attribute, and the reason the industry keeps putting it in tokens is that the request arrives phrased as "just add one claim" rather than "please make identity a distribution channel for finance's org chart."

Employment status. Yes, yes, yes — but only after you narrow it. HR's employment status is a rich enum: active, on leave, garden leave, notice period, secondment, contractor-via-agency, rehire pending. Identity does not want that enum, because every value in it is a domain concept identity would then have to understand, and keep understanding as HR redefines it. Identity wants the derived fact: may this principal authenticate, and at what assurance — active, suspended, terminated. Mapping HR's enum onto that triple is business logic; it belongs on the HR side or in the sync, never in claim mapping. And Q2 is yes here even though HR is upstream: a security administrator must be able to suspend an account now, ahead of HR, during an incident. That unilateral power is the tell that identity owns the derived fact rather than mirroring the source one.

Licence entitlement. Splits in a way that's easy to miss. Seat assignment — is Alice licensed for the analytics module — is slow, low-cardinality, and gates access: a good identity attribute, provided somebody owns the assignment and the sync. Seat count — are we within our contracted 500 — is an aggregate over a population, not a fact about Alice, and identity has no business computing it at issuance. When a contract requires the count to be authoritative at the moment of grant, that's a genuinely live check, and it belongs in the containment pattern at the end of this article rather than in a claim.

Contract or plan tier. Yes, no, yes. It gates feature access, the billing system owns it, and it changes at contract speed. A textbook replica: hold it, project it, put it in the token, and make it read-only in the identity admin UI so nobody can grant a customer the enterprise tier by editing the wrong screen.

Per-user feature flags. This one fails hardest and gets requested most, usually by a platform team who notice that the token already reaches every application. Q2 is no — the flag system owns flags. Q3 is emphatically no, and the reason is worth stating in mechanism rather than principle: a feature flag in a token has a rollback latency equal to your token lifetime. The entire operational value of a kill switch is that it takes effect in seconds. Ship the flag as a claim and the kill switch now takes fifteen minutes, or however long the longest live access token has left, and it takes that long during the incident where you're using it. Nobody discovers this in design review. They discover it at 03:00 while watching a graph refuse to bend.

Manager. Almost always no on Q1 — it's requested for approval workflows, org charts, and out-of-office routing, all domain features. The case where it genuinely gates access ("a manager may view their reports' timesheets") needs the reporting graph, not a single manager string, and a graph in a token is the cardinality mistake described in the authorization-engine piece.

The pattern: the attributes that fail aren't the ones that look like junk. They're the ones that feel identity-adjacent, arrive through an identity protocol, and have a plausible one-line access story attached. Which is exactly why a written test beats intuition.

A replica needs an owner, or you get two writers

Question 2 does more work than it looks, and the mechanical consequence is what most teams skip.

If identity holds a replica of an attribute owned elsewhere and the admin console lets someone edit that field, you have two writers to one value. Nobody designed that either; it happens because the admin UI renders every schema attribute as an editable field and the sync is a separate subsystem by a separate team.

The failure isn't a conflict error. It's silent and delayed: an administrator fixes a wrong department during a support call, everyone confirms the fix worked, and the nightly sync reverts it. Access changes back at 02:00. The ticket is reopened by someone who now believes the platform is haunted, and the audit trail shows two legitimate changes by two legitimate principals with no indication that either overwrote the other.

Three rules make replicas safe, and they cost almost nothing at the time you build the sync:

  • Mark replicated attributes read-only in every write surface — admin console, admin API, self-service profile. If a human can type into it, the sync will eventually fight them.
  • Record the source and the source event time on every replicated value. Not the time you wrote the row: the time the authoritative system recorded the change. That's the field the entire staleness argument below depends on, and it cannot be reconstructed later.
  • Define the break-glass path explicitly. There will be an incident where a value must be overridden ahead of the source. Make that an audited, expiring override the sync knows about and refuses to clobber — not an admin quietly editing a field.

The staleness budget is a number, and somebody has to say it

Here is the framing that makes the rest of the migration tractable.

"Is this attribute fresh enough?" is unanswerable. It has no units, it invites opinions, and in every meeting I've watched it resolves to whoever is most senior saying "it should be pretty much real time" — which is not a requirement, it's an anxiety.

"This attribute must be correct within fifteen minutes" is a design input. It sizes the sync interval, tells you what to alert on, and tells you whether a claim is viable at all. It's negotiable, because it has a cost on both sides, and it can be agreed — which means it can also be violated, detected, and escalated. So for every attribute you replicate, negotiate a budget with the consumers and write it down next to the attribute.

Two pieces of arithmetic make these conversations concrete.

First: total staleness is not sync lag. When a consumer reads an attribute from a token, the value they see was true at the moment the source system changed it, delayed by the sync pipeline, then frozen for as long as the token lives. The number that matters is:

effective staleness ≈ source-to-projection lag + token lifetime

A fifteen-minute budget with a fifteen-minute access token leaves roughly zero budget for the sync — another way of saying the attribute cannot be a claim. This is the most common error in these designs: teams size the sync against the budget and forget the token is a second, usually larger, staleness term stacked on top. If the budget is tight relative to token lifetime, your options are a shorter token (now paid for at the token endpoint on every refresh; see Identity Is Mostly Read Traffic) or moving the attribute out of the token into a lookup at the point of use.

Second: the binding budget is the minimum across consumers, not the average. If four systems read employment_status and three tolerate a day while the fourth must revoke within five minutes, your budget is five minutes — or you split the attribute into two delivery paths with different guarantees and say so out loud. Discovering that fourth consumer after building a nightly sync is the standard way this project goes wrong.

The resulting table belongs in the same repository as the sync:

Attribute Source of authority Budget Who agreed it Enforcement
employment_status Workday 5 min Security + HRIS Event-driven; alert at 3 min lag; deny on missing
licence.analytics Entitlement service 1 h Product ops 15-min scheduled sync; omit claim on staleness breach
plan_tier Billing 24 h Finance + platform Nightly; stale value served, staleness surfaced on dashboard
department Workday 24 h Nobody would commit Not an access input — display only, no gate permitted

That last row is the most useful one. When nobody will commit to a number, nothing important depends on the attribute — you've just discovered it fails question one.

The migration: fetched-at-login to synced-ahead-of-time

The architectural argument is easy; the removal is not. Seven steps, in order, with the failure mode of skipping each.

1. Inventory the readers, and what each does with the value

Not "who reads this claim" — what does each reader do when the value changes, and what would it do if the value were absent. Display, filter, and gate are three different answers and only the third constrains you.

The hard part is that you cannot grep for consumers of a JWT claim. It's in a token you hand to other people's software, including software you don't own. Three techniques beat asking:

  • Instrument the consumers you can reach, and accept the result as a lower bound.
  • Emit a canary value. For a small, consenting population, set the claim to something obviously distinctive and see who breaks. Crude, effective, best run against a pre-production tenant wired to real integrations.
  • Announce a dated removal and watch the ticket queue. The teams that surface in the two weeks after a deprecation notice are your real consumer list. In the story above, three of five readers were found this way and one was never found at all.

Skip this step and you get the eleven-month ticket.

2. Get a staleness budget from each reader

One number per consumer; take the minimum. If a consumer can't produce a number, push: what is the worst thing that happens if this value is an hour old? "Nothing" means write down a day and move on. "Someone keeps access after termination" means you have both a budget and a much more important conversation.

3. Establish the source of authority and stop the second writer

Name one system as writer of record, then go find the others — there is always at least one: the admin console field, the quarterly CSV import, the support tool that "fixes" values. Make the attribute read-only everywhere except the source.

This step frequently reveals that the source of authority doesn't exist yet — the value only ever existed as the output of a function the callout happened to call. That's a real project, and better found now than at cutover.

4. Build the sync, and make lag a first-class metric

SCIM push, scheduled pull, or an event subscription; the choice mostly depends on what the source system can do, and each carries an operational cost worth pricing first (the cost of SCIM synchronization). What matters architecturally is invariant across all three:

  • Every projected record carries the source event timestamp, not the write timestamp.
  • Staleness is exported in seconds, per tenant, per attribute. "40,000 messages behind" means nothing to anyone; "eleven minutes stale for tenant 44" means something to your on-call engineer and to your auditor.
  • The alert threshold sits below the negotiated budget, so a breach is a page rather than a postmortem finding.

5. Shadow-compare — the step everyone skips

Run both paths simultaneously. Keep the login-time fetch as the authoritative value, compute the synced value alongside it, and diff them on every issuance.

flowchart LR
    L["Token issuance"] --> F["Live fetch<br/>(still authoritative)"]
    L --> S["Synced projection<br/>(shadow)"]
    F --> T["Token: cost_center"]
    F --> D{"Differ?"}
    S --> D
    D -->|"yes"| R["Divergence record<br/>class, tenant, user, both values"]
    D -->|"no"| N["Counter++"]

This is the phase that finds the surprises, and the one that gets cut when the project runs late. What it turns up, reliably:

  • Normalization drift. The callout returned "FIN-100"; the sync produces "fin_100". Both are correct. One of your consumers does an exact string comparison.
  • Population gaps. The sync covers employees. The callout covered anyone the HR API would answer for — contractors, service accounts with a human sponsor, and forty people in a recently acquired subsidiary sitting in a different HR instance. The most common finding, and the one most likely to fail open at cutover.
  • The implicit fallback. The old callout had a timeout, and on timeout returned a default: "UNKNOWN", the empty string, or last year's value from a local cache. Consumers have been treating that default as a real value for years, sometimes with a special case for it. Your sync has no such behaviour, so the distribution of values shifts at cutover in a way nobody predicted.
  • Effective dating. HR records a transfer on the 3rd, effective the 20th. The callout queried "current" and got the old value; the sync ingested the change record and applied it immediately, so a few dozen people change department eighteen days early. Understanding effectivity is business logic — which is precisely why it belongs in the sync rather than in claim mapping.

Run the comparison for a full business cycle, not a week: payroll, month-end, the quarterly access review, and the onboarding wave each produce distinct divergence classes. A clean week proves you sampled a quiet week. The exit criterion is that every divergence class is explained and either fixed or accepted in writing — not "the diff rate is under 1%," because 1% of a large directory is a lot of people and they will not be randomly distributed.

6. Decide the missing-and-stale behaviour, and make it fail secure

Two distinct cases, and they deserve distinct answers:

Missing. No projection exists for this user. Newly provisioned, sync never ran, tenant onboarding incomplete. The answer is deny, not fetch — the synchronous fallback read is exactly the dependency you are removing, and it will be exercised for the first time during the incident where hammering the source system is the worst available move. This is the same rule as what an identity platform should do when its database is down: a missing projection means denied, logged, and visible.

Stale beyond budget. A projection exists but its source event timestamp is older than the negotiated number. Here the budget earns its keep: you already agreed what the value is worth, so you can decide between serving it with staleness surfaced, omitting the claim, or refusing the scope.

One non-obvious requirement spans both: omitting a claim is only fail-secure if the consumer denies on absence. Most don't — most treat an absent claim as "not applicable" and proceed. If your fail-secure design is "we stop emitting the claim," you have built a fail-open system with a fail-secure name on it. Either the consumer denies on absence and you have tested that it does, or issuance itself refuses the affected scope so the token can't be used for the gated operation at all.

The first-login race is guaranteed and needs a designed answer rather than a discovered one: someone authenticates at 08:58, their projection lands at 09:03. Deny-and-retry is usually right. Whatever you pick, pick it now — you will not enjoy improvising it on a new hire's first morning in front of their manager.

7. Cut over, then actually delete the callout

Flip the source, watch the divergence counter, and hold for one more business cycle with the old path disabled but present.

Then delete it. Not commented out, not behind a feature flag, not retained as "an emergency fallback" — delete the code, the credentials, the network policy, the firewall rule. A dormant synchronous callout with a flag in front of it gets re-enabled during an incident by someone who reads the flag name and reasonably concludes it's a safety net, and the entire availability argument returns in one deploy. Leave a dated comment in the config saying the attribute is synced and why; that's the artifact that stops this being re-litigated in two years.

This migration removes a dependency from the login path, which is the topology argument in Identity Is Not Your Integration Layer and I won't re-run its arithmetic. It also does something that argument doesn't cover: it forces a written answer to who owns this fact — the question that made the attribute unremovable in the first place.

The honest counterpoint: attributes that must be live

Some attributes genuinely cannot be pre-computed, and an architect who claims otherwise loses the room.

Values whose whole content is "right now." Risk and fraud signals computed from this device, this IP, this velocity. A pre-computed risk score from four hours ago answers a different question.

Contractually authoritative checks at the instant of grant. Some entitlement models require the seat count to be correct when the token is minted, because an audit clause says so. "Eventually consistent" isn't always defensible to a vendor's auditor.

Revocations with a near-zero budget. A break-glass suspension is meant to take effect immediately. If your arithmetic says the value can be fifteen minutes stale, the suspension is fifteen minutes stale, however elegant the pipeline.

The containment pattern has one move people consistently miss, and it isn't about timeouts: if an attribute must be live, that is strong evidence it should not be an attribute at all. A value that must be correct at the moment of use should be evaluated at the point of use, by the system that owns the decision — where the failure is scoped to one feature, the fallback is a domain decision rather than a login failure, and the dependency runs between two systems that already know about each other. Moving a live fetch into the login path doesn't make it fresher; it makes its failure everyone's failure. The procurement tool at the top of this article should have been calling finance's org-structure API, where a timeout means "approval routing is degraded" rather than "nobody in the company can log in."

For the residue that genuinely must be synchronous at issuance, the containment rules are the ones in the integration-layer piece — tight timeout set backwards from the login budget, a breaker that opens fast, a pre-agreed decision rather than an error when it trips. One attribute-specific rule to add: scope the live check to the audience that needs it. An authoritative licence check should run when minting a token for the licensed application, not on every authentication in the estate. Most "must be synchronous" requirements shrink by an order of magnitude once you ask which token requests they actually apply to.

And one asymmetry worth keeping: a stale value is safe when the attribute is restrictive and dangerous when it is permissive. Last-known-good on a deny-list entry fails in the safe direction; last-known-good on an entitlement grant is a privilege bug with a timestamp on it. Same mechanism, opposite conclusions, decided per attribute.

Where the line falls

The boundary is short enough to keep in your head. An identity platform answers who are you, and what may you access. It does not answer what is this person's cost center according to SAP. The second question is legitimate and important, and belongs to a system that isn't in the authentication path.

The three questions tell you which one you're being asked: does an access decision depend on it, do you have authority to change it, does it move more slowly than a session. The non-obvious answers fall out immediately — employment status is yours but HR's enum isn't, licence assignment is yours but seat count isn't, and a feature flag in a token is a kill switch with a fifteen-minute delay welded to it.

The migration itself is unglamorous bookkeeping: inventory, budgets, authority, sync, shadow-compare, fail-secure, delete. Shadow-compare is what separates teams who ship this from teams who roll it back, because it's the only phase that tells you what the old path was actually doing rather than what its documentation claimed.

And the artifact that survives all of it is the smallest one — a table with an attribute name, an owner, and a number of minutes. Every argument here can be re-derived from that table. Almost nobody has one.


Disclosure: I work on ClavionX, which takes the hard-line version of this — no customer code executes in the identity runtime, so there is no login-time callout to write in the first place. The sanctioned answers are declarative claim mapping, a policy engine, async events off the critical path, and pre-computing attributes via SCIM or a scheduled sync; the runtime never synchronously calls the control plane during request execution, and a missing projection fails secure rather than reaching back for a fresh read. Those constraints are what forced us to work the migration above out in detail, because "just fetch it at login" isn't something we can offer. It's a real constraint with real costs, and the reasoning stands whether you run Keycloak, Entra, Auth0, or something you built.