Identity Isn't Your User Database
This is the most violated boundary in the whole category, and it's violated for the most sympathetic reason: the identity platform already has a table with a row for every person, and you need somewhere to put a field.
This is the most violated boundary in the whole category, and it's violated for the most sympathetic reason: the identity platform already has a table with a row for every person, and you need somewhere to put a field.
So the field goes there. Then another. Then a team builds a feature that reads from it, and now the identity store is on the critical path of something that has nothing to do with authentication.
Nobody makes this decision. It's made by accretion, one attribute at a time, and by the time it's visible it's structural.
The line
The identity platform owns facts required to authenticate you and decide what you may access. Roughly: your identifier, credentials, authentication factors, tenant membership, coarse role or group, lifecycle state, and the audit trail of changes to those things.
Your application owns facts about who you are as a customer. Profile, preferences, billing tier, avatar, notification settings, business relationships, usage history, anything domain-specific.
The distinguishing question isn't "is this about a user?" — everything is about a user. It's:
Would this field ever participate in an authentication or access decision?
Job title: no. Timezone: no. Preferred language: no, unless you're rendering the login page, which is a real and narrow exception. Department: maybe, if it drives access — but see the warning below. Employment status: yes, because terminated means no access.
Most fields fail that test, and most fields end up in the identity store anyway.
How it happens
The progression is always the same.
You need a display name on the login page. Reasonable — put it in the identity store.
Then the application needs the display name too. It's already in the identity store and it's in the token, so read it from there. Now the identity store is the source of truth for a display field.
Then someone needs a phone number for SMS MFA. Legitimately identity. But now there's a phone number in the identity store, so when the billing team needs a contact number, they use that one. It's already there.
Then a customer asks to sync their org chart. Manager, cost center, employee ID — all arriving via SCIM, all landing in the identity store because SCIM is an identity protocol and that's where SCIM writes.
Two years later, the identity store contains the company's most complete record of every human, and four teams read from it. It is now, functionally, the master customer database. Nobody decided this. Nobody wrote it down. And it can never be changed again, because four teams depend on the shape of it.
What it actually costs
Every read is now a dependency on your most availability-critical system. The identity platform is designed to be available during your worst incidents — that's its job. But it wasn't sized on the assumption that the profile page, the billing screen, and the notification service all query it on every page load. You've taken a component with strict uptime requirements and given it a workload it wasn't scoped for.
The token bloats. Once attributes live in the identity store, the obvious move is to put them in the token so applications don't have to call back. Then the token contains a display name, a department, a timezone, and a list of group memberships. Tokens travel in headers, and HTTP header size limits are real — 8KB is a common default, and a user with many group memberships plus a few profile fields gets there faster than anyone expects. The failure mode is memorable: authentication works fine for everyone except your largest customer's most senior administrator, who belongs to the most groups.
Change velocity collapses. Identity systems change carefully and deliberately, on a slow, audited cadence — appropriately, given what they protect. Product data changes constantly. Once product data lives in the identity store, product changes inherit the identity system's release discipline. The marketing team wants a new preference field and it's a schema migration on the system that gates every login.
GDPR gets harder. A deletion request now spans a system whose whole purpose is keeping an immutable record of what happened. Untangling "delete this person's profile" from "retain the audit trail proving who accessed what" is much harder when both live in the same store — and the retention requirements point in opposite directions.
Your identity provider becomes unswappable. This is the one that has real strategic cost. Migrating identity providers is difficult but tractable if the identity store holds identity. If it holds your customer master data, migration means simultaneously moving your customer database, and that's the kind of project that doesn't get approved.
The attribute trap specifically
A warning about the case that looks most innocent.
An attribute like department arrives via SCIM from a customer's HR system. It's genuinely useful. Then somebody writes:
if user.department == "finance":
allow(financial_reports)
The identity attribute has silently become an authorization rule. Now HR renaming the department to "Finance & Operations" is a production access incident, and the people who made the change had no idea they were administering your access control.
Attributes are legitimate inputs to policy — "members of the finance department may access financial reports" is a fine rule. But the rule has to exist as a rule, in a place where you can enumerate who has access. If the only expression of your access policy is a string comparison against an HR-sourced field, you cannot answer "who can see financial reports?" without reading source code, and your security model is downstream of someone else's naming conventions.
What to do instead
Keep a user record on both sides, joined by the identity provider's subject identifier.
flowchart LR
IdP["Identity store\nsub, credentials, factors,\ntenant, role, lifecycle"]
App["Application store\nprofile, preferences,\nbilling, domain data"]
IdP -->|"sub (stable key)"| App
The application creates its own user row on first login, keyed on sub. Identity data stays in the identity store; everything else lives where it belongs.
Two details that matter more than they look:
Key on sub, never on email. Email addresses change, and at some organizations they get reassigned — a new hire receives a departed employee's address. Joining on email means that new person inherits the old one's account. The subject identifier is stable and opaque precisely so this can't happen.
Provision lazily on first login, not eagerly. Trying to keep a full mirror of the identity store synchronized into your application is a sync problem you don't need. Create the local record when the user first appears, refresh the handful of fields you care about from the token on subsequent logins, and let the identity store remain authoritative for the things it owns.
The honest counterarguments
Two lookups instead of one. Real, and usually irrelevant — the application read is a local query against your own database, not a network call to the identity provider. If you were reading profile fields out of the token you weren't making a lookup at all, so this does add work. It's typically a single-digit-millisecond join against data you already own.
"Who is the source of truth for a display name?" Genuinely ambiguous, and it's fine to keep a copy in both places for different purposes — the identity store's copy renders the login page, the application's copy renders the application. Duplication with a clear owner beats a shared field with none.
Small applications don't need this. True. If you have one application, no multi-tenancy, and no plans to change identity providers, putting everything in one place is a reasonable simplification, and the separation is overhead you're paying for flexibility you may never use. The boundary starts paying when you have a second application, a second identity source, or a migration on the horizon.
Some things genuinely sit on the line. Employment status is an identity fact (terminated means no access) and a business fact. Language preference matters for the login page and the application. There isn't a clean answer for every field — the test tells you where most of them go, and for the rest, pick an owner deliberately and document it.
The signal to watch for
You've crossed the line when a product team files a ticket against the identity platform to add a field.
That's the moment the identity system stopped being infrastructure and started being a shared database. It's worth treating that ticket as a design question rather than a schema change — because the field itself is never the problem, and the fortieth one always is.