Who Owns Identity in Your Org Chart?
The quarterly business review had two slides in it, twenty minutes apart, and nobody in the room noticed they were about the same system.
The first was from the security organization. It showed nine controls shipped in the quarter: step-up authentication on eleven sensitive operations, a device-posture check, MFA enforcement extended to three more tenant tiers, session lifetime reduced from twelve hours to four, an IP allow-list capability, a re-authentication prompt before any change to payment details. The slide was green. The team had hit every commitment. Somebody said "great quarter" and meant it.
The second was from the product organization. It showed activation down 4% quarter over quarter, with a footnote that the drop was concentrated in the first-session funnel. There was a hypothesis about onboarding copy and an experiment planned for next quarter. Support ticket volume in the category "can't log in" was up 60%, but that category was owned by the support org's dashboard and didn't appear on either slide.
Both teams were doing their jobs correctly. The identity team — which sat under the CISO — had no metric on any dashboard anywhere that got worse when logging in got harder. Not one. Every number they were measured on improved monotonically as they added friction, and the number that would have contradicted them lived in a different org, on a different slide, attributed to a different cause.
That's the whole problem, and it isn't a story about security people being obstructive. Flip the reporting line and you get a mirror image of the same failure with different symptoms. I want to argue something more specific and less comfortable than "identity needs a good owner": the reporting line is usually not the variable that matters. The measurement is. A team that owns identity and is measured on only one half of the outcome will fail in the direction of its metric, reliably, regardless of which VP it reports to — and reorganizing without changing the measurement just relocates the failure.
This is about steady-state operational ownership of a system that is already running: who holds the pager, who approves a policy change, who says yes to onboarding an app. The separate and equally messy question of who decides things during a project — the governance structure, the decision latency, the six stakeholders who must all sign off before a line of code is written — is covered in why enterprise identity projects fail before anyone writes code. Different problem, different failure mode. This one starts the day after that project ships.
The four arrangements, and why three of them are defensible
There are essentially four places identity lands in a real org chart.
| Model | Typical trigger | Measured on | Fails toward |
|---|---|---|---|
| Security-owned | A pentest finding, an audit, a breach at a competitor | Controls deployed, findings closed, audit posture | Friction that nobody's dashboard punishes |
| Platform-owned | Identity started as a library, then a service | Uptime, latency, developer throughput | Authorization debt that never pages |
| Dedicated team | Enough apps and integrations that neither of the above can carry it | Its own SLOs, usually | A queue everyone eventually routes around |
| Nobody | Nothing. That's the point | — | Uncontrolled accretion by whoever needed something |
Three of these are legitimate choices with real trade-offs. The fourth is not a choice at all, and it is the one most organizations actually have — because "nobody owns it" is what you get by default when a system grows past the team that built it and no one picks up the deed.
What follows is not a ranking. It's an attempt to explain each failure mode by the incentive that produces it, because the incentive is portable and the stereotype is not. If you only remember "security teams over-restrict," you'll misdiagnose the security-owned team that doesn't, and miss the platform team failing the same way for the opposite reason.
The mechanism underneath all of it: falsifiability
Here is the general law, and everything below is a special case of it.
A team drifts toward whichever half of its mandate has a falsifiable metric.
Not the important half. Not the half leadership talks about. The half where the number can be shown to be wrong.
Availability is falsifiable. If login is down, the graph goes to zero, the pager fires, and no amount of narrative fixes it. Latency is the same — the p99 is the p99. These numbers have the property that reality argues back.
Security posture, in the general case, is not falsifiable. "How many breaches did we prevent this quarter" has no observable answer, because the counterfactual is unavailable. You cannot count the attacks that didn't happen, and you certainly cannot attribute them to specific controls. So the metric degrades, necessarily, into a proxy: controls deployed, findings closed, coverage percentage, policies enforced. Every one of those proxies is monotonic. There is no mechanism by which shipping another control makes the number worse.
This is the engine. A metric that only goes up produces a roadmap that only accumulates. A metric that can be falsified produces a roadmap that responds to reality. Put those two next to each other inside one team and the falsifiable one wins the sprint planning argument every time, because it's the one that can be demonstrably violated on a Tuesday.
That's why the two most common ownership models fail in opposite directions. It isn't culture. It's which half of the job has a number that can be proven wrong.
Security-owned: a roadmap made entirely of controls
Put the identity team under the CISO and you get a team whose inbound signal is audits, pentest findings, questionnaire responses, and customer security reviews. Every one of those channels generates work of exactly one kind: add a control, tighten a setting, close a gap.
Notice what never arrives through that channel. No pentest report has ever contained the finding "step-up authentication fires too often and users are abandoning checkout." No auditor has flagged that your session lifetime is too short. The enterprise security questionnaire has a hundred and forty questions and not one of them asks whether legitimate users can get in. The team's entire inbound queue is one-directional, and its roadmap is therefore one-directional, without anyone ever deciding that it should be.
The second-order effect is the interesting one. Because the proxy metric is monotonic, the team has no principled way to remove a control. Controls accumulate; they don't retire, because retiring one means arguing for a number going down with no offsetting number going up. Ten years in, you find step-up prompts guarding operations that stopped being sensitive in 2019, MFA challenges on read-only endpoints, an IP allow-list that predates the company being remote-first. Nobody defends these individually and nobody can remove them either, because "we made authentication weaker" is a sentence with a career cost and "we made authentication less annoying" is a sentence with no owner.
The diagnostic takes ten minutes: read six months of the team's tickets and count how many have a reduction in friction as their success criterion. In the backlogs I've looked at the answer is zero to two, and the two are accessibility fixes filed by someone else.
The honest defense of this model: it's the right call when the actual risk is severe and immediate, and when the organization has a demonstrated history of shipping identity that isn't safe. A quarter of pure controls work after a real incident is a correct allocation. The problem is that the arrangement outlives the reason for it, and the metric never signals that the emergency is over.
Platform-owned: the debt that never pages
Now flip it. The identity service lives with the platform team, alongside the service mesh, the CI system, and the internal developer portal. The team is competent, on-call, and measured on uptime, latency, and how fast product teams can ship. Login is fast. Deploys are clean. Everybody likes them.
And security debt accrues invisibly, for a reason that is precise and worth stating carefully, because it is the single most useful thing in this article.
Availability failures are self-reporting. Authorization failures are not.
If a policy is too tight, you find out within minutes: users cannot log in, they say so loudly, the support queue fills, error rates climb, the pager fires. The system tells you. If a policy is too loose, nothing happens. The request succeeds. It appears in the logs as a success. It is indistinguishable, at every layer of your telemetry, from the request that should have succeeded. There is no error rate that moves, no latency spike, no user complaint — because the person the loose policy admitted is not going to file a ticket about it.
I've made the technical version of this argument in every identity bug is a trust bug: the two failure directions are not symmetric, and only one of them ends when you deploy the fix. The organizational consequence is what matters here. A team's backlog is written by its inbound signal, and only one of identity's two failure modes generates one. A platform-owned identity team isn't neglecting authorization. It is responding, correctly and diligently, to every signal it receives — and half the problem space emits no signal at all.
This produces a characteristic shape. The team has excellent latency dashboards, a real latency budget, good load tests, thoughtful caching — and also a wildcard redirect URI added for a demo in 2022, a scope that grants more than its name suggests, a service account with admin because the narrower role didn't exist yet, and a token lifetime someone set to thirty days during an incident and never set back. Each was a locally correct decision that unblocked someone. None ever produced an alert.
The fix is not "care more about security." Caring is not a mechanism. The fix is to manufacture the missing signal — to build something that pages when a policy is too loose, so that authorization debt enters the backlog through the same door as everything else. That's what half the dashboard section below is about.
The dedicated team: correct at scale, and a queue
At sufficient size, both of the above stop being viable and you staff a dedicated identity team. This is the right answer, and I want to be clear that the criticism that follows is a criticism of the default implementation, not of the model.
A dedicated team owns a service every other team depends on and no other team can change. That is a bottleneck by construction. Whether it's a problem depends on one ratio: how much demand the team serves through self-service versus through its request queue. Serve 90% via self-service and the team scales sublinearly with application count. Serve 90% through a queue and it scales linearly, and eventually crosses the threshold where routing around it becomes rational.
That threshold is arithmetic, not attitude. An engineering team with a launch date compares two numbers: the cost of waiting for the identity team's queue, and the cost of not waiting. When the queue is two days, nobody routes around it. When the queue is six weeks and standing up your own SSO tenant takes an afternoon with a credit card, a competent team acting in the company's interest will route around it — and will be right to, given the information they have. Shadow infrastructure is not an indiscipline problem. It is a queue-length problem with a predictable trigger point.
The symptoms are observable, and you should go look for them, because the team being routed around is structurally the last to know:
- A second SSO tenant nobody registered. Usually found in the SaaS spend report before it's found in the architecture diagram.
- A shadow IdP. A team that needed three custom claims, couldn't get them scheduled, and now runs its own OIDC provider federated to nothing.
- Service accounts with human owners. A user account with a person's name, a password in a vault, and a cron job using it — created because provisioning a real machine identity required a form. This is one of the ways non-human identity quietly becomes the majority of your principals without anyone tracking it.
- Hard-coded credentials with a comment. Sometimes literally:
// TODO: move to the identity service when they have capacity. That comment is a dated receipt for a queue length. - Applications integrating against the database directly, or against a stale user-export job, rather than the API.
- A per-team "auth helper" library with three forks and no owner — what happens when the shared path is too slow to change.
Each is load-bearing structure built out of the shortest available material, and each is now yours forever whether you know it or not. Remediation costs an order of magnitude more than serving the original request would have, which means a six-week queue isn't saving the identity team six weeks of work. It's deferring it and multiplying it.
The design response: convert queue work into self-service work, and treat any request type appearing more than a few times a quarter as a product gap rather than a ticket. Delegated administration is the biggest lever, because "add a user," "reset a factor," and "grant a role" are the highest-volume queue items everywhere and none should reach a central team.
Nobody: the worst one, and the most common
There's an intuition that unowned systems sit still. They don't. Identity configuration is not a codebase that stays where you left it — it's a mutable control surface that dozens of people have legitimate reasons to touch.
So when nobody owns it, the changes still happen. A new SAML connection added by whoever ran that customer onboarding. A relaxed password policy for a tenant that complained. A scope widened during an incident at 3am. A group created for a migration and never deleted. A token lifetime extended because a batch job kept failing.
Each is made by someone competent, for a real reason, under time pressure. None is wrong in isolation. What's missing is not judgment at the point of change — it's reconciliation: someone whose job includes holding the union of all the changes and asking whether the resulting system is still the system you meant to have. Reconciliation is the only function that has no local sponsor. Every individual change has someone who wants it. Nobody wants the audit.
The result is a configuration that is a fossil record with no reader, and it degrades in the characteristic way described in why identity systems fail gradually and then all at once. There is no incident to point at until there is a very large one.
"Nobody" also produces a risk-acceptance vacuum. Each of those changes carries risk, and risk nobody accepts explicitly is still accepted — implicitly, by whoever was nearest the keyboard, on behalf of an organization that never learned it had taken it on. When that surfaces in a postmortem, the finding is the one from who owns the incident when every layer worked: everyone acted correctly and the outcome was still bad, because the property that failed belonged to no one.
If you take one thing from this section: unowned identity is not a system standing still. It is a system being changed continuously by a large number of uncoordinated hands with no one holding the sum.
The claim: fix the measurement before you move the box
Here's the part I actually want to defend, because it's the non-obvious one and it saves reorgs.
Every failure mode above is a metric failure wearing an org chart costume. Security ownership fails because the security half of the mandate has an unfalsifiable metric that degrades into a monotonic proxy. Platform ownership fails because only one of identity's two failure modes emits a signal. The dedicated team fails because it's measured on the reliability of what it runs, not on the latency of what it's asked for. Nobody's ownership fails because no metric exists at all.
Which means the intervention with the best return is not moving the team. It's giving whoever owns it paired metrics that pull in opposite directions, with both halves on the same dashboard, in the same review, owned by the same person.
The pairing is the point. A single metric can be optimized monotonically; a pair cannot, because every improvement in one has to be defended against movement in the other. That forces the trade into the open, where people with the context can argue it. Same logic as an error budget: the value isn't the number, it's that someone now has standing to argue the other side.
| Control half | Friction half | The conversation it forces |
|---|---|---|
| Policy drift: rules that have loosened since last audit | Login success rate, by tenant and by method | "This got safer — did anything get worse for anyone?" |
| Median time-to-revoke access | Median time-to-onboard an application | "We can kill access in 90 seconds and it takes six weeks to add an app" |
| MFA / phishing-resistant coverage | Step-up prompt rate per successful session | "Coverage is 94% — how often are we interrupting people to get it?" |
| Privileged accounts without a named human owner | Self-service ratio: requests served without a ticket | "Who owns these, and why is everything still a ticket?" |
| Credential age distribution (secrets, keys, certs) | Rotation-induced incidents in the last two quarters | Whether your rotation story is real or theoretical |
| Authorization coverage: endpoints with no policy | Support tickets in the "can't access" category | Both halves of the same boundary, side by side |
Two notes on building this.
Policy drift is the load-bearing one, because it is the manufactured signal for the failure mode that doesn't self-report. Concretely: snapshot the effective policy set — token lifetimes, redirect URIs, scope definitions, MFA requirements, role grants, session limits — on a schedule, diff consecutive snapshots, and route anything that loosened to a human. Not a report; a ticket, in the same queue as everything else, with the same triage. That single mechanism converts invisible authorization debt into ordinary work with an inbound signal, which is exactly what the platform-owned team was missing. It is also cheap: it's a diff and a rule about which direction is "looser."
Time-to-onboard-an-application is the load-bearing one on the other side, because it's the number whose growth predicts shadow infrastructure. Track it as a distribution, not a mean — the p90 is where teams decide to route around you. If you measure nothing else from this table, measure that one, and treat a rising p90 as an incipient outage in a system you don't monitor yet.
Once both halves sit on one dashboard with one owner, the reporting line matters much less than it appears to. A security-owned team with a login-success SLO behaves very differently from one without; a platform-owned team with a policy-drift queue stops accumulating silently. Both changes are cheaper than a reorg and reversible, which is not a small property.
On-call, and the question that decides ownership
Everyone asks "who carries the identity pager?" It's the wrong question, or at least the shallow version of the right one.
The right question is: at 3am, who is authorized to make the trade?
The identity incidents that matter are exactly the ones where the two halves conflict. A device-posture check is failing closed and forty thousand people cannot log in. Turning it off restores the business in ninety seconds and lowers your posture for the duration. Somebody has to decide now, with partial information.
If your on-call engineer can make that call inside their authority, you have functioning ownership. If they must wake a security director for permission — or worse, don't know whether they need to — you have two half-owners and a coordination cost paid in minutes of downtime during the only moments that count.
Three things make this workable, all decided in advance:
- Pre-authorized degradations. A written, short list of controls the on-call engineer may disable unilaterally during a declared incident, with a mandatory time bound and automatic re-enablement. Not a judgment call at 3am — a menu approved in daylight.
- A named risk-accepting role with a phone number, for anything outside that list. One person, escalating to one deputy. If the escalation path is a Slack channel, there is no escalation path.
- Mandatory review of every degradation within one business day, which is what stops the temporary session-lifetime extension from becoming permanent. The failure mode isn't the emergency change; it's that nothing forces it back. Wire the re-enablement to a timer rather than an intention. Operational specifics of running this rotation are in running an identity platform on call.
Where the boundary sits with application teams
Ownership questions get answered badly when the scope is vague, and identity scope is unusually easy to get wrong in both directions. The clean split follows the protocol boundary, and it's worth stating as a rule:
The identity team owns the assertion. Application teams own what they do with it.
The identity team is accountable for who this principal is and what has been verified about them — the assertion's correctness, its issuance, its lifetime, its revocation, and the availability of the system that produces it. Application teams are accountable for what that principal may do here, because that question is business logic and belongs with the business logic. This is the organizational form of authentication is not authorization, and drawing it here saves you from the failure described in identity isn't your authorization engine, where the identity team accumulates every product's permission model and becomes a change-approval board for features it does not understand.
| Identity team owns | Application team owns |
|---|---|
| Authentication methods, MFA, enrollment, recovery | Which operations require step-up (identity provides the mechanism) |
| Session and token issuance, lifetime, revocation | Honoring revocation promptly; not caching decisions past their stated life |
| Federation, SAML/OIDC connections, IdP onboarding | Consuming the assertion correctly; validating it every time |
| The policy mechanism and its floors | The policy values within the permitted range |
| Provisioning and deprovisioning plumbing (SCIM) | Its own resource-level authorization decisions |
| Audit of identity events: authn, grants, admin actions | Audit of business events |
The two rows most often gotten wrong: step-up, where the identity team should own the primitive and the application should own the trigger list, because only the application knows which operations are sensitive; and policy floors, where the identity team sets a permitted range and application or tenant owners choose within it. That floors-and-ranges model is the same one that makes delegated administration safe, and it's what keeps "you may configure this" from meaning "you may configure this to zero."
One boundary trap worth naming: application teams routinely ask the identity team to store one more attribute, because it already has a user record and is already in the request path. Say no. It's the cheapest yes in the building, and it's how identity becomes everyone's user database and then everyone's integration layer — at which point every product team's schema change needs your review and your queue explodes for reasons unrelated to identity.
The honest counterpoint: at small scale, the right answer is "part time"
Everything above describes organizations past a threshold. Below it, the correct arrangement genuinely is: the platform team owns identity, part time, alongside six other things, and nobody is titled for it.
This isn't a compromise; it's correct, for structural rather than budgetary reasons. With eight engineers and four services, the coordination cost of a formal boundary exceeds the cost of the failure modes it prevents. Reconciliation happens implicitly because three people were all in the same standup. The queue is zero because there is no queue. A dedicated team at that size doesn't reduce risk — it creates handoffs where there was shared context, and process that must be maintained by people needed elsewhere.
Over-formalizing early has its own recognizable failure mode: an identity "team" of one and a half people, a request form, an intake process, a review board, and now a two-day change takes two weeks in an organization whose risk profile never required any of it. That's the queue problem, self-inflicted, without the scale that justified the structure.
So the useful question is not "which model" but "have we crossed the threshold yet." These are the signals I'd trust, and any two together mean the answer is yes:
- Identity work has exceeded roughly a third of one person's time for two consecutive quarters. Not a spike — a floor. This is the single most reliable indicator, because it means the part-time arrangement is now a full-time job being done in the margins of another one.
- More than one external IdP integration exists. One is a feature. Two is a compatibility surface, and the second always reveals the first was IdP-specific rather than standards-compliant — the identity tax arriving on schedule.
- The first question you can't answer from memory. "Which applications can read this scope?" — and finding out takes an afternoon. Implicit reconciliation has stopped working.
- A compliance regime with named evidence requirements. The moment someone must attest to a control quarterly, that control needs an owner with a name, because "the platform team, collectively" cannot sign anything.
- The first shadow integration. Somebody stood up their own auth because yours was inconvenient — a queue forming before anyone called it a queue.
- The person who built it is the only one who understands it. Not an ownership model but a single point of failure with a laptop, realized on their notice period rather than on an incident.
Crossing the threshold does not mean hiring an identity team on Monday. The cheapest correct first move is almost always to name a single owner — one person, part time, with the paired metrics above on their dashboard and the authority to say no. That converts "nobody" into "somebody," which is the largest single improvement available in this entire article, and it costs one line in a document.
What to actually do
Four things, in order of return relative to cost.
Name an owner, even a part-time one. The gap between "nobody" and "one person, 30% allocated, with authority" is larger than the gap between any two formal models.
Give that owner both halves, measured. One dashboard, both columns, one meeting, one person. This is the intervention that changes behavior without changing anyone's manager.
Manufacture the missing signal. Policy-drift detection, routed to a real queue with real triage. Half of identity's failure modes are silent by construction, and a team works only the tickets it receives.
Track time-to-onboard-an-application at p90 as a leading indicator of shadow infrastructure. When it climbs, teams are already routing around you; you just haven't found the evidence yet.
Then leave the reporting line alone unless you have a reason beyond the failure modes described here. Moving the box is the most visible intervention available and usually the least effective — it consumes a quarter of organizational attention, relocates the incentive problem into a new department, and produces a fresh set of predictable failures with the same root cause. Metrics are cheaper to change, faster to change, easier to reverse when you get them wrong, and they were determining the outcome the whole time.
Whoever owns identity has to be able to lose. If nothing on their dashboard can get worse because of a decision they made, they aren't owning the system — they're owning one half of it, and the other half is accumulating somewhere you're not looking.