Identity Isn't Your Audit System
The question arrived from outside counsel and it was one sentence long: between 1 March and 30 June, which of our employees accessed the records of these 340 customers?
It landed on the identity team, because the identity platform was where the audit log lived. Everyone knew that. It was the system with the seven-year retention, the append-only store, the WORM archive, the SOC 2 evidence, the whole apparatus.
Two days later the identity team sent back a careful answer. Over the four-month window the platform had recorded 1.14 million authentication events. Filtered to the internal support console — the only application that can display a full customer record — that was forty-one distinct employees, supplied with timestamps, source addresses, factors used, and the policy version in force for each. Every record accurate, every record immutable, none touched since the moment it was written.
Counsel then asked the follow-up, which is the only follow-up there ever is: which of those forty-one saw which of the 340 customers?
There was no answer to that, and there was never going to be, because the identity platform had not been present when any customer record was read. It watched forty-one people walk through a door. It had no idea what any of them did on the other side of it.
It got worse in the way that generalises. Three of the 340 customers had been flagged as never-accessed, on the strength of the identity platform showing no console logins from the team that owned them. That conclusion was wrong. The reads had come from svc-support-console, the console's own backend, which authenticated once on 27 February and then refreshed its credential silently every twelve hours for five months. In identity's records, five months of machine access to customer data appears as a single event dated four days before the window opened — so the answer was not merely incomplete. It was confidently, quietly negative about a thing that had happened forty thousand times.
The application's own data-access log had the real answer, for thirty days. By the time counsel asked, March, April and May had rolled off.
That is the boundary this article is about, and it is subtler than it looks, because I have argued the other half at length. Immutable audit logs are an engineering feature makes the case that an identity platform must produce rich, tamper-evident, reconstructible records, and I stand behind every word. This is the complement: identity is an authoritative producer of a narrow, well-typed class of audit events, and it is emphatically not the system of record for the enterprise's audit trail. Two different jobs — and most organisations have quietly assigned the second to whichever system did the first one well.
The resolution limit
Start with what makes this structural rather than a matter of effort, because almost nobody states it and it decides everything downstream.
An identity platform's audit trail is sampled at the session boundary, not at the action. It records the moment a credential is exchanged for an authorization artifact. Everything that artifact is then used for is invisible to it by design — that is the entire point of issuing a token rather than proxying every request.
Put numbers on it: a support agent, eight-hour shift, fifteen-minute access token, rolling session, an application making ten API calls per screen at forty screens an hour.
| Events identity observes | Actions the user takes | |
|---|---|---|
| Interactive login | 1 | — |
| Token refresh (15 min, 8 h shift) | ~32 | — |
| Screens viewed | 0 | 320 |
| API calls / record reads | 0 | ~3,200 |
Thirty-three identity events; roughly 3,200 accesses. About one to a hundred, and that is the favourable case, because refresh is chatty. Make the token lifetime an hour and it is one to four hundred. Give a service a long-lived refresh token rotated daily, as the support console had, and it is one to forty thousand. Issue a client-credentials token to a batch job and the ratio is unbounded — one event, arbitrarily much activity.
Notice which direction the incentives run. Every good thing you do to your token architecture makes identity's audit trail coarser. Longer-lived tokens reduce load on the issuer. Self-validating JWTs remove the introspection call that would otherwise leave a trace. Caching at the gateway — the correct design — deliberately eliminates per-request contact with the identity platform. Each is right, and each widens the gap between what happened and what identity saw.
The resolution of an identity platform's audit trail is set by its token lifetimes, and every performance and availability improvement you make lowers it.
So "route the audit question to the IdP" is not a slightly lossy strategy that better logging would fix. The information was never in the building. The only way to put it there is to route every customer-record read through the identity platform — a proposal we'll come back to, because someone always makes it.
What identity can actually testify to
There is a clean statement of identity's audit scope, and it is narrower than most teams' mental model:
Identity can testify to the issuance, use-for-issuance, and lifecycle of credentials and the policy state that governed them. Nothing else. It is a first-hand witness to exactly the events it participated in.
That is not a small remit — it covers questions nothing else in the estate can answer, and answering them well is hard. But it is a closed set, and the useful exercise is to write it down as a table you can route questions through.
| Question | Who can answer it | Why |
|---|---|---|
| Who authenticated to this application, when, from where? | Identity | It performed the authentication |
| Which factors were satisfied, at which assurance level? | Identity | Sole observer |
| Which policy version was in force on the node that decided? | Identity | Nobody else has the per-node projection state |
What was issued — scopes, audience, lifetime, jti? |
Identity | It minted the artifact |
| Who was granted or removed from an admin role? | Identity | Its own state transition |
| Who changed a tenant's MFA policy, from what to what? | Identity | Its own configuration |
| When was this credential revoked, and by whom? | Identity | Its own lifecycle |
| Who read customer 8812's record? | Application | Identity was not in the request path |
| Who approved the refund, and against which limit? | Business system | A domain event with domain state |
| What changed in the orders table at 03:00? | Database / CDC | Identity has no view of your data plane |
| Which files left via the API? | Gateway / DLP | Byte-level observation identity never makes |
| Which of these forty-one people saw record 8812? | Application, joined to identity | The join is real work, and it is somebody's job |
| Was this action taken by a human or an agent acting for one? | Identity, if you propagated the chain | See Auditing Agent Actions |
The last row cuts the other way and is worth naming. Delegation structure — sub stays the user, actors accumulate in act — is something only identity can attest to, because it issued every token in the chain. Resource servers record what they were told; the authorization server records what it asserted, independently. That is a genuine, underused custodial role, and a small and specific one: not the record of what happened, but the record of what authority existed.
Witness and custodian are different jobs
Here is the framing I would most like to survive this article.
An audit capability has two separable roles. The witness observes an event first-hand and produces testimony about it. The custodian holds the corpus, keeps it interpretable, makes it queryable, and produces it under legal process. Identity must be an excellent witness and should not be the custodian — and the reasons aren't preferences. They are four properties a custodian needs and identity structurally cannot have.
1. Retention, and the fact that three systems disagree
Everyone knows the seven-year number. What gets missed is that the same event is subject to three incompatible retention regimes at once, in three systems with three different cost curves.
| Tier | Typical retention | Optimised for | Cost driver |
|---|---|---|---|
| Identity's own operational store | days to weeks | serving the runtime, supporting the console | live database, indexed, replicated, in the availability path |
| SIEM, hot | 30–90 days | detection and correlation across sources | ingestion and indexing — the expensive one |
| Compliance archive, cold | 3–10 years | production under process | object storage plus retrieval |
The Hidden Cost of Audit Logs does the arithmetic on why the middle row dominates the bill, and I won't re-derive it. The boundary point is different: these are not three tiers of one system. They are three systems, with different owners, failure modes and deletion schedules — and the reason identity's own store must be the short one is that it lives inside the thing that has to stay up. A seven-year table in an operational identity database is an unbounded growth curve attached to your most availability-critical component, whose vacuum, backup, failover and schema-migration times all scale with it. You discover this during a major-version upgrade, when the table you never query is why the maintenance window doesn't fit in a weekend.
There is a second retention fact that is much less well known, and it determines whether your seven-year archive is worth anything at all.
The effective retention of an answerable question is the minimum retention across every system needed to interpret the record — not the maximum you paid for.
An identity audit record from 2021 says client_id=svc-7741, src_ip=10.42.7.19, groups=[g_8842, g_1190], subject=u_44119. Turning that into a sentence a regulator can read takes four lookups, each with its own lifecycle:
| Field in the seven-year record | What you need to interpret it | Typical retention of that |
|---|---|---|
client_id |
the client registry, with history | current state only |
src_ip |
DHCP / VPN / NAT logs | 30–90 days |
groups |
group definitions as they were then | current state only |
subject |
the user record, post-erasure request | possibly deleted |
Every one of those is a current-state store. The audit record is historical; its referents are not. So you keep seven years of records whose meaning expires in ninety days, and nobody notices until an investigation reaches back eighteen months and produces a list of opaque identifiers no living system can resolve.
The fix is cheap and has to happen at write time: denormalise the human-readable interpretation into the record itself. The client's name and owner, not just its ID. Group names alongside group IDs. The resolved network zone alongside the IP. It offends every normalisation instinct you have, and it is correct, because an audit record is not a row in a relational model — it is a statement that must remain readable after everything it refers to has changed. That is a witness's obligation, and one of the few things on this list identity should do more of.
2. The workload is inverted
Identity's audit store is optimised for a write pattern that is very close to the ideal case: append-only, monotonic in time, partitioned by tenant and day, essentially never read. Ninety-nine point something percent of records are written once and never touched again. That store wants time-ordered append, aggressive compression, and roughly two indexes.
An investigation is the opposite workload in every dimension.
| Identity's audit store | An investigation | |
|---|---|---|
| Access pattern | append, sequential | ad-hoc, random |
| Read:write ratio | ~1 : 10⁶ | all reads |
| Predicates | known at design time | unknown until the question is asked |
| Time range | recent | arbitrary, often distant |
| Cardinality | partition by tenant + day | filter on IP, device, subject, resource, none of which partition |
| Joins | none | across four systems that share no key |
| Latency tolerance | none (it's in the pipeline) | minutes is fine |
| Frequency | continuous | a few times a year |
Rows four and five are where the incompatibility bites. You can serve an investigation from an append-optimised store if you index the fields it will filter on. But you don't know those fields, and the ones investigators actually use — source IP, device identifier, user agent, session ID, resource ID — are all high-cardinality. Every one you index is a permanent multiplier on the storage and write cost of a stream you pay for continuously, to serve a query someone runs four times a year.
That trade is fine in a system built for it. A SIEM or columnar warehouse spends its money on exactly this: columnar layout so a filter on one high-cardinality field doesn't read the other forty, bloom filters and zone maps so most partitions get skipped without decompression, and a query engine allowed to take ninety seconds. An operational identity database has none of that, and adding it means running a second storage engine inside your identity platform — at which point you have built a SIEM with worse economics and one customer. The query you cannot anticipate should not be served by the system that cannot afford to index for it.
3. The correlation problem, which is somebody's job but not identity's
The forty-one names in the opening story are useless alone, and joining them to the application's access log is not a small task. Identity holds subject identifiers — an opaque sub, a session ID, a token jti. The application holds business keys — a customer ID, a case number, an employee number that came from the HR system four years ago and has no relationship to sub at all.
Nothing joins them unless someone deliberately built the join. The carrier is the token: sub and sid are the only identifiers that cross from the identity platform into the resource server, which is why a resource server that logs its local principal object and discards sid has silently destroyed the join. Across a redirect chain there is no global identifier at all — Debugging Across Four Layers works through why, and its answer is the same technique an investigation needs.
The boundary point is about ownership, not mechanism. Somebody must own the join, and identity is the wrong owner for a structural reason: the join is defined by the application's key space, which changes at application speed, while identity's contribution is a single opaque identifier that must never change. If identity owns the correlation it acquires a mapping table that grows with every application in the estate and breaks every time one re-keys. If the consuming system owns it, identity's obligation collapses to one rule: emit sub, sid, jti and a transaction identifier on every event, stably, forever, and let the join live where the other half of it lives.
4. Nobody can be the sole custodian of the evidence against themselves
This is the reason that has nothing to do with cost or query planning, and it is the one that ends the argument.
An identity platform is the highest-value target in the estate and the most likely subject of an investigation. When an incident involves an insider, a compromised admin credential, or the platform's own defect, the records that matter are the ones identity wrote about itself. If the only copy lives inside the system under investigation, controlled by the team that operates it, those records prove nothing — including for the people they would exonerate. 6.5 makes this argument about mutability; the custodial version is about possession, and possession is the stronger constraint. A perfectly immutable log whose only copy is administered by the party in question still has the wrong shape, because the question an investigator eventually asks is not "were these records altered?" but "are these all the records?"
The answer is not a better hash chain. It is a second custodian: records streamed continuously to a system operated by someone else — the customer's SIEM, the security team's warehouse, an escrowed archive — so that the copy outside the identity platform's blast radius is the one that gets produced. Merkle-root anchoring makes the two copies comparable; the export is what makes there be two.
There's a vendor-shaped corollary. If your identity platform is SaaS, treating it as the system of record means your enterprise audit trail is a folder in someone else's account — fine for detection, poor for everything downstream.
The proposal that always gets made
At some point someone proposes the elegant fix: if identity can't see what happened, let's send it what happened. An /audit endpoint on the identity platform. Applications POST their events, identity writes them into the same immutable store with the same retention, and now there is one audit system, one schema, one query surface.
It is a genuinely appealing idea, and it is wrong for three reasons of increasing severity.
It puts a firehose on a Tier-0 system. Identity's audit volume is bounded by session count; the enterprise's is bounded by action count, which by the arithmetic above is two to four orders of magnitude larger. You have proposed multiplying write load on your most availability-critical component by several hundred, generated by teams with no visibility into your capacity plan. The failure lands during a batch job.
It becomes a synchronous dependency the moment anyone makes it durable. As soon as an application needs its audit write acknowledged before it commits — and compliance-relevant events do — identity's audit endpoint sits inside that application's transactions. That is Identity Is Not Your Integration Layer inverted: rather than acquiring outbound dependencies, identity becomes an inbound dependency for everyone, with the same multiplicative arithmetic and nobody able to see the whole graph. And the endpoint's schema is a permanent public API the day the second team integrates — every extension point is.
And it destroys the only property identity's records actually have. This is the one I'd lead with in a design review:
An audit store that accepts events it did not observe converts first-hand testimony into hearsay, and there is no field that marks which is which.
Every record in your identity audit log has a property worth more than its retention: the system that wrote it was present at the event. Nobody can inject a false authentication record without compromising the authenticator itself. Open a write endpoint and that guarantee is gone globally, not locally — a reader six months from now cannot tell, from within the corpus, which records identity observed and which it merely accepted from a caller whose integrity is unknown. You can add a provenance field. Investigators will not read it, and the first time someone builds a compliance narrative on "our audit log is immutable," the narrative will quietly be describing a store where half the rows arrived over HTTP from an application anyone with a client credential could impersonate.
The right shape is a shared bus, not a shared store. Everyone publishes; a custodian consumes; identity is one producer among many and holds no one else's records.
flowchart LR
ID["Identity platform<br/>auth · issuance · lifecycle · config"] -->|"typed events"| BUS["Event stream"]
APP["Applications<br/>data access · approvals"] -->|"typed events"| BUS
DB["Databases (CDC)"] -->|"typed events"| BUS
GW["Gateway / DLP"] -->|"typed events"| BUS
BUS --> SIEM["SIEM — hot, 90d<br/>detection, correlation"]
BUS --> WH["Warehouse — warm<br/>investigation, joins"]
BUS --> ARC["Archive — cold, 7y<br/>WORM, legal hold"]
ID -.->|"short operational window<br/>days–weeks, for the console"| OPS["Identity's own store"]
Note what identity keeps: a short operational window sized for its own console and runtime, not for anyone's investigation. Everything long-lived is downstream, in systems whose job it is.
Identity's real obligation is an export, not a query API
If the custodian lives elsewhere, identity's audit obligation reduces to one thing most platforms do badly: a well-specified export. Not a search UI. Not a REST endpoint with ?from=&to=. A stream contract.
Take the query-API version seriously for a moment and cost it. A paginated admin log API, generously rate-limited at 10 requests per second with 500 records per page, moves 5,000 records/second — 432 million a day, in theory. Now pull 90 days from a deployment producing 86 million events a day: 7.7 billion records, or about eighteen days of continuous, error-free polling, against an endpoint sharing a rate-limit budget with your provisioning jobs. Nobody does this. They narrow the query until it fits — which means the investigation's scope is set by the export mechanism's throughput, a fact that appears in no design document and shapes the outcome of every incident.
Worse, polling a log API by timestamp is silently incomplete, and this is the detail I'd most want a platform team to internalise. Records have two times: when the thing happened (occurred_at) and when the store accepted it (recorded_at). Any pipeline with a queue, a retry, or an outbox has a gap between them. Page by occurred_at and every late-arriving record — precisely the ones written during an incident, when the pipeline was backed up — falls behind your cursor and is never returned. You get a complete-looking export with a hole in it exactly where the interesting events are. So: paginate by monotonic sequence or recorded_at, never by occurred_at, and let the consumer sort. If your vendor's API only offers occurred_at, you have a completeness bug you cannot fix from the outside.
A good export contract has eight properties. None are exotic; the value is in stating them as a contract rather than discovering them.
- A stable, versioned schema. Additive changes only, explicit version field, no repurposed fields. A record written today must parse in tooling written in three years, and vice versa.
- A closed, documented event-type taxonomy.
authentication.succeededmeans one thing forever; new behaviours get new types, not new meanings for old ones. - A stable idempotency key — a UUID per event — because delivery is at-least-once. Omit it and every consumer invents a different fragile hash of the payload.
- A monotonic per-partition sequence number, which makes gaps detectable rather than merely possible.
- Correlation and causation identifiers. Correlation says which flow this belongs to; causation says which event caused this one. Most schemas conflate them, and the difference separates "these forty events are related" from "this happened because of that."
- At-least-once delivery with an explicit lag SLO. "Within 60 seconds at p99" is a commitment a consumer can alert on; "eventually" is not.
- A documented gap-detection mechanism — below, and the one everybody skips.
- Minimum viable PII. Pseudonymous subject identifiers rather than email addresses; hashes and references rather than values. Not fastidiousness: an export goes to a SIEM readable by dozens of people and retained for years, and every identifying field is a copy you must find again when an erasure request arrives.
The heartbeat, and why silence is ambiguous
Property 7 deserves its own paragraph because it fails in a way that is invisible until it matters.
A consumer receiving no events cannot distinguish "nothing happened" from "the pipeline is broken." For an identity stream that is a realistic ambiguity — a small tenant genuinely might produce nothing overnight — and the failure is asymmetric: a broken export looks exactly like a quiet Sunday, and stays broken until someone runs an investigation and finds the hole. I have seen an export sit dead for nineteen days over a Christmas period for precisely this reason.
Two mechanisms fix it, and you want both. Sequence continuity: per-partition monotonic sequence numbers let the consumer detect loss positionally — it received 4,401 then 4,403, and can say so immediately without knowing anything about expected volume. And a periodic watermark emitted even when nothing happened: a heartbeat every minute carrying the partition, the current high-water sequence, and a count of events since the last watermark. Silence becomes a detectable condition, and the count is an independent reconciliation total that catches events lost before they got a sequence number.
Together they turn "did we get everything?" from an act of faith into a query, for a few bytes a minute.
The legal dimension, which is not a footnote
There is a set of obligations that attaches to a system of record and to nothing else, and it is the reason the custodian question can't be settled purely on engineering grounds.
Legal hold is a write-path capability, not a policy. When a hold is issued, deletion must stop for a named set of subjects or a named time range — and stay stopped, provably, until it is lifted, while normal retention continues expiring everything else. Now check whether your identity platform can suspend deletion for one subject. Almost none can: retention in an IdP is typically a single global TTL, sometimes per-tenant, applied by a background job with no concept of an exception. So the moment a hold is issued against an IdP-as-system-of-record you are in one of two bad positions — disable retention globally and accumulate everything forever, or let a compliance job spoliate evidence on a schedule. Neither is fun to explain.
Chain of custody wants a custodian who can testify. Producing records under process usually requires someone who can attest to how they were generated, stored and extracted. If the records live in a vendor's multi-tenant SaaS, that person works for the vendor, is not your employee, and is not obliged to appear. What you can produce yourself is an export you received and held.
Erasure and retention point in opposite directions, and the conflict is resolved at write time. An Article 17 request against an append-only seven-year store has two engineering answers, both decided before the first record is written: pseudonymous references, so identifying data lives in one deletable place, or per-subject encryption keys, so crypto-shredding renders content unrecoverable without touching record structure. Neither is retrofittable — and the more systems hold copies, the more places each has to be implemented, which is a real argument for concentrating custody in one designated system. Just not the identity platform. Data residency compounds it: a system of record inherits every jurisdictional constraint of every subject in it, and a regionally-deployed identity platform acquires a legal topology it was not designed around.
None of this makes a vendor-hosted IdP a bad choice. It makes it a bad system of record, which is a much narrower claim: use it to authenticate, stream its events out, and let the custodian be a system whose retention, hold and production semantics you control.
The operational tells
You have merged witness and custodian if any of these is true:
- A legal or compliance request for non-authentication activity gets routed to the identity team.
- Your identity platform's audit table is measured in years rather than weeks.
- The IdP's admin console is where investigators start.
- There is an endpoint on the identity platform that applications POST audit events to.
- Nobody can state the export lag, and there is no alert on it.
- An investigation's time range is decided by what the export API can move rather than by the facts.
- You cannot place a legal hold on a single subject without disabling retention globally.
- The identity platform's audit records reference identifiers that no current system can resolve.
- A database upgrade's maintenance window is dominated by a table nobody queries.
- The only copy of the records about your administrators is administered by your administrators.
And the failure modes, by misplacement:
| Misplacement | How it fails | When you find out |
|---|---|---|
| IdP treated as system of record | "No login events" read as "no access"; machine sessions invisible | An investigation with a lawyer attached |
| Seven-year retention in the operational store | Upgrades, backups and failovers scale with dead data | A major-version upgrade that won't fit the window |
| Investigation queries against the audit store | Full scans on an unindexed high-cardinality field, or new permanent indexes | The first serious incident |
Export paginated by occurred_at |
Silent gaps exactly during pipeline backlogs | Never, which is the problem |
| No heartbeat / sequence numbers | A dead export looks identical to a quiet weekend | Weeks later, if ever |
| Applications POSTing to an IdP audit endpoint | Write-load multiplication; testimony becomes hearsay | A batch job, then an audit |
| Identifiers not denormalised at write time | Old records are unreadable; referents have changed | Any investigation older than a quarter |
| Records held only by the party under investigation | Evidence proves nothing, including innocence | The insider case |
| Global-TTL-only retention | Legal hold forces all-or-nothing | The first hold |
What to do on Monday
- Route a real question from your last investigation through the table above. If it doesn't land on identity, check whether the system it does land on has retention that outlives your discovery time. That comparison is usually the whole finding.
- Measure your export lag, alert on it, then unplug the consumer for ten minutes in a test environment and confirm you notice. If a stopped export raises nothing, your audit trail's real retention is "until the next silent failure."
- Check what your export paginates by.
occurred_atis a completeness bug; a sequence number only helps if the consumer checks continuity rather than merely storing it. - Read one audit record from eighteen months ago aloud as a sentence. Every identifier you can't resolve is a write-time denormalisation you should have done — and the fix only helps records written from now on.
- Ask whether you can hold one subject's records against deletion, and confirm a copy of the stream lives somewhere your identity administrators cannot alter. Not a backup they can restore over — a downstream custodian with its own credentials.
Disclosure: I work on ClavionX. Its relevant constraint here is a general one rather than a feature: the runtime consumes projected state via events and never calls the control plane synchronously, which is why the audit stream is a first-class output rather than a table someone queries — and why the per-node policy version has to be stamped into each record, since there is no single global answer to "what was the policy at 09:11." The argument in this article holds whatever you run.
Identity has one audit obligation and it is narrow: be a scrupulous witness to the events it actually attended, describe them in a schema that will still parse in 2033, and hand them off promptly to someone whose job it is to keep them. The temptation is always to do more, because identity is the system that already writes good records and already has the compliance story, and building a second thing feels wasteful.
But an audit trail is not a place. It is a chain of testimony from many witnesses, assembled by a custodian, and the value is in the assembly. A witness who volunteers to hold everyone else's evidence has stopped being a good witness — and is now the only party who cannot be asked to produce it.