The Cost of SCIM Synchronization

At 02:00 on a Saturday, a customer's HR system finished its weekly export and something in it decided that the Field Operations division should now be called Field Services. One string, one field, one person's Friday-afternoon ticket.

The identity platform did what it was built to do. It diffed the export, found 38,400 user records whose department attribute had changed, recomputed the dynamic groups those users belonged to — three each — and enqueued the resulting changes for the eleven downstream applications that tenant had connected over SCIM.

By Monday morning the queue had 1.4 million operations in it and was draining at the only rate the slowest target would accept, which was 100 requests per minute. The dashboard showed the pipeline healthy: no errors, no retries in the dead-letter queue, throughput exactly at the configured ceiling. It was working perfectly. It was going to be working perfectly for another two days.

The part that mattered was not in the dashboard. On Friday at 16:40 — nine hours before the HR export ran — a contractor had been terminated. The active=false operation for that contractor's account in the tenant's document-management system was item 1,118,204 in a FIFO queue. It was dispatched on Tuesday afternoon, 91 hours after the termination, into a product the platform's own marketing page described as offering "near-real-time deprovisioning."

Nobody had under-provisioned anything. There was no incident, no page, no error budget burned. The failure was a queueing failure: the one class of event with a security SLA attached to it was sharing a lane with an unbounded volume of events that had no SLA at all, and one person renaming a department was enough to bury it.

That is the shape of SCIM economics, and it is why this is the least glamorous and most persistently expensive subsystem in enterprise identity. The infrastructure bill is genuinely trivial — I'll show you it's under $200 a month at real scale. The cost is somewhere else entirely, and so is the risk.

This is the integration chapter of a series doing arithmetic on identity infrastructure. The Hidden Cost of Audit Logs set the method; The Cost of Session Storage found a memory bill that was really a write bill; The Cost of Password Hashing priced a security control as a deliberate purchase; The Cost of Token Introspection vs Self-Validating Tokens found payroll dominating both options.

SCIM has a shape none of those have. Every other line item in this series scales with your traffic, runs on your hardware, and is bounded by decisions you control. SCIM's cost is a product of three numbers — events, integrations, tenants — and is bounded by the availability, rate limits and spec compliance of software written by other companies who owe you nothing.

The assumptions, published so you can re-run it

Parameter Value Note
Enterprise tenants 200 Same scenario as the rest of the series
Identities under management 600,000 Median tenant 1,200; largest 40,000
Group memberships (edges) 7.2M ~12 groups per identity
SCIM target types supported 14 Distinct vendor implementations
Targets connected per tenant 6 average (1–30 range) → 1,200 sync channels
Apps interested in a given user 4 of 6 Not every app gets every user
Annual employee turnover 15% Plus 10% internal moves
Downstream rate limit 100 req/min per channel Median across the 14; range 20–1,000
Downstream p99 write latency 1.2 s mean, 4 s p99, 30 s worst target
vCPU $25/vCPU-month Blended, as elsewhere in the series
Log ingestion + indexing $0.50/GB Mid-range negotiated
Fully-loaded engineer $16,700/month $200k/yr

Two of those carry more weight than they look like they do. Rate limit is not your parameter — it is the downstream vendor's, it varies by two orders of magnitude across your target set, and some vendors change it without telling you. Targets connected per tenant is the multiplier that turns a linear problem into a quadratic-feeling one, and it is the number that grows every time sales closes a deal.

The event volume nobody counts

Ask a team what their SCIM system does and you'll be told: joiners, movers, leavers. Ask them how many of those there are and you'll get an answer derived from turnover.

At 600,000 identities and 25% combined turnover-plus-moves, that's 150,000 lifecycle events a year — about 410 a day. Four hundred and ten. You could dispatch those by hand.

Now count what actually flows through the pipe.

Source of change Daily volume Why
Joiners / movers / leavers 410 25% annual churn on 600k
Attribute changes ~12,000 2%/day of records carry a changed synced attribute
Group membership edge changes ~36,000 0.5%/day of 7.2M edges
Total source events ~48,400
Of which security-relevant 410 (0.8%)

The 2% attribute-churn figure is the one people dispute, and it's the one I'd defend hardest. It is not 2% of employees changing job title every day. It's this: HR systems and directories do not emit deltas. They emit records, and your differ decides what changed. A record carries a title, a manager, a cost centre, a location, a phone extension, a license flag, a preferred name, and a dozen fields somebody added for a payroll integration in 2019. Any one of them moving makes the whole record dirty. Add a nightly job at the customer's end that normalises phone numbers, or a manager hierarchy that recomputes when anyone anywhere is promoted, and 2% is conservative. I have seen tenants where the honest figure was 100%, every night, because the export included a lastSyncedAt timestamp inside the payload.

Group edges are worse and are almost never modelled. Dynamic group rules mean a single attribute change fans out into membership changes before it ever reaches SCIM. One person moving teams can be one attribute change and eleven edge changes.

So: 0.8% of your events are the ones with a security SLA, and 99.2% are cosmetic. Hold onto that ratio. Almost every design error in SCIM pipelines comes from treating the two categories as one workload.

Why the protocol makes this expensive

SCIM is a per-resource REST protocol. One user, one URL, one HTTP request. That is a fine design for a management API and a poor one for a synchronization bus, and the gap between those two things is where the money goes.

The spec anticipated this. RFC 7644 §3.7 defines a /Bulk endpoint that posts many operations in one request. In practice, the number of major SaaS targets implementing it usefully is close to zero, and those that do often cap it low enough that it saves round-trips rather than rate-limit budget. Probe for it in /ServiceProviderConfig, support it where it exists, and don't build a cost model that assumes you'll get it.

So each logical change costs you more than one HTTP request:

Logical change HTTP operations against a typical target
Update one user attribute 1 PATCH — if you already know the target's internal resource ID
Update one user attribute, ID unknown 1 GET ?filter=userName eq "..." + 1 PATCH
Add a user to a group 1 PATCH on the Group (or, on targets that reject group PATCH, 1 GET + 1 PUT of the entire member list)
Deprovision 1 PATCH active=false, or 1 DELETE, depending on the vendor's chosen semantics
Create a user 1 POST, plus 1 GET first on targets that return 400 rather than 409 on duplicates

Take an average of 2 HTTP operations per logical change and multiply through:

48,400 source events/day
  × 4 target apps interested per user
  = 193,600 logical operations/day
  × 2 HTTP requests
  = 387,200 requests/day
  = 4.5 requests/second, averaged

Four and a half requests per second. That is the number that makes every architect in the room relax, and it is the most misleading figure in this article — a daily average over a workload that neither arrives evenly nor can be dispatched evenly.

The peak is 26× the mean, and the mean is irrelevant

Enterprise directory churn is nocturnal and batched. Roughly 80% of a day's changes land in a twenty-minute window when the customer's HR export runs, and customers cluster into a handful of business timezones — call it 45% of your tenants in the largest cluster.

193,600 ops/day × 0.80 in batch windows   = 154,880
  × 0.45 largest timezone cluster          = 69,700 ops
  ÷ 1,200 seconds                          = 58 logical ops/s

Against a daily average of 2.2. A peak-to-mean ratio of about 26×, which is far worse than the 3.3× you'd size a login tier for, and which arrives on a schedule you don't control and can't smooth by asking nicely.

Now put the rate limits back in. Each sync channel — one tenant, one target — is capped at 100 requests/minute, so 1.67 requests/second per channel, 0.83 logical operations per second. Your global capacity is not a number you provision; it is the sum of 1,200 independent allowances, and it is completely non-fungible. A channel with spare budget cannot lend it to a channel that's saturated. You can have 99% of your capacity idle and one tenant's queue three days deep at the same time, which is exactly what happened in the story at the top.

Here is that reorg priced properly:

38,400 users × (1 attribute change + 3 group edge changes) = 153,600 source events
  → against ONE target app, at 2 HTTP requests each        = 307,200 requests
  ÷ 6,000 requests/hour (100/min)                          = 51 hours

Fifty-one hours of queue, per target, from one text field. The eleven targets drain in parallel, so the wall-clock is ~51 hours rather than 561 — but every one of those channels is saturated for two days, and anything else that needs to traverse them waits.

The single highest-leverage knob is time, not throughput

You cannot raise the rate limit. You can change how many logical operations a given amount of change turns into, and the lever for that is coalescing: hold changes for a short window and merge everything targeting the same (user, target) pair into one operation.

Naive:      153,600 events × 2 HTTP  = 307,200 requests → 51 hours
Coalesced:   38,400 users  × 1.2 HTTP =  46,080 requests →  7.7 hours

A 60-second debounce window collapses a 51-hour storm into under 8 hours, and it costs you a hash map. SCIM cost is set by how you batch in time, not by how fast you send — and the fastest queue is the one you never enqueue into.

The catch is the obvious one: a coalescing window is latency you are choosing to add. Which brings us to the actual architecture.

Two lanes, and per-target bulkheads

The design error underneath the opening story is that the pipeline had one queue. Fixing it does not require more capacity; it requires admitting that you are running two workloads with different economics and different SLAs, and that one of them is allowed to starve.

flowchart LR
    S["Directory / HR<br/>change events"] --> C{"Classify"}
    C -->|"lifecycle: create,<br/>deprovision, suspend<br/>410/day"| L["Priority lane<br/>no coalescing<br/>reserved budget"]
    C -->|"attribute + group churn<br/>48,000/day"| A["Bulk lane<br/>60s coalescing<br/>best effort, preemptible"]
    L --> B["Per-target concurrency<br/>+ rate budget<br/>(bulkhead per vendor)"]
    A --> B
    B --> T1["Target A"]
    B --> T2["Target B"]
    B --> T3["Slow target C"]
    R["Reconciler<br/>weekly full compare"] --> A

Three properties matter here, and none of them are about scale.

The priority lane is never coalesced and never queued behind bulk work. 410 events a day against a 100/min budget is 0.005% utilization. You can reserve 10% of every channel's rate budget for lifecycle events and lose nothing.

Bulk work is preemptible and droppable. If the bulk lane is 300,000 operations deep, the correct action is often to discard it and let the reconciler fix it, because a 300,000-item backlog of attribute updates and a weekly full-compare converge to the same end state — one just gets there without holding queue slots for two days.

Per-target bulkheads exist because your worst vendor will otherwise consume your entire concurrency pool. This is Little's Law used as a weapon. A shared pool of 200 worker slots, and one target whose p99 is 30 seconds:

5 ops/s to the slow target × 30 s latency = 150 slots occupied

That target is carrying about 8% of your traffic and holding 75% of your global concurrency. Every other tenant's sync — including deprovisions — is now bounded by a vendor you have no relationship with and cannot page. Your throughput is set by your worst endpoint unless you build a wall around it, and the wall is a fixed per-target slot budget, sized as target_throughput × target_p99, enforced independently of the global pool.

Note the asymmetry that makes this worth doing: you can bulkhead your way out of a slow vendor, but you cannot bulkhead your way out of a wrong one. Which is the next section.

Reconciliation is a permanent line item, and it is not optional

Drift is guaranteed. Not likely — guaranteed. The mechanisms:

  • A 5xx you retried, that had actually succeeded.
  • An admin who edited the user directly in the downstream app.
  • A vendor deployment that dropped externalId on a subset of records.
  • A tenant that revoked and reissued your token, silently discarding queued work.
  • An operation you dead-lettered eight months ago and nobody triaged.
  • meta.lastModified with second granularity and no ordering guarantee, so your delta watermark either misses changes that landed in the same second or re-sends them. (Use >= with idempotent writes and dedupe, never >. This bug is subtle, silent, and everywhere.)

So you must periodically compare full state. Price it:

600,000 identities × 4 channels each = 2.4M records to fetch
  ÷ 100 records per page              = 24,000 page requests
  spread across 1,200 channels        = 20 pages per channel
  at 100 req/min                      = 12 seconds per channel
2.4M records × 2 KB                   = 4.8 GB per full reconcile

Weekly, that's 250 GB/year of transfer and about twelve seconds of wall-clock per channel. Reconciliation is essentially free. State that plainly, because teams routinely defer building it on cost grounds, and cost is not the reason it's hard.

The reason it's hard is that reconciliation forces you to write down a policy you have been avoiding: when local and remote disagree, who wins? Blindly overwriting the target destroys legitimate local administration and generates a fan-out storm of its own. Blindly accepting the target means your identity platform is no longer authoritative. The honest answer is per-attribute and per-tenant, it requires a conversation with the customer, and it is a product decision wearing an infrastructure costume.

And there is a second trap. Some targets don't support filter on externalId at all — a few don't support filter beyond userName eq. Against those you cannot ask "do you have user X?"; you can only list everything and compare locally. That means you must maintain your own mapping table of local identity → remote resource ID, for all 2.4M pairs, durably, forever. Lose it and your only recovery is a full enumeration of every target. It is a small table with an outsized blast radius, and it belongs in your backup and DR plan explicitly.

"SCIM is a standard" is half true, and the other half is your payroll

Here is the infrastructure bill for everything above.

The work is IO-bound, not CPU-bound. At a 58 ops/s peak with 1.2 s mean downstream latency, Little's Law gives ~70 concurrent in-flight requests. On an async runtime that's a couple of vCPU; on thread-per-request, a few hundred MB of stacks. The change journal — every operation with its request and response captured, because vendor support will demand it — runs about 6 KB per operation, so 250,000 operations/day is 1.5 GB/day, roughly $23/month to ingest and index at $0.50/GB. Add the queue, the mapping table, egress.

Line item Monthly
Worker compute (2–4 vCPU, HA) ~$100
Queue + mapping store ~$40
Change journal ingest/index (30-day hot) ~$23
Egress (reconcile + operations) ~$10
Infrastructure total ~$175

A hundred and seventy-five dollars a month to synchronize 600,000 identities across 1,200 channels. This is why SCIM never shows up in a FinOps review, and why nobody notices what it actually costs.

What it actually costs is 14 vendor dialects, forever.

SCIM 2.0 is a real standard with real conformance, and it is still true that every target you integrate is a distinct maintenance stream. A partial list of the divergences you will write code for:

Divergence What breaks
PATCH with filtered path (members[value eq "x"]) unsupported Group updates degrade to GET + full-list PUT — O(group size) per membership change
Deprovision as active=false vs DELETE Wrong choice leaves a licensed, sometimes still-loginable account
groups written on User vs members written on Group Spec says groups is read-only on User; several vendors require writing it
startIndex 0-based instead of 1-based Off-by-one that silently skips or duplicates the first record of every page
externalId not persisted Your only correlation key is gone; you now own the mapping table
Duplicate create returns 400, not 409 Retry logic can't distinguish "already exists" from "bad payload"
userName case sensitivity Duplicate accounts for A.Smith and a.smith
meta.lastModified second-granularity Delta sync watermarks are unreliable by construction
Rate limit undocumented or per-org rather than per-token You discover it during a customer's onboarding

None of these are exotic. Every one is a branch in your code, a test fixture, and a thing that breaks when the vendor ships.

Steady-state maintenance runs about half an engineer-day per integration per month once you include vendor changes, tenant-specific behaviour, and support escalations. Fourteen integrations:

14 × 0.5 engineer-days/month = 7 days/month ≈ 0.35 FTE = ~$5,800/month
Plus on-call and support triage for 1,200 channels ≈ 0.4 FTE = ~$6,700/month
                                          Total payroll ≈ $12,500/month

Infrastructure is 1.4% of the cost of this subsystem. Payroll is 98.6%. The Cost of Token Introspection vs Self-Validating Tokens reached a similar conclusion for a different reason — there, payroll dominated because validation middleware was distributed across forty services. Here it dominates for a structural reason that no amount of engineering removes: you are maintaining a compatibility layer against software you don't control, and the layer grows monotonically with your customer list. Introspection cost can be engineered down. Dialect cost can only be amortised.

The planning consequence: the honest unit of SCIM cost is not per-user or per-event. It's per-target-type, per-year. Roughly $5,000–$8,000 a year, permanently, for each vendor integration you support, mostly in salary. Sales asking for three new connectors is a $20,000 annual recurring commitment, and it will be presented as a two-week task.

Pricing the deprovisioning SLA

Everything above is throughput. The number an enterprise buyer actually cares about is latency, on one event class: how long between HR marking someone terminated and their access disappearing everywhere.

Price the ladder.

Target Architecture Marginal build Marginal run
24 hours Nightly full compare, one cron 2 weeks ~0.1 FTE
1 hour Hourly delta sync with watermarks +3 weeks (watermark correctness) ~0.2 FTE
5 minutes Event-driven push, priority lane, per-target bulkheads, retry/DLQ, lag observability +2–3 months ~0.5 FTE
< 60 seconds, guaranteed Not achievable — —

The last row is the important one and it is usually missing from the contract. You do not own the downstream. If a target is down, rate-limiting you, or has a 30-second p99, no architecture on your side produces a bounded time-to-effect. What you can bound is time-to-dispatch.

Write the SLA as: time-to-dispatch (yours, bounded, measurable) + a published reconciliation interval (your backstop) + per-target effect visibility (theirs, reported not promised). Any deprovisioning SLA phrased as a single number is promising someone else's uptime, and the first time a vendor has a bad afternoon you will discover you sold it.

Now the risk side, because the ladder is only worth climbing if it retires something.

At 15% turnover on 600,000 identities: ~246 leavers per day. Exposure-hours per day = leavers × mean deprovisioning latency:

SLA Exposure-hours/day vs 24h baseline
24 hours 5,904 —
1 hour 246 −96%
5 minutes 20.5 −99.7%

That table is seductive and I want to argue against reading it too literally. Aggregate exposure-hours are the wrong risk model, because the overwhelming majority of leavers are voluntary, cooperative, and have already handed back the laptop. Nine thousand exposure-hours belonging to people who have emotionally moved on is not nine thousand hours of risk.

The risk lives in the tail. Of 90,000 annual leavers, perhaps 8% are involuntary, and of those a handful are terminations conducted while the person is still logged in and holding a grudge — a few dozen a year where deprovisioning latency is the only control between an angry person and a document repository.

You do not buy 5-minute deprovisioning for the average leaver. You buy it for the twenty or thirty events a year where minutes are the whole control. And that reframing changes what you build, because those events don't need throughput — 30 events a year is 0.000001 ops/second. They need a bounded-latency priority path, which is cheap to design in and close to impossible to retrofit onto a FIFO queue. The most expensive mistake in this entire subsystem is not a capacity mistake. It's building one queue and discovering, during an actual termination, that the security event is behind a department rename.

Three architectures, honestly

Batch (nightly full compare) Event-driven push Just-in-time lookup at login
Deprovisioning latency Up to 24 h Seconds to minutes Never fixes it — see below
Attribute freshness 24 h stale Seconds Perfect
Load shape One enormous spike Smooth-ish, storm-prone Per-login, on the critical path
Downstream cost High (full enumeration) Low (deltas) Zero
Login availability Unaffected Unaffected Coupled to the source system
Login latency Unaffected Unaffected + source system p99
Failure mode Silent staleness Queue depth, drift Login failures
Engineering cost Low High Moderate — then very high
Honest verdict Fine below ~20 tenants The right answer, with lanes Solves the wrong problem

Batch is genuinely fine at small scale and is dismissed too fast. If you have 15 tenants and two target types, a nightly full compare is a hundred lines of code with no queue, no watermark bugs, no drift, and no reconciler — because it is the reconciler. Do not build an event pipeline to serve a workload of 410 daily events until the tenant count forces you to.

Just-in-time lookup at login is the interesting one, because it's the answer that keeps getting proposed and it is wrong for a reason people usually get only half right.

The half everyone gets: it puts a third-party system on your authentication critical path. A login that calls the customer's HR API inherits its latency (enterprise HR APIs have p99s measured in seconds, not milliseconds) and, worse, its availability. Multiply 99.9% by 99.5% and you have 99.4% — about 52 hours a year of login failures attributable to a system with no on-call rotation for logins. It also makes your latency SLO unmeasurable, since it's now a function of somebody else's deployment schedule. That argument is developed at length in The Extensibility Trap and Identity Is Not Your Integration Layer; I won't re-run it.

The half that gets missed is more decisive: just-in-time lookup cannot solve deprovisioning at all, because deprovisioning is not a read problem.

JIT enrichment makes your view of a user fresh at the moment they authenticate through you. It does nothing about the account sitting in the customer's document-management system, which has its own user table, its own local sessions, and its own API keys. That account does not become stale-free because you looked something up during a login that, by definition, is not happening — the user was terminated. The stale state lives in someone else's database, and the only way to change someone else's database is to send them a write. There is no read-time solution to a write-shaped problem.

Which is worth stating as a rule: JIT lookup addresses attribute freshness; only push addresses lifecycle. They are not competing designs for the same problem, and a team that adopts JIT thinking it has retired its SCIM obligation has retired the cheap half and kept the expensive one.

This is the reasoning behind a design stance I should disclose an interest in: ClavionX — the platform I work on — declines to execute customer code or make synchronous outbound calls in the authentication path at all, and its sanctioned answer for "I need an attribute from an external system in the token" is to synchronize it ahead of time so that at login it is a local read. The runtime never synchronously calls the control plane either; configuration arrives as projections built from events, and missing state fails secure rather than reaching back for a fresh read. I bring it up because it makes the trade-off legible rather than because it's the only way to do it: pre-computation buys bounded login latency and availability independent of third parties, and the bill for that choice is precisely this article. Staleness windows, queue depth, reconciliation, drift, and fourteen vendor dialects are the cost of not fetching at login. It is a real cost, and pretending otherwise would be dishonest. The argument is that it's the cheaper of two bills, and that it lands on the axis you can engineer — throughput and lag — rather than the one you can't, which is somebody else's uptime during your users' logins.

Where the cost actually concentrates

Scale Dominant cost The thing that bites
Small (< 20 tenants, 1–3 targets) 100% engineering — building the first two connectors Discovering that "SCIM support" in the vendor's docs means a partial /Users implementation with no /Groups
Mid (~200 tenants, 6–14 target types) Dialect maintenance + on-call, ~$12.5k/month The reorg storm arrives; queue design becomes a security control; drift becomes real and the reconciler gets built in a hurry
Large (2,000+ tenants) Fairness and support surface One tenant's 40,000-user onboarding backfill starving 1,999 others; N×M distinct ways for a customer to say "sync is broken"

The large-scale line deserves a number. Onboarding a 40,000-user tenant across 6 targets is 240,000 create operations; at 100 req/min per channel that's 6.7 hours per channel even running all six in parallel — and with an unbounded global worker pool, those six channels will happily consume every slot you have for most of a business day. Fair queueing across tenants stops being an elegance and starts being the difference between a new logo onboarding and every existing customer's deprovisioning SLA. Weight the scheduler by tenant, cap per-tenant concurrency, and let the big backfill take twelve hours instead of seven.

Notice what does not appear in that table at any scale: compute, storage, bandwidth. This subsystem never gets expensive in the way infrastructure gets expensive. It gets expensive in headcount and it gets dangerous in latency, and both of those are invisible to the tooling most organizations use to find cost.

What to measure

Lifecycle-event dispatch latency, p99, split out from everything else. Not "sync lag." The p99 time from termination event to dispatch of the deprovision operation, per target. If you measure one number from this article, measure this one — it is the number you are selling.

Queue depth by lane, and the ratio between them. A bulk lane 300,000 deep is fine. A priority lane 50 deep is an incident.

Events per source change. Your amplification factor. If one HR attribute change produces 40 outbound operations, your problem is dynamic group rules, not capacity.

Per-target concurrency occupancy. The share of your worker pool held by each target. When one vendor holds 60% of your slots for 8% of your traffic, you have found your bulkhead candidate.

Rate-limit rejection rate per channel, and time-to-drain. 429s are not errors; they are your capacity signal. Time-to-drain at current rate is the honest version of your deprovisioning SLA on any given day.

Drift found per reconcile, by cause. If it's not trending down after you fix each cause, your delta path has a correctness bug, and the reconciler is quietly papering over it forever.

The closing thought

SCIM is the subsystem that never makes anyone's architecture diagram look impressive and never makes anyone's cloud bill look bad. A hundred and seventy-five dollars a month, four requests a second, and no error budget to burn.

And it is, at most enterprises, the single largest gap between what identity infrastructure promises and what it delivers — because the promise is a latency guarantee on a security event, and the delivery mechanism is a queue shared with an unbounded stream of somebody's phone extension changing.

The cheap fix is not more capacity. It is the recognition that 0.8% of your events carry all of the risk and none of the volume, and that they should therefore never, under any circumstances, wait behind the other 99.2%.

Rename that department any day you like. Just not in the same lane.