The Cost of Multi-Region Identity
The business case was one slide, and it was a good slide.
Thirty percent of the user base was European, authenticating against a US-East identity stack. Login felt slow abroad — measurably slow, 900ms click-to-dashboard against 380ms domestically. A second region in Frankfurt would fix it, improve the availability story for the enterprise deal in the pipeline, and get ahead of the residency question that a German customer's legal team kept raising. Three problems, one project. The slide had three bullets and everyone in the room nodded at a different one.
They shipped it in five months. Afterward, three things were true.
European click-to-dashboard went from 900ms to 660ms — real, but a third of what the slide implied, because most of the gap was never in the origin. Measured availability for the year went down, from 99.95% to 99.93%, entirely on the back of two incidents that could not have happened in a single-region system. And the German customer's legal team read the architecture document, saw that EU user records were being replicated to us-east-1 for failover, and said no — the thing they had actually needed was for the data to not be in two places, which is a different project with a different shape.
The interesting failure here isn't engineering. Every component worked. The failure is that "we need multi-region" is three unrelated purchases wearing one budget line, and they have almost nothing in common: different cost curves, different architectures, and — this is the part that surprises people — an ordering of value that is roughly the reverse of the order in which they get cited.
So let's price them separately.
The three purchases, and why conflating them is expensive
| Latency | Availability | Residency | |
|---|---|---|---|
| Nature of the requirement | Economic — buy as much as it's worth | Economic — buy as much as it's worth | Categorical — you comply or you don't sell |
| What it actually demands | A read path close to the user | A failover path that works when untested | Data that is absent from other regions |
| Correct topology | Edge termination, then read replicas | Warm standby with rehearsed, deliberate failover | Partitioning — separate stacks, no replication |
| Cost shape | Cheap at the edge, expensive at the origin | Fixed cost + a rehearsal budget most teams skip | N × fixed floor, near-zero replication cost |
| Typical error | Buying a write region to fix a network problem | Counting the nines you add, not the ones you remove | Assuming it means replication, and pricing a full mesh |
Read the third column carefully, because it contains the finding I'd most like to land: the only one of the three that is a genuine requirement is also, architecturally, the cheapest form of multi-region. Residency forces partitioning, and partitioning has no cross-region replication, no split brain, no global write consistency, and no failover logic. Teams price residency as though it means active-active because "multi-region" is a single word in their heads, then flinch at a number they didn't need to pay.
Meanwhile latency — the reason cited most often, and the one on most slides — is the worst buy on the page, and about eighty percent of it is available for two percent of the price from a component that isn't a region at all.
The assumptions, published so you can re-run it
Same price anchors as the rest of this series, so the numbers compose.
| Parameter | Value | Note |
|---|---|---|
| Peak logins | 50/s (≈1.5M/month) | Mid-sized B2B SaaS, few hundred enterprise tenants |
| Token/session validations | 400/s average | Across all applications |
| Concurrent sessions | 500,000 | |
| Audit events per login | 20 | Per 13.1; ~86 GB/day at 1 KB |
| European share of users | 30% | |
| US-East ↔ Frankfurt RTT | 90ms | Good day, warm path |
| vCPU | $25/vCPU-month | Blended |
| Cross-region transfer | $0.02/GB | AWS-shaped |
| Log ingestion + indexing | $0.50/GB | Negotiated mid-range |
| Fully-loaded engineer | $16,700/month | $200k/yr |
| Target CPU utilization | 40% | Because queueing |
| Single-region identity infra | ~$6,500/month | $3,900 fixed floor + $2,600 variable |
That last row needs unpacking, because it's the axis the whole article turns on.
The fixed floor of one identity region, assembled from the components the earlier pieces in this series priced individually — the four data stores each at their HA minimum, plus the runtime:
| Component | Minimum viable footprint | Monthly |
|---|---|---|
| Operational database (record of intent) | 3-node HA, cross-AZ | $700 |
| Session store | 3-node in-memory cluster (13.2) | $450 |
| Runtime auth tier | 6 × 4 vCPU, 3 AZs | $600 |
| Introspection / validation tier | 6 instances + cache (13.4) | $900 |
| Audit ingest + hot index | pipeline, 30-day hot tier | $600 |
| Load balancers, NAT, DNS, baseline egress | $250 | |
| Observability agents, log shipping, secrets | $400 | |
| Fixed floor per region | ~$3,900 |
Every number in that table is the price of existing, not the price of serving. It is the same $3,900 at 5 logins per second and at 50. Both of the previous cost pieces in this series arrived at that same shape independently — introspection's $900 HA floor, the session cluster's $300–900 floor — and multi-region is the decision where the floor stops being a footnote and becomes the entire argument.
Because a second region does not add a fraction of a stack. It adds a whole one.
Purchase one: latency, and the hop you can't buy back
Start with the honest version of the saving, because the slide version is always wrong in the same direction.
A login is not a request. It's a redirect chain — five browser round trips and a back-channel call, as the latency budget walks through. The naive model says a European user talking to a US origin pays 90ms of ocean per hop, therefore a Frankfurt region saves five times ninety, therefore 450ms. That model is wrong twice, in opposite directions, and the errors don't cancel.
It overstates, because most of the WAN cost is connection setup, and connection setup is not a regional problem. The first contact with your identity domain from a cold client is a DNS lookup, a TCP handshake and a TLS handshake before a single application byte moves — three round trips, 270ms, of which zero is your server. An edge PoP terminating TLS in Frankfurt collapses all of it: DNS resolves to a nearby anycast address, TCP and TLS complete in ~10ms locally, and only the request itself rides a warm, pre-established backhaul to the origin. You keep one crossing per hop instead of three.
It understates, because it counts only the login. The login happens once per session. Token validation, session checks, userinfo, JWKS fetches and refreshes happen thousands of times per session, and if any of them cross an ocean, that's where the aggregate user-latency actually lives.
Put both corrections in and price the three things you could buy:
| Edge TLS termination | Read-replica region (validation path only) | Full active-active write region | |
|---|---|---|---|
| What's deployed | Anycast PoPs, TLS offload, warm backhaul | Stateless validation pods + replicated cache + JWKS; no writes | Everything in the fixed-floor table, twice |
| Monthly cost | ~$150 | ~$1,200 | ~$14,000 (derived below) |
| Login latency saved (EU user) | ~180ms | ~180ms (inherits edge) | ~270ms |
| Validation latency saved | 0 | 90ms per cross-region validation | 90ms |
| New failure modes introduced | ~none | one (stale replica, read-only) | many (see next section) |
Now convert each into the same unit — aggregate user-seconds of latency removed per month per dollar — which is the only way to compare a saving that happens once per session with one that happens continuously.
EU logins per month: 1.5M × 30% = 450,000. EU cross-region validations per month, assuming a gateway or introspection design where a cache miss actually crosses: 400/s × 30% × 2% miss × 2.6M seconds ≈ 6.2M.
| Purchase | Seconds saved / month | Monthly cost | Cost per user-hour of latency removed |
|---|---|---|---|
| Edge TLS termination | 450,000 × 0.18 = 81,000 s (22.5 h) | $150 | $6.70 |
| Read-replica validation region | 6.2M × 0.09 = 558,000 s (155 h) | $1,200 | $7.70 |
| Full write region (marginal, on top of the above) | 450,000 × 0.09 = 40,500 s (11.25 h) | $12,800 | $1,140 |
That's the whole latency argument in one row: the incremental latency a full write region buys, over an edge PoP and a read replica you could have bought instead, costs about 150× more per unit of user time.
Two caveats I owe you, because both change the numbers and one of them may change your answer.
If you're on self-validating JWTs, the read-replica row largely evaporates — resource servers validate locally and never cross a region, which is precisely the geographic argument for JWTs that 13.4 makes. What remains is JWKS fetches, session-status checks, refresh calls and userinfo, which is a smaller number but not zero. And if your traffic is one-request-per-token — webhooks, batch integrations — the miss rate isn't 2%, it's 100%, and the read-replica row gets an order of magnitude better rather than worse.
Second: this analysis prices the mean. Regional distance also fattens the tail, and a 90ms RTT with a 4% retransmit rate on a congested transatlantic path produces p99s that are not 90ms worse but 400ms worse. If your complaint is "login is occasionally unusable abroad" rather than "login is consistently slower abroad," you have a tail problem, and the honest answer is that a region does fix tail problems that an edge PoP cannot — because the edge only removes handshake round trips, and the tail lives in the crossing itself. That's a legitimate reason to buy a region. It is a different reason than the one on the slide, and it should be argued with a p99 histogram rather than a mean.
The general form, which is worth stating plainly: a write region optimizes the rarest flow you serve. Login is the one identity operation a user performs once per session; everything else in the identity workload is reads, at three orders of magnitude more volume, on a path that replicates cheaply. Buying a full region to make login faster is spending the expensive part of your budget on the low-volume path — and identity is mostly read traffic is not just an architecture observation. It's a purchasing one.
Purchase two: availability, and the nines you remove
The availability case is always made forward: we're at 99.95%, a second region takes us to 99.99%, that's 21.9 minutes a month down to 4.4. It's arithmetic, it's on a slide, and it is wrong in three separate places.
Error one: most of your downtime isn't regional
A second region only helps with failures that are confined to one region. Go through the incident log of any identity platform that's been running for three years and classify the downtime minutes by cause. The shape is remarkably consistent:
| Cause | Share of downtime minutes | Does a second region help? |
|---|---|---|
| Bad deploy or config change | 35% | No — it propagates to both |
| Schema migration, data corruption, poison record | 15% | No — it replicates |
| Dependency failure (upstream IdP, KMS, expired cert, DNS) | 20% | Rarely — certs and DNS are global |
| Overload / capacity / a tenant's login flood | 10% | Partially |
| Single-region infrastructure fault (AZ, network, provider) | 15% | Yes |
| Operator error during a maintenance window | 5% | No |
Roughly 15–20% of your downtime minutes are addressable by a second region. Everything else is software you wrote, config you propagated, or a dependency that is itself global. This is the single most-skipped step in the multi-region business case: teams apply the availability improvement to 100% of their downtime when it applies to a sixth of it.
Error two: failover is a probability, not a property
Having a second region and successfully cutting over to it under duress are different claims. Failover for identity is deliberately non-automatic in most sane designs — automatic failover on a partitioned identity system produces two regions issuing tokens and revoking into a void, which is worse than being down — so the path is: detect, decide, execute, verify. Call it 15–40 minutes end to end when it goes well.
And the code that executes it runs once a year, in an exercise, if at all. It is, structurally, the least-tested code you own, invoked exclusively during your worst incidents, by whoever happens to be on call. Empirically, unrehearsed annual failover succeeds first-try somewhere around 60% of the time; quarterly-rehearsed failover gets to around 90%. The failure modes are mundane and identical everywhere: a secret that only existed in the primary, a DNS TTL nobody shortened, a replica that was never actually promotable, a runbook whose first step references a console that requires SSO to the system that is down.
That last one deserves its own sentence, because it's specific to us: identity failover has a circular dependency risk that other systems don't. If your operators authenticate to the cloud console through the identity platform that is currently down, your failover procedure requires the thing it exists to restore. Every identity team should have a break-glass path that does not traverse its own product, and most discover they don't have one at 3am.
Error three: multi-region adds failure modes that single-region cannot have
New, permanent, and entirely self-inflicted:
- The global routing tier is now in front of 100% of traffic. Whatever decides which region serves a tenant — DNS policy, a routing table, an edge worker — is a new component with global blast radius. You have not eliminated the single point of failure; you have moved it from a region to a router, and the router is smaller, newer, and less rehearsed than the region was.
- Replication lag produces authentication anomalies, not just stale reads. A password changed in Frankfurt and not yet visible in us-east means the old password still works. A user enrolls MFA in one region and is challenged as unenrolled in the other. A session created in one region and validated in the other doesn't exist yet. These present to users as "identity is broken," to support as unreproducible, and to your dashboards as nothing at all — every component is green.
- Split brain during a partition. Two regions accepting writes to the same identity, reconciled later by a policy someone wrote in a hurry. For a user profile that's a merge conflict. For a credential or a revocation, it's a security decision.
- Config propagation skew — its own section below, because it's the one that generates the most support tickets.
Putting a number on it
Expected downtime minutes avoided per month:
21.9 min × 0.18 (regional share) × 0.6 (unrehearsed failover success)
× 0.7 (fraction of an incident actually truncated by a 20-min failover)
≈ 1.7 minutes/month
Expected downtime minutes added in year one, from the new failure modes: two self-inflicted multi-region incidents at 45 minutes is 90 minutes a year — 7.5 minutes/month.
Net year one: −5.8 minutes/month. Availability goes from 99.95% to 99.93%, which is exactly what happened to the team in the opening story, and it is the ordinary outcome rather than the unlucky one.
Year three, with quarterly rehearsals, matured routing and a failover path someone has actually driven:
avoided: 21.9 × 0.18 × 0.9 × 0.8 ≈ 2.8 min/month
added: one incident/year at 25 min ≈ 2.1 min/month
net: +0.7 min/month → 99.955%
You bought 8.4 minutes of annual downtime for roughly $168,000 a year: about $20,000 per minute of downtime avoided. Whether that's a good trade is a real question with a real answer — if you're an exchange, a payments processor, or the SSO in front of a hospital's clinical systems, a minute is worth far more than $20,000 and you should buy it without further discussion. If you're B2B SaaS with a 99.9% SLA whose credit is 10% of a monthly invoice, run the number: the credit exposure on 21.9 minutes is almost certainly a fraction of the region.
The uncomfortable corollary: the availability case only becomes positive if you fund the rehearsal. Quarterly failover drills — six engineers, six hours, four times a year — cost about $14,000 a year in engineer time and are the difference between p=0.6 and p=0.9, which is the difference between a negative and a positive expected value on a $168,000 purchase. Teams reliably buy the region and skip the drill, which is buying the asset and declining the thing that makes it work.
Purchase three: residency, where the cost is an input
The first two purchases are optimizations: you decide how much they're worth and buy that much. Residency isn't. A contract clause saying EU personal data does not leave the EEA is not a latency target you can partially hit. You comply, or you don't sign, or you sign and are non-compliant. The cost is a hard input to the deal, not a variable you're optimizing.
Which is why the interesting question about residency is never whether, but what shape — and here the intuition that "multi-region means replication" costs teams real money.
Residency demands absence. The requirement is not that a copy exists in Frankfurt; it's that a copy does not exist in Virginia. Replication topologies are built to create copies everywhere; that is their entire purpose. So the correct architecture for residency is the one the architecture companion calls tenant partitioning — each tenant lives in exactly one region, and residency becomes a property of the topology rather than a filter someone can forget.
Compare the two, line by line, for a two-region deployment:
| Line item | Active-active replication | Partitioned stacks |
|---|---|---|
| Fixed floor | 2 × $3,900 | 2 × $3,900 |
| Capacity sizing | Each region sized for 100% of load (it's a failover target) | Each sized for its own tenants |
| Cross-region replication traffic | Continuous, all state classes | ~None — only the routing table |
| Global consistency problem | The whole dataset | A small, slow-changing tenant→region map |
| Split brain | Possible, must be designed for | Impossible — one writer per tenant |
| Failover logic | Required, complex, untested | None (or in-region only) |
| Revocation propagation window | A security window you must measure | Zero — revocation is local |
| Audit consolidation | Central aggregation | Federated query, or none |
| Cost at our scale | ~$14,000/month | ~$8,500/month |
Partitioning is roughly 40% cheaper and structurally simpler, and it is the topology the only categorical requirement actually asks for. That's the inversion worth carrying out of this article: the reason with the strongest claim on your budget is served by the cheapest architecture, and the reason with the weakest claim is the one that demands the most expensive one.
Two things partitioning does cost you, stated honestly. Roaming users pay a full cross-region round trip on every hop of their login when they travel, which is the trade 5.4 covers and which people tolerate better than architects expect. And you lose capacity fungibility: each region's headroom is stranded in that region, so a partitioned deployment carries more total idle capacity than a pooled one. At our scale that's the $2,600 of variable capacity, provisioned twice at 60% utilization instead of once at 75% — call it $900/month of stranded headroom, which is the price of the guarantee.
The break-even tenant count, which sales should know
Residency has a number that almost never gets computed and should appear in every deal review: the fixed floor divided by the margin on the tenants who require it.
A dedicated EU stack costs $3,900/month floor plus perhaps $1,200 of variable capacity and $2,000/month of amortized operational burden — call it $7,100/month, $85,000/year, to serve however many EU-resident tenants you have. If your enterprise tier is $3,000/month at a 70% gross margin, each tenant contributes $2,100:
| EU-resident tenants | Monthly contribution | Region cost | Net |
|---|---|---|---|
| 1 | $2,100 | $7,100 | −$5,000 |
| 3 | $6,300 | $7,100 | −$800 |
| 4 | $8,400 | $7,100 | +$1,300 |
| 10 | $21,000 | $7,100 | +$13,900 |
The first residency deal loses $60,000 a year, and it is still frequently correct to take it — because it's the deal that unlocks the region, and the region is what makes the next nine sellable. But the company should know it's a market-entry investment rather than a profitable contract, and the number to defend it with is "we break even at four," not "we should probably support EU."
The variant that goes wrong: a single mid-market customer with a residency clause and a $600/month contract. That deal costs $85,000/year to service and there is no queue behind it. The right answer is a price that reflects the floor, or a polite no.
The line items nobody budgets
Five costs that appear in no business case I have ever read, ordered by how badly they're mis-estimated.
The replication egress is the smallest number on the page
This is the line everyone warns about, and it's a rounding error. Our audit stream is 86 GB/day. At $0.02/GB cross-region that's $1.72/day — $52/month. Even at ten times the scale it's $520. Session state should stay regional and doesn't replicate at all; user records are small and change rarely. Replication bytes are not the multi-region cost, and the roadmap bullet that says they are — mine did — is repeating folklore.
What is expensive is what happens to those bytes when they land. 13.1's central finding was that storage is cheap and searchability is expensive; multi-region takes that and multiplies it. If each region ships its audit stream into a central log platform that charges on ingestion, you pay the indexing price a second time on data you already indexed regionally: 860 GB/day at $0.50/GB is $13,000/month — 250 times the transfer cost of the same bytes.
So the audit rule for multi-region is: move the bytes, don't re-index them. Land regional streams in regional object storage with a shared schema and a federated query layer over the top, and accept that a cross-region investigation takes ninety seconds instead of two. You are paying for search latency during incidents you have a handful of times a year, at a price of $13,000 a month.
And in the residency case you may not have a choice at all: if EU audit records are personal data that cannot leave the EEA — which they usually are, since they contain identifiers, IPs and device fingerprints — then central aggregation is not merely expensive, it's the compliance violation. A compliant user database attached to a non-compliant log pipeline is the most common residency failure there is. Federated query stops being a cost optimization and becomes the only lawful design, which is a rare and pleasant case of the cheap answer and the correct answer being the same one.
The fixed floor makes multi-region regressive
Take the $3,900 floor and divide it by logins.
| Platform | Logins/month | Single-region cost/login | Two-region cost/login | Increase |
|---|---|---|---|---|
| Small (5 logins/s peak) | 150,000 | $0.029 | $0.086 | +197% |
| Our scenario (50/s) | 1,500,000 | $0.0043 | $0.0093 | +116% |
| Large (500/s) | 15,000,000 | $0.0011 | $0.0016 | +45% |
| Very large (2,000/s) | 60,000,000 | $0.00055 | $0.00061 | +11% |
Cost per login roughly triples for the small platform and rises a tenth for the very large one, from the identical architectural decision. Multi-region is a regressive tax: it is priced almost entirely by the number of stacks you run, not by the traffic you serve.
Which inverts a piece of advice you hear constantly. "Start single-region, go multi-region as you grow" is usually said as though multi-region is the advanced move you graduate into. The economics say something stronger and more specific: multi-region is only affordable once you're large, because the entire cost is a floor that scale amortizes and small scale cannot. A 20-person company adding a second region for latency is making the most expensive possible per-login decision at the point in its life when it can least absorb it.
Config propagation: the ClavionX case, and what it costs
In a design with strict runtime/control-plane separation — ClavionX's ADR-0002, where the runtime never synchronously calls the control plane and works only from projected state, failing secure when the projection is missing — a config change is not a write. It's an event that must be projected into every runtime region independently.
flowchart LR
CP["Control plane<br/>(admin writes here)"] -->|event| P1["us-east projection<br/>lag: 1.2s"]
CP -->|event| P2["eu-central projection<br/>lag: 3.8s"]
CP -->|event| P3["ap-southeast projection<br/>lag: 47s ⚠"]
P1 --> R1["Runtime: serving"]
P2 --> R2["Runtime: serving"]
P3 --> R3["Runtime: fail-secure<br/>on missing state"]
Single-region, that lag is a second and nobody thinks about it. Across N regions it becomes N independent propagation windows, and three costs fall out:
A support cost. "I disabled that user ten minutes ago and they just logged in" and "I changed our MFA policy and it isn't live in Sydney" become recurring ticket categories. At even four such tickets a month, each consuming an hour of engineering plus support time, that's a few thousand dollars a year of pure coordination overhead — and each one is genuinely hard to diagnose, because every dashboard is green.
An availability cost you chose on purpose. Fail-secure on missing projected state is the right call — a runtime that serves stale security config is worse than one that refuses — but it converts replication lag into request failures in whichever region was furthest behind, which is by definition the region you were watching least. The mitigation is the per-region freshness watermark: each region tracks how stale its projection is and degrades deliberately when it exceeds a stated window. That's engineering work that only exists because you have more than one region.
An admin UX cost. Every configuration screen now owes the operator an answer to "is this live everywhere yet?" Either you build per-region propagation status into the admin surface, or your operators learn to wait an unspecified amount of time and refresh, which is how confident administration turns into superstition.
Keys, certificates and the operations multiplier
13.6 established that rotation cost is set by key class, not cadence — automated discoverable keys cost about a dollar per operation, coordinated bilateral ones cost weeks of counterparty lead time. Multi-region multiplies the operation count by N in the class that can least absorb it. Per-region signing keys keep material in-jurisdiction and give you a real blast-radius boundary, and they also mean every validator's trust configuration, every SAML federation, and every customer's pinned certificate is now an N-way relationship. The automated class shrugs. The coordinated class, whose p99 counterparty change lead time is measured in months, does not.
The rest of the operational multiplier is unglamorous and adds up: on-call coverage that must now reason about which region an incident is in, a deploy pipeline whose wall-clock doubles or whose canary value halves, every schema migration permanently constrained to be forward- and backward-compatible across a window when regions run different versions, and a staging environment that either mirrors the topology (expensive) or doesn't (and therefore never tests the failure modes that only exist in the topology). Budget 0.3–0.5 of an engineer per additional region, which at our anchor is $60,000–100,000 a year and is larger than every infrastructure line in this article except the fixed floor.
Testing costs grow faster than region count
Failover paths grow as N(N−1). Two regions is two paths; three regions is six, plus the question of what happens when the region you're failing to is the one that's degraded. Nobody tests six paths. Most teams test zero and discover the p=0.6 number the hard way, during the only incident where it mattered.
The honest budget line is the drill, priced above at ~$14,000/year, and it is the single highest-return item in this entire article — because it is the only expenditure that moves the availability calculation from negative to positive.
The ladder, and where to stop on it
Multi-region is not a binary. It's a ladder with four rungs, and the industry habit is to skip to the top because that's the rung with a name.
| Rung | Cost/month | Buys | Stop here if |
|---|---|---|---|
| 1. Edge termination | ~$150 | ~80% of the achievable login latency improvement; zero new failure modes | Your case is latency and you can't articulate a p99 problem |
| 2. Read-replica region (validation, JWKS, discovery — no writes) | ~$1,200 | Continuous read-path latency; survives primary write outage for validation | Your case is latency or graceful read-only degradation |
| 3. Partitioned stacks (tenant pinned, no replication) | ~$8,500 | Residency as a topology guarantee; no split brain, no failover logic | Your case is residency — which it usually is |
| 4. Active-active write region | ~$14,000 + drills | Low-latency writes globally for the same tenant | You genuinely need it and can state your revocation window as a measured number |
Rung 2 deserves more attention than it gets, because it does something the availability section didn't credit: a read-replica region keeps token validation and session checks serving when the primary's write path is down. Most identity outages are write-path or control-plane outages, and what identity should do when its database is down argues that continuing to validate existing tokens while refusing new logins is usually the correct degraded mode. Rung 2 buys that mode geographically, for 8% of rung 4's price, without a failover decision, without split brain, and without any of the anomalies that replication of writable state produces.
The thresholds, stated as numbers you can argue with:
- If your identity infrastructure bill is under $10,000/month, a second full region at least doubles your cost per login. Don't, unless a contract requires it.
- If your case is latency, buy rung 1 first and re-measure. If the remaining gap is under 200ms on the mean, stop. If it's a p99 problem, keep going — but bring the histogram.
- If your case is availability, price the drill before the region. Without quarterly rehearsed failover the expected value is negative, and you will be able to prove it from your own incident log in eighteen months.
- If your case is residency, partition — and compute the break-even tenant count before the deal closes, not after.
What to measure
- Your downtime minutes by cause, classified as regional or not. If you can't produce this from your incident log, you cannot compute the availability case at all, and every nine on the slide is imaginary.
- Measured failover time and success rate, from a rehearsal with a date on it. Not the runbook's estimate. The stopwatch. It's the term the whole expected-value calculation is most sensitive to.
- The share of your identity latency that is connection setup rather than crossing. One packet capture. It tells you whether an edge PoP solves your problem for $150.
- Cost per login, per region, with the fixed floor separated from the variable. The ratio tells you where you sit on the regressive curve and therefore whether the decision is affordable at all.
- Per-region projection freshness, as a watermark with an alert. Staleness that isn't measured is staleness that isn't bounded — and in a fail-secure design it's also an outage waiting in the region nobody is watching.
The closing thought
The team from the opening story ended up somewhere they could have started. They kept the edge PoPs, which had delivered most of the latency win at a rounding-error price. They demoted Frankfurt from an active-active peer to a read-replica region serving validation and discovery, which halved its cost and removed every anomaly that had caused their two self-inflicted incidents. And they built a separate, small, deliberately unreplicated EU stack for the German customer — the thing the contract had actually asked for, which turned out to be the cheapest piece of the entire program.
Three purchases, three architectures, one of which they needed. The slide had them as three bullets under one heading because that's how they arrived — three people wanting three different things, agreeing on a word.
Which is the pattern this series keeps finding. 13.6 ended on a single policy number being a uniform price applied to a non-uniform population. This is the same error one level up: a single architecture applied to three unrelated requirements, priced as though they were one. The fix isn't cleverer engineering. It's refusing to let "we need multi-region" stand as a requirement, and asking the only question that separates the three — what specifically breaks if we don't? Slow login abroad, an outage you can't survive, and a contract you can't sign are three different sentences with three different price tags, and only one of them is a fact about the world rather than a preference about milliseconds.