Who Owns the Incident When Every Layer Worked Correctly?
Over eleven days, an attacker tried 4.2 million passwords against 380,000 accounts at a company I'll leave unnamed. They got into 1,900 of them. The pattern was textbook credential stuffing: one attempt per account, rotating residential proxies, human-plausible timing, user agents drawn from a reali
Over eleven days, an attacker tried 4.2 million passwords against 380,000 accounts at a company I'll leave unnamed. They got into 1,900 of them. The pattern was textbook credential stuffing: one attempt per account, rotating residential proxies, human-plausible timing, user agents drawn from a realistic distribution.
The postmortem ran for three weeks. Here is what each team correctly reported.
The WAF team: no rule was violated. The traffic was within volumetric thresholds, distributed across tens of thousands of source addresses, and contained no attack signatures. The WAF is configured to block known-bad patterns and it blocked every known-bad pattern it saw.
The API gateway team: rate limits were enforced as specified — 100 requests per minute per IP. No source came close. The limits were the ones documented, reviewed, and signed off in the last security review.
The identity team: every one of the 4.2 million attempts was authenticated correctly. 4,198,100 were rejected because the password was wrong. 1,900 were accepted because the password was right. The system's job is to verify credentials, and it verified 4.2 million credentials with no errors.
The application team: every session that reached them carried a valid token from the identity provider. There is no signal at their layer that distinguishes a legitimate login from a successful stuffing attempt, and they trust the identity layer's assertion by design — that's what a trust boundary means.
Every statement is true. Every team behaved competently. No configuration was wrong relative to its stated purpose. And 1,900 accounts were compromised.
The incident review ended with no owner and three action items, all assigned to "security" as an abstraction. Eight months later a near-identical attack succeeded, and the postmortem observed that the action items from last time had not been completed, because they had never belonged to anyone with a backlog.
This is a structural consequence, not a failure of diligence
The reflex is to read that story as a story about lazy teams or bad management. It isn't, and reading it that way guarantees you repeat it.
Layered architecture is correct. Separating the edge from the gateway from the identity service from the application is how you get independently deployable, independently reasoned-about, independently ownable components. Clear component ownership is one of the highest-value organizational patterns in software. Nobody in that story should be criticized for having a well-defined layer with a well-defined responsibility.
But there's an arithmetic to it that nobody plans for. Each layer is given a local responsibility — block known attacks, enforce rate limits, verify credentials, trust the token — and each discharges it. The security outcome, though, is a property of the composition. And composition, by construction, belongs to no component.
Put formally: you can decompose a system into components with complete ownership coverage and still have zero ownership coverage of the system's emergent properties. Every layer is 100% owned. The composition is 0% owned. This is not a gap someone forgot to fill; it's a category that the decomposition itself does not produce.
The tell is a specific sentence in a postmortem: "working as designed." When every layer worked as designed and the outcome was bad, you are not looking at a bug. You are looking at a specification gap at a seam — and seams have no team.
Why this is worse in identity than almost anywhere else
Three properties of identity systems make them the canonical case.
The signal is distributed but the decision is local. The only place the attack was visible was in the aggregate: 380,000 distinct accounts, one attempt each, one failure each. The identity service saw 4.2 million individually unremarkable events. The WAF saw traffic from tens of thousands of individually unremarkable addresses. The evidence existed, distributed across two layers, and neither layer's job description included joining it. Detecting credential stuffing requires correlating across the dimension that each layer partitions by — which means it is structurally invisible to every layer acting alone.
Success and failure are indistinguishable downstream. Once the identity service says "this credential was valid," the application has no basis to disagree. A successful stuffing attempt and a legitimate login are the same event with the same shape. The trust boundary that makes the architecture clean is also what prevents the layer with the most business context from participating in the decision.
Each layer's SLIs are locally green during the incident. The WAF's block rate was normal. The gateway's rejection rate was normal. The identity service's error rate was zero — every request was processed correctly. Login success rate dropped slightly, which is exactly the kind of small change that gets attributed to a marketing campaign or a client release. There was no dashboard, anywhere, whose needle moved in a way that demanded attention, because the metric that would have moved — authentication failure rate across distinct accounts per source ASN — is a cross-layer metric and therefore belonged to nobody's dashboard.
The fixes that don't work
Worth naming, because these are the reflexes and each has a real cost.
"Collapse the layers." Put all the security logic in one place with one owner. This trades a clear problem for a worse one: a monolithic security layer that every team must coordinate through, that becomes a deployment bottleneck, and whose failure is total rather than partial. The layers exist for good reasons. Ownership ambiguity is a real cost of composition, and the correct response is to pay it deliberately, not to abandon composition.
"Make security own it." This is what that postmortem did, and it's the most common outcome. It fails because a central security team has no authority to change the WAF rules, the gateway limits, or the identity service's behaviour — they can only file requests against four backlogs owned by teams with their own priorities. Accountability without authority produces reports, not fixes. If you do this, you must also give that team either write access or a prioritization mandate, and almost nobody does.
"Add a layer." A detection and response system above all four. Sometimes right — but if you add it without answering the question this article is about, you've created a fifth layer with a local responsibility and the same seams, now with more of them.
"Have every team think about the whole system." A worthy sentiment that survives no contact with a quarterly planning cycle. Teams are measured on their layer. Asking for cross-layer vigilance without a structure that rewards it is asking people to do unpaid work with no clear success criterion.
What actually works
Four things, in rough order of leverage. The pattern across all of them: make the composition a first-class object with a name, an owner, and instrumentation — because you cannot assign responsibility for something that has no representation in your organization or your telemetry.
Own outcomes alongside components
The move is not to reassign the components. It's to add a second ownership dimension.
Components have owners: the WAF team owns the WAF. Add outcome owners: someone owns "account takeover rate," someone owns "a legitimate user can log in," someone owns "credential-based attacks are detected within an hour." These are cross-cutting, they don't map to any layer, and they are the things the business actually cares about.
An outcome owner's job is not to implement everything. It's to be the person for whom the outcome degrading is their problem — to notice, to convene, to allocate the work across the layers, and to be accountable for whether the number moves. In SRE language, it's an SLO whose error budget spans four services, and someone owns that budget.
The critical design detail: the outcome owner must have either budget or prioritization authority over the component teams. Without one of those, you've recreated "make security own it" with a new title. This is the part that requires an executive decision rather than an architectural one, which is precisely why it usually doesn't happen.
Build the cross-layer metric first, because it doesn't exist by accident
Every layer instruments what it does. Nobody instruments what happens across layers, because it requires joining data owned by different teams in different systems with different retention.
For the story above, the metric that would have fired on day one is: distinct accounts with failed authentication, grouped by source ASN, per hour. No single layer can compute it. The identity service has the accounts and failures but not enriched network context. The edge has the network context but not the account outcome. It requires a join, which requires a correlation ID that survives every hop, which requires all four teams to agree on a header and propagate it.
That agreement is the actual deliverable, and it's the highest-value thing to come out of an incident like this. Not a rule change — a shared identifier and a shared event stream that makes cross-layer questions answerable at all. Everything else is downstream of it.
A rough sufficiency test: pick a user and a five-minute window, and ask whether you can reconstruct every decision every layer made about their requests, in order, from one query. If you can't, no cross-layer detection is possible, and no amount of ownership clarity will substitute.
Write down the assumptions each layer makes about the others
The gap in that incident had a precise location. The gateway's rate limits were designed against one threat model (a single source hammering one endpoint). The attack used a different one (many sources, one attempt each). Nobody had written down which threats the gateway's limits were intended to address, so nobody could notice that a threat had fallen between the gateway's model and the identity service's.
The artifact that closes this is unglamorous and cheap: a per-layer statement of what this layer defends against, what it explicitly does not defend against, and what it assumes another layer handles. One page per layer.
Its value is entirely in the third column. When the WAF's page says "assumes the identity layer detects credential-based attacks" and the identity layer's page says "assumes the edge detects distributed automated traffic," you have found the gap in a meeting rather than in an incident. This is the mundane version of a threat model, and it's the version that actually gets maintained, because each team only has to write about their own layer.
Review the seams on a schedule, with all owners present
Architecture reviews look at components. Seam reviews look at the spaces between them, and they need the four owners in one room because no smaller group can see the gap.
The agenda is short: for each pair of adjacent layers, what does the upper assume the lower has done? Has that assumption ever been verified? What has changed at either layer since the last review? Quarterly is enough. The output is usually one or two "wait, we both thought the other one did that" discoveries, which is a very high return for ninety minutes.
The uncomfortable part
Here's what makes this genuinely hard, and why I don't think there's a clean answer.
Clear component ownership is unambiguously good. It's why teams can deploy independently, reason locally, and be accountable for something they control. Every argument in favour of it is correct.
And its cost is that the sum of well-owned parts has no owner. Not because someone was careless — because ownership was defined over components and the outcome is not a component. You cannot get the benefits of decomposition without generating this class of gap. You can only choose whether to name it and staff it, or to discover it in a postmortem where four teams each correctly report that their layer worked.
The failure mode to watch for isn't a team saying "not my problem." It's four teams each saying, accurately and in good faith, "my layer behaved correctly." When you hear that sentence in an incident review, the finding is not about any layer. It's that you have a system whose behaviour nobody is responsible for, and the action item is not a config change.
It's a name on the outcome.