Identity Isn't Your Fraud Engine

The ticket was closed as "cannot reproduce," and that turned out to be the most alarming sentence anyone wrote that quarter.

A customer had been locked out of her account on a Tuesday in March. Correct password, correct TOTP code, then a screen saying her sign-in couldn't be completed. She called support twice, got nowhere, and eventually wrote to the regulator, which is how the ticket acquired a due date. In June an engineer sat down to answer the only question that mattered — why was this login denied? — and did the obvious thing: pulled the request from the logs and replayed it against staging with the same device fingerprint, IP, account, and time-of-day bucket.

It was allowed. Cleanly, in 41 milliseconds.

He tried production shadow mode. Allowed. He tried reconstructing the account state as of March. Allowed. The audit log had faithfully recorded what happened — decision=deny, reason=risk, risk_score=0.83, threshold=0.75 — and that record turned out to be worth almost nothing, because the thing that produced 0.83 had been retrained eleven times since March. The model that made the decision no longer existed. Its features came from a store whose 30-day aggregates had rolled forward. The threshold had been moved twice, once at 2am during an incident, by someone who has since left.

The company could prove that a decision had been made. It could not say what made it, could not repeat it, and could not tell a regulator why one of its customers had been denied access to her own money for four days. Not because anyone was negligent — the fraud team's model governance was genuinely good, better than the identity team's — but because a probabilistic system with a weekly release cadence had been placed inside a deterministic system whose value proposition is that you can reconstruct any decision it ever made.

That's the boundary this article is about. Authentication and fraud detection look like neighbours: they score the same login, consume overlapping signals, and both can say no. Underneath they have opposite engineering properties, and merging them fails in the same four ways every time.

Why the merge keeps happening

It isn't ignorance; every step is locally sensible.

The fraud team has the better apparatus. A feature store, an offline/online consistency story, an experimentation platform, labelled data, and someone whose full-time job is watching a precision-recall curve. The identity team has a rules file and a risk score somebody weighted by hand. When the question is "should this login be scrutinized more carefully," the fraud team is visibly better equipped to answer it, and they'll say so.

The signals genuinely overlap. Device, IP, ASN, velocity, time-of-day, account age — both systems want all of them, and building the same device-recognition pipeline twice is obviously wasteful. So the first proposal is always the reasonable-sounding one: the fraud service already computes this, let identity call it.

And the incident that triggers the conversation is always real. An account takeover ran through login and drained a balance, and the retro asks whether it could have been caught at the front door. The honest answer is sometimes, at a cost you haven't priced; the answer that gets written in the action items is "integrate fraud scoring into the authentication flow." Six months later there's a synchronous HTTP call to a model service inside the login handler, and nobody remembers agreeing to it.

Two systems, opposite properties

Put the two side by side and the incompatibility isn't a matter of degree. Almost every row inverts.

Property Authentication Fraud / abuse detection
Decision shape Binary, on the critical path Score, usually off it
Latency budget Tens of milliseconds, hard Seconds to minutes, soft
Determinism Same inputs → same output, forever Same inputs → different output after any retrain
Failure default Fail secure: deny Fail permissive: let it through, review later
Correctness metric Pass/fail against a spec ROC curve, precision at a chosen recall
Ground truth None. You never learn a login was the attacker Chargebacks, disputes, confirmed reports — 30–120 days late
Change cadence Quarterly, reviewed, protocol-constrained Weekly or faster, by design
Adversary loop Slow — protocols change on RFC time Fast — the adversary retunes against you continuously
Explainability Must state a reason a support agent can read "The model scored it 0.83"
Reconstructability Any decision, years later Best-effort, current model only
Context needed The account and the request The transaction, the counterparty, the money

The row that does the most damage is Failure default, because it's the one nobody notices until the fraud service is down.

Identity's fail-secure posture says: when in doubt, deny. Correct for a credential check — a password you can't verify is a password that didn't verify. But a fraud score you can't obtain isn't evidence of fraud, it's an absence of evidence, and applying the fail-secure default to it turns a fraud-vendor outage into a total authentication outage for a system that was up the whole time. Fail open instead and the control is decorative: present exactly when it isn't needed, absent whenever an attacker can knock over the scoring service or simply wait for someone else to. There is no correct answer, which is the tell. A dependency with no correct failure behaviour is a dependency in the wrong place — the same argument that rules out customer code in the auth path, arriving from a different direction.

The latency arithmetic

Take a 200ms budget for the credential-submission request — the number I'd defend — of which the server-side portion is maybe 60ms, most of it deliberately burned in the password hash. There's real room, but not much.

Now add a fraud call. People estimate this by asking the fraud team for inference latency, and the answer is encouraging: gradient-boosted trees over a few hundred features, 2–4ms. That number is honest and irrelevant, because inference is the cheapest part. The expensive part is assembling the feature vector, and a serious model's features live in four or five places: a counter store for velocity, a warehouse-derived table for 30- and 90-day behaviour, a device-graph service, a payments store for instrument history, and often a third-party enrichment call.

Assume you parallelise all of it, and each store is fast: p50 of 8ms, p99 of 80ms. The fan-out doesn't get you the p99 of one store, it gets you the p99 of the slowest of six independent calls:

P(all six under 80ms) = 0.99^6 = 0.9415

So 5.9% of logins wait more than 80ms on feature assembly alone, against a total server budget of 60ms. Your login p95 — not p99, p95 — is now set by a subsystem whose owners consider 80ms excellent, because in their world it competes with a payment authorization that has three seconds.

The usual response is a timeout: 50ms, then proceed. That's worse than it sounds, for a reason that isn't latency. A timeout converts your security control into a coin flip whose bias is set by unrelated infrastructure load, and it does so silently and non-uniformly: the logins most likely to breach the timeout are the ones arriving during a traffic spike, which is precisely when an attack is running. You have built a control that degrades exactly when it is needed, and the metric that would reveal it — timeout rate correlated with attack traffic — is one nobody graphs, because it lives between two teams' dashboards.

There's a second-order cost. Once the call is in the login path, the fraud service inherits the identity platform's availability requirement — change-controlled, on-call, several nines — while keeping its own cadence of a few deploys a week. Nobody tells the fraud team this has happened. They find out during their first bad deploy, when the incident channel fills with people who cannot log in to debug it.

The false-positive economics, in the currency your dashboard doesn't count

Here is the calculation that should end most of these conversations, and I've never seen it done in the meeting where the decision gets made. A 0.5% false-positive rate is a good model; fraud teams would take it happily. Run it against two different consumers.

On a fraud review queue. 40,000 transactions a day, 0.5% flagged in error = 200 reviews. At three minutes each that's ten analyst-hours: one and a bit full-time people. The customer whose transaction was held gets a call or a slight delay, and the money moves. It's a staffing line item.

On a login endpoint. Two million logins a day, 0.5% denied in error = 10,000 people who cannot get into their accounts. No queue, no analyst, no recovery path — the user's next action is a retry, then a support call.

Compare that against the availability target somebody signed. A 99.9% SLO permits 0.1% of requests to fail: at two million a day, a budget of 2,000 failed logins for everything — deploys, failovers, the lot.

A 0.5% false-positive rate on the login path spends five times your entire error budget, every day, and none of it appears on your availability dashboard, because a wrongly denied login is a 200 with authentication_failed in the body.

Your SRE dashboards count 5xx; your fraud model produces 200s. The outage is invisible by construction. I have watched a team run an eight-week experiment in which login success rate dropped 0.4% and nobody raised it, because the graph everyone watched was error rate, and error rate was flat.

Two amplifiers make it worse than the raw number.

Retries manufacture the evidence that justifies the block. A user denied at login doesn't stop; they try again, then again with a different password in case they mistyped, then from their phone. Call it 3.7 attempts. Ten thousand false positives become roughly 37,000 additional authentication attempts, disproportionately from accounts the model has already flagged, arriving in tight bursts from one device. Velocity rules — the deterministic kind identity legitimately owns — now fire on that burst, and the account gets locked. Reviewed afterwards, the evidence trail reads exactly like credential stuffing against that account, because in signal terms it is. The false positive has generated its own corroboration, and the only thing that distinguishes them is a piece of context nobody logged: whether the first denial was the model's.

The errors are unrecoverable in one direction only. A held transaction can be released. A denied login has no equivalent — there is no queue to appeal to, and the support agent taking the call sees risk_score=0.83 and nothing else. This is why the right action for a probabilistic signal at authentication time is almost never deny but challenge, which a legitimate user can pass; the base-rate arithmetic behind that is a separate argument I won't repeat. What's specific to this boundary is that fraud engines are built to emit decline, because in their native habitat decline is recoverable. Wire that verdict straight into login and you've imported a verb whose semantics didn't survive the journey.

Reproducibility, and why the audit log is lying to you

An identity platform's audit trail carries an unusual obligation: answer why, years later, in front of someone hostile. Deterministic controls satisfy it almost for free. "Denied: account locked after 10 failed attempts, rule LOCKOUT_V3, attempts at 14:02:11–14:02:39" is reconstructible forever, because the rule is a version-controlled artifact and the inputs are in the record.

A model score satisfies none of it. To reproduce 0.83 you need four things: the exact model artifact by content hash, retained for the audit period; the feature vector as computed at decision time, not recomputed later once the 30-day aggregates have moved; the threshold and policy in force at that instant; and the preprocessing code, usually a notebook-derived pipeline that also ships weekly.

Almost every deployment logs the score and none of the four. A score without them isn't evidence; it's a number with no referent — a pointer into a heap that was freed six retrains ago.

The retention mismatch is what makes this structurally hard rather than merely neglected. Identity audit records are commonly kept seven years. Feature vectors are wide — a few hundred floats per decision — so at two million logins a day, snapshotting them runs to roughly a terabyte a month, forever, for a system that was supposed to be a lookup; and you need every model artifact that ever served traffic alongside them. That cost is why nobody does it, and not doing it is why the March ticket closed as "cannot reproduce."

The fraud system doesn't have this problem: a disputed transaction settles inside a 120-day window, after which nobody asks again. The two systems have retention obligations that differ by a factor of twenty, and the merged system inherits the longer one and the more expensive record.

The related trap is change control. Identity changes go through review because a mistake locks out a customer base. Retraining is automated because a stale model is a bad model, and holding a retrain for two weeks during a live fraud fight is a real loss. Both cadences are right for their own system. Put an automatic weekly retrain inside the login path and you have a component that changes production authentication behaviour with no change ticket, no canary measuring the thing that matters, and nobody on the identity rotation aware it happened. You notice when login success steps down 0.3% on a Tuesday, and it takes three days to find the deploy — because the deploy wasn't in your repo.

The labelling asymmetry, which is the deepest reason

This is the argument I'd make if I only got one, and it's the one I most rarely see stated.

Fraud detection has ground truth. It arrives late — a fraud chargeback typically lands 30 to 90 days after the transaction, and card network dispute windows run to 120 — but it arrives, unambiguous and tied to a specific event. That delayed label is the engine of the whole discipline: it's what lets you retrain, measure drift, and compare a challenger against a champion with an honest number.

Authentication has no ground truth at all. You never find out that the successful login at 22:15 was the attacker. Nothing in the system ever emits "that one was wrong." The nearest thing to a label is an account-takeover report, which arrives weeks later, is attached to an account rather than a session, and covers a vanishing fraction of the actual incidents.

So when you put a model on the login path, it cannot be trained on authentication outcomes — there aren't any. It's necessarily trained on downstream fraud outcomes, and this is where it goes quietly wrong:

A login-blocking model trained on downstream fraud labels has a feedback path for its false negatives and none for its false positives. Only one side of its error surface can ever correct itself.

Trace both directions. A login the model wrongly allowed produces a transaction, then a chargeback, then a label, which trains the model to be stricter. That loop closes in 30 to 90 days, automatically.

A login the model wrongly denied produces no session, so no transaction, so no chargeback, so no label — ever. The user gives up, or calls support and becomes a row in a system the model has never read. From inside the training pipeline that decision isn't merely unlabelled, it's invisible: indistinguishable from a correctly blocked attack.

Now iterate. Every retrain pushes in one direction, because one side of the ledger is populated and the other is structurally empty. This is a selection-bias loop of exactly the shape a credit model has when it never observes the loans it declined — well understood in that world, with a name (reject inference) and an expensive countermeasure: deliberately approve a random holdout of applications you'd otherwise decline, purely to buy unbiased labels.

The equivalent at login is: allow a random sample of the logins your model wants to block, and see what happens. Say that out loud in a security review. It is methodologically correct and organizationally unsellable, which is why the drift is never corrected, and why login-path models get monotonically stricter until somebody notices the complaint volume and moves the threshold by hand — restarting the clock rather than fixing the loop.

Fraud engines don't have this problem, because their declines do generate feedback: the customer calls, an analyst reviews, the transaction is released or confirmed, and that outcome is a label. The feedback path exists in the fraud system's native environment and is destroyed by moving the decision to login.

The context fraud needs, and why identity mustn't hold it

Everything a fraud model is actually good at depends on facts an identity platform has no business storing.

What fraud needs What it is Why identity shouldn't hold it
Transaction amount, currency, velocity of value Payments state Changes per-second; identity is mostly read traffic sized for logins
Payment instrument history, BIN, issuer Card data adjacency Drags the IdP into PCI scope — the single most expensive attribute on this list
Chargeback and dispute history 120-day-lagging outcome data An outcome store inside a system with no outcomes
Counterparty and beneficiary graph Money-movement graph Fraud's core asset; meaningless without the payments domain
Merchant category, basket contents Commerce data Pure domain state — 12.4 in its clearest form
KYC/AML status, sanctions screening Regulated identity-of-record Different legal regime, different retention, different auditors
Account age, tenure Genuinely identity's Emit it. This one belongs to you.

The last row matters as much as the others: identity does hold a handful of facts nobody else can produce, and the correct move is to publish them.

Note the direction of travel, too. Fraud is largely a post-hoc discipline — its best features summarise weeks of behaviour and its labels describe what happened afterwards. Identity is pre-hoc: it decides now, with what it has. Asking identity to hold fraud's context asks a system optimised for a point-in-time lookup to become one optimised for retrospective aggregation over months. Different databases, different shapes, and attempting both in one store is how an identity platform stops being replaceable.

There's a regulatory sting on top. Once a model participates in an access decision, the model risk governance regimes banks apply to credit models (SR 11-7 and its equivalents) start pointing at your authentication service: independent validation, documented conceptual soundness, annual review, ongoing monitoring. And GDPR Article 22, plus the various adverse-action doctrines, have a good deal to say about automated decisions with legal or similarly significant effects on a person — "you may not access your account" being a strong candidate. None of it is fatal. It is an entire compliance apparatus attaching itself to your login path because of one HTTP call, and it never appears in the design doc.

The legitimate overlap, drawn honestly

The purist version of this argument is wrong, and it's worth being specific about where. Velocity checks, impossible travel, credential-stuffing detection and device recognition all genuinely belong near the login endpoint, and some belong inside identity. 12.1 drew the first cut — decidable from the request alone goes to the edge, requires account knowledge stays in identity — and it still holds. This boundary needs a second cut inside the account-knowledge half:

Identity may act synchronously on a signal it can compute from state it already reads during login, using a deterministic rule whose output is explainable from its inputs and reproducible from the audit record. Everything else it emits, and does not wait for.

Three clauses, all load-bearing. State it already reads bounds the latency, because the marginal cost is a field on a row you were fetching anyway. Deterministic rule preserves reproducibility. Explainable from its inputs means the support agent and the regulator get the same sentence.

Run the overlapping signals through it:

Signal Who computes it Who acts synchronously at login Why
Failed-attempt count for this account Identity Identity Its own state, deterministic, rate-limiting is a capacity problem
Password spraying across accounts Identity Identity (tenant-scoped throttle) Only identity sees across accounts; still a counter and a threshold
Known-device / registered-credential match Identity Identity The strongest allowlist signal it has, and it's a lookup
Breach-corpus password match Identity (or a specialist feed) Identity Deterministic set membership; k-anonymity makes it cheap
Country change vs account history Identity Identity, coarsely A country set on the account row; do not attempt city-level distance
IP reputation, ASN class, bot score Edge / specialist Identity consumes as input Not identity's to compute — 12.1
Device graph: this device seen on 400 other accounts Fraud Nobody, synchronously Requires a cross-account graph identity doesn't keep
Behavioural model score Fraud Nobody, synchronously Non-deterministic, unreproducible, unexplainable
Money-movement anomaly Fraud Nobody, at login The transaction doesn't exist yet
Chargeback-derived account risk Fraud Nobody, synchronously — consumed as a stored verdict Arrives days late; it's state, not a call

Two rows deserve a note.

Impossible travel is the one people expect to find in identity, and mostly shouldn't, at least not in the form it's usually built. Country-set membership is a clean deterministic check on data identity already has. Great-circle-distance-over-elapsed-time against IP geolocation is a physics calculation over a data source wrong at city level often enough to make the output noise, and it fires on VPNs, carrier egress changes, and anyone on a plane. Keep the coarse version; the sophisticated version is a fraud feature and belongs where features live.

Credential stuffing splits instructively. Detecting the campaign — a spike in unknown-device attempts across a tenant, an unusual ASN mix — is population-level work identity is uniquely positioned to do, being the only system that sees across accounts. Acting on it per-login synchronously is where it goes wrong: the right response is a tenant-level control (throttle, mandatory challenge, temporarily disable a weak factor) applied as a state change, not a per-request model consultation.

The shape that works

Identity emits; fraud consumes; fraud writes back verdicts, not scores; identity reads verdicts from its own store.

flowchart LR
    L["Login request"] --> ID["Identity runtime<br/>deterministic checks<br/>+ verdict lookup (local)"]
    ID -->|"allow · challenge · deny"| U["User"]
    ID -.->|"typed events: auth.succeeded,<br/>device.registered, factor.changed,<br/>recovery.completed"| BUS["Event stream"]
    BUS -.-> FR["Fraud platform<br/>models · features · labels · queues"]
    PAY["Payments / app events"] -.-> FR
    CB["Chargebacks, disputes,<br/>analyst outcomes"] -.-> FR
    FR -.->|"coarse verdict + TTL + reason<br/>narrow versioned API"| VS["Verdict store<br/>(identity-owned)"]
    VS --> ID

Solid arrows are the login path, and every one of them stays inside the identity platform or heads back to the user. Every dotted arrow is asynchronous and allowed to fail, lag, and retry — the same structural move as runtime/control-plane separation: the fast path reads projected state, it never calls out.

Six details decide whether it works.

The events must be well-typed and stable, because they are a public API. authentication.succeeded with a versioned schema — subject, tenant, timestamp, factor used, acr, device identifier, whether the device was newly registered, IP as passed from the edge, correlation ID — is worth more to a fraud team than any synchronous hook, because they can join it to everything else they hold and compute features you'd never have thought to expose. Emit authentication.failed too, with a reason code; the failure stream is where campaigns are visible. What you must not do is add fields on request, one per fraud experiment, until the schema is a feature vector: every extension point is a permanent public API, and an event schema is an extension point.

The verdict vocabulary must be small, coarse, and enumerated. Four or five values, no scores: require_step_up, require_high_assurance_factor, restrict_session (authenticated but read-only), account_restricted, clear. A score crossing this interface is a bug, because it forces identity to hold a threshold — and a threshold without the model that calibrated it is a magic number that will be wrong within a month. Coarse verdicts also survive a vendor swap.

Every verdict carries a TTL and a reason code, and the TTL is mandatory. This is the detail that gets skipped and causes the worst incident of the lot. With no expiry, a fraud system that dies mid-incident — or a bad batch job — leaves accounts restricted permanently and with no process to clear them, because the system that set the flag is the only one that knows how to unset it, and it's gone. Bound it: 24 hours for a step-up requirement, seven days for a restriction, plus a documented manual override with an audit record and a named actor. Same principle as the stale-window discipline in 12.2 — a degraded upstream should decay your posture toward normal, deliberately, at a rate somebody signed off on.

The verdict is read locally, not fetched. Identity reads it from its own store, on a row it was already fetching. That converts an availability dependency into a staleness dependency, which is the trade you want: a fraud outage makes verdicts stale rather than stopping logins. State the staleness budget as a number.

No code crosses the boundary. The tempting shortcut is a hook: let the fraud team ship a scoring function that runs inside the auth path, and all this event plumbing becomes unnecessary. Same offer as every plugin proposal, same costs — unbounded latency, a permanent interface, an availability dependency on someone else's release cadence, and no correct answer when it fails. ClavionX declines customer code in the identity runtime for exactly these reasons; I work on it, so take that as disclosure rather than recommendation. The fraud case is where the offer is most persuasive, because the party asking is internal, competent, and right about the problem. The answer is still events out, verdicts in.

And the real fraud decision belongs at the action, not the login. This is what resolves the incident everyone is trying to prevent. Authentication answers is this the account holder? Fraud answers is this action going to cost us money? The second isn't answerable at login, because the action doesn't exist yet — at authentication time you know the least you will ever know about intent. Move the high-assurance requirement to the moment the payee is added or the amount entered, where the fraud engine has the transaction in hand, a decline is recoverable, and the challenge is proportionate and comprehensible. That's step-up authentication doing the job it exists for, and it beats blocking at the door: the legitimate user is never interrupted for logging in, and the attacker meets the wall at the exact moment they attempt the thing they came for.

The tells

You've merged them if any of these is true:

  • The login handler makes an outbound call to a service your team doesn't deploy.
  • Nobody can state what happens to authentication when that service returns 503 — or two people state different things.
  • Your audit log contains a risk score with no model version beside it.
  • Authentication behaviour changed with no corresponding ticket in the identity repo.
  • A model retrain requires a login-path canary — or, worse, doesn't have one.
  • Login success rate is on the fraud team's dashboard and not the identity team's.
  • A support runbook contains the step "ask the fraud team to allowlist the user."
  • Identity's data model has a field whose only consumer is a model.
  • An identity incident and a fraud incident are always the same incident.

And the failure modes, by misplacement:

Misplacement How it fails When you find out
Synchronous model call in login p95 login latency doubles; timeouts spike under load Your next traffic peak — i.e. during an attack
Fail-closed on the fraud service Fraud outage becomes total authentication outage The fraud team's first bad deploy
Fail-open on the fraud service Control silently absent whenever it matters Never, which is the problem
Model verdict wired to deny Error budget spent 5× over, in invisible 200s A support queue, weeks later
Score logged without model version No decision can be reconstructed The first regulator or legal hold
Training on chargeback labels only Threshold drifts monotonically stricter Complaint volume, months in
Fraud context stored in identity PCI scope creep; the IdP becomes unswappable The migration that gets cancelled
Verdicts with no TTL Accounts permanently restricted by a dead system The morning after a batch job fails

What to do on Monday

  1. Find every outbound call in your login path and write down its failure behaviour — not the intended one, the actual one, established by running the test. If it's a fraud or risk service, that line is the whole conversation.
  2. Graph login success rate beside error rate on the identity team's dashboard, with an alert on a 0.2% step change. If a model is denying logins today, this is how you see it, and it costs an afternoon.
  3. Check whether any risk score in your audit log has a model version beside it. If not, you cannot answer "why was this denied" for any decision the model made, and it's better to know before someone asks.
  4. Enumerate your verdict vocabulary and put a TTL on every entry. Restrictions set by a system that can't unset them are a lockout incident scheduled for an unknown date.
  5. Ask the fraud team what they'd want from an event stream and build that instead of the hook they were about to request. In my experience they prefer it: a well-typed event they can join to their own data beats a 50ms window inside someone else's request path, and they know it.

Neither system is being demoted here. Fraud detection is harder than authentication and better instrumented, and the fraud team is usually right that they can see things the identity platform can't. The mistake is concluding that the model should therefore move to the login path. It's the signal that should move — outward, asynchronously, as events — and what comes back should be a small, durable, expiring fact identity can read from its own store, explain to a support agent, and reproduce in front of a regulator three years later.

Authentication's product is a decision you can defend. Fraud detection's product is a probability you can improve. Both are valuable. Neither survives being asked to be the other.


Disclosure: I work on ClavionX. It appears once above, for its refusal to execute customer code inside the identity runtime, because that constraint is exactly what forces the events-out/verdicts-in shape described here — and it's a constraint with real costs, not a feature I'm selling. The argument applies identically whatever you run.