Impossible Travel Is Full of False Positives
The queue had forty-one items in it when she picked it up on Monday morning, and by the third coffee she had a rhythm.
Alert 7: user in Bangalore, then Singapore, eleven minutes apart. Resolution: the company's egress for SaaS traffic is in Singapore, and the user's laptop had reconnected to the VPN after a wifi drop. Alert 12: London to Dublin in four minutes. Resolution: the mobile app refreshed its token while the handset moved from wifi to cellular, and the carrier's gateway is where it is. Alert 19: Frankfurt to Ashburn, Virginia in ninety seconds. Resolution: the user clicked a magic link and their employer's mail security product opened it first, from a datacentre, to check whether it was a phishing page. Alert 23: the same user as alert 7, twice more.
By item thirty-four she had stopped writing resolutions and started using a saved comment: known egress pattern, no action. The last seven took ninety seconds in total. None of them were attacks. None of them had been attacks the previous Monday either, or the Monday before that, and everyone on the rota knew it, which is why the queue was the one thing that got skipped when a real incident landed on the same day.
Somewhere in the eleven months this detector had been on, it had almost certainly fired on something real. Nobody could say which item that was, because nobody had ever gone back through the closed ones, and the closure notes all said the same thing.
That is the normal condition for this control, and it is not because the team was sloppy. It is because of what the control actually computes, which is not what its name suggests.
Impossible travel is not a measurement. It is a chain of three inferences — IP to place, place to person, two places and a clock to a claim about a human body moving through the world — and each link carries its own error, of a kind that compounds rather than cancels. The output is a weak hypothesis wearing the costume of an observation, and virtually every deployment failure follows from treating it as the observation.
This piece is about that chain, why the errors in it are structural rather than fixable, what the base rate does to whatever survives, and — the part that matters, because a purely negative article is useless — the two places where the same signal is genuinely worth having, both of which are somewhere other than the login gate.
I'm not going to re-derive risk-based authentication in general. Adaptive MFA: Security Should Follow Risk, Not Rules covers the full signal set, how to combine signals without labels, and the arithmetic of low-precision detectors; this is a deep read on one signal from that list, the one people reach for first and understand least. Identity Isn't Your Fraud Engine draws the boundary about which of these computations belong inside an identity platform at all, and lands on roughly the position this piece argues for from the measurement side. Consider both prerequisites I'm building on rather than repeating.
What the detector actually computes
Written out honestly, the algorithm is: take two authentication events for the same account, resolve each source address to a coordinate pair, compute the great-circle distance between them, divide by the elapsed time, and compare the implied velocity against a threshold that is usually a commercial airliner's cruise speed with some slack.
Every one of those steps is a modelling decision, and the composition looks like this:
flowchart TB
IP["Source IP as recorded<br/>at the auth endpoint"]
L["A coordinate pair"]
P["Where the human was"]
V["'This person could not<br/>have travelled that far'"]
IP -->|"inference 1: database lookup"| L
L -->|"inference 2: no intermediary"| P
P -->|"inference 3: same person, one body,<br/>accurate clocks"| V
A1["assumes: the record is<br/>current, granular, and about<br/>this address today"] -.-> L
A2["assumes: no VPN, proxy, CGNAT,<br/>carrier gateway, cloud client,<br/>scanner, corporate egress"] -.-> P
A3["assumes: a session is a person,<br/>both events are logins,<br/>the interval is meaningful"] -.-> V
The dotted boxes are the whole article. Take them in order.
Inference one: an IP address is not a location, it's a business record
The first thing to internalise is that IP geolocation databases are not maps. They are compilations of administrative and commercial facts: registry allocations, WHOIS registrant addresses, BGP announcement patterns, reverse-DNS naming conventions, latency triangulation from measurement networks, data purchased from partners, and — increasingly, and this is the good part — self-published feeds in which network operators state where they intend a prefix to be used. RFC 8805 standardised the format for exactly that self-publication, because the guessing had become bad enough that large content providers were asking ISPs to please just tell them.
Notice what none of those inputs are: an observation of a device's physical position. The database answers "where is this network administratively and topologically anchored?" and you are asking it "where is the person?" Those questions have the same answer often enough to be useful for choosing a CDN edge or a default currency, and they diverge in ways that are neither random nor rare.
Three properties follow that matter for a physics calculation and are usually discarded before it runs.
Accuracy is structurally uneven, not uniformly noisy. It varies by country, by operator, by allocation size, and above all by connection type. Fixed residential broadband in a country with many regional ISPs and stable, small allocations resolves well. A large mobile operator that anchors a whole country's traffic at two or three gateways resolves to those gateways. Business ranges resolve to a registrant address that may be a headquarters in another city. Hosting ranges resolve to wherever the provider says, which is often right about the datacentre and says nothing about who is driving the machine. Anyone quoting a single accuracy figure for "IP geolocation" is quoting an average over populations that behave completely differently, and your users are not distributed like that average. The honest characterisation is qualitative: country-level resolution is usually right and sometimes badly wrong; city-level resolution is a different quality of data in different parts of the world and for different connection types, and you should assume it is weakest exactly where your least typical users are.
The databases return an uncertainty, and almost every integration throws it away. Commercial providers ship a confidence or accuracy-radius field alongside the coordinates — this point, plus or minus some kilometres. That radius is frequently large enough to swallow a metro area, and for some records large enough to swallow a small country. A velocity calculation that consumes the point and drops the radius has converted an estimate into a measurement by omission. If you keep only one thing from this section, keep that one: the uncertainty was in the response and your code deleted it.
Fallbacks manufacture coordinates that look precise. When a database can place an address in a country but no more finely, many pipelines fill in a national or regional centroid rather than returning null. The result is a perfectly well-formed latitude and longitude for a point no user is at and some users are hundreds of kilometres from — and there is a well-known, rather sad genre of story about ordinary residential addresses that sit at such a default and receive years of misdirected attention from people tracing IPs. The part that matters here is that the centroid is not marked as a fallback by the time it reaches your risk engine. A user moving between two ISPs whose fallbacks differ has "travelled" without moving, at a velocity determined by cartography.
Then there is the error nobody schedules for: the answer changes over time. These databases update continuously — weekly is typical — and prefixes get reallocated, re-announced and re-anchored. An alert raised in March and investigated in May may not reproduce, because the address now resolves elsewhere. If your audit record stores the IP but not the resolved location and the database version that produced it, you have logged a fact that cannot be replayed — an instance of the general problem in the hidden cost of audit logs: the field you kept is not the field the decision was made on. Store both, and store which build said so. It is also the difference between an alert an analyst can close with confidence and one they close with a saved comment.
And underneath all of it, one question that has nothing to do with geography: is the IP you are geolocating even the client's? It arrived through a load balancer, a CDN and a reverse proxy, most likely as a header, and headers are attacker-influenced unless every hop in the chain strips and rewrites them correctly. What happens when the layer below you lies works through this class in general; the version relevant here is that a risk control keyed on X-Forwarded-For with sloppy trust configuration is a control an attacker can steer. Not merely evade — steer, into flagging the legitimate user.
Inference two: a location is not the user's location
Grant, for a moment, a perfect geolocation database. The second link still breaks, because the address you observe belongs to whatever last touched the packet, and modern networks are full of things that touch the packet.
Corporate egress and split tunnelling. An always-on VPN sends everything through a concentrator that may be on another continent. A split tunnel is worse for our purposes, because it sends some traffic through it, selected by destination: a user reaching an internal IdP over the tunnel and a SaaS application directly generates two observed locations for two halves of one login. SASE and ZTNA generalise this — the point of presence a user egresses from is chosen by the vendor's routing, per-destination, and can change mid-session for reasons internal to the vendor. The user did not move. The routing policy did.
Enterprise SSO across two networks. Related and underappreciated: in a federated login, the identity provider and the service provider each see the browser from wherever their own path egresses. If both run impossible-travel detection — and both often do — they can compute different velocities for the same authentication event, and neither is wrong about what it saw. When a customer reports an alert you don't have and you can't reproduce, this is a leading candidate.
Carrier-grade NAT and gateway placement. Mobile networks do not hand each handset a routable address; traffic is anchored at a gateway and NATed behind a comparatively small pool of public addresses. Where that gateway sits is a capacity and peering decision, not a geographic one, and it can be hundreds of kilometres from the handset. A user commuting between two cities may egress from one gateway in the morning and another in the afternoon, generating movement that is real but not theirs. Fixed-line CGNAT does the same at smaller scale, and is common in networks that ran out of IPv4 addresses — which, structurally, means it is more prevalent in some regions and operator types than others, so its effect on your false-positive mix depends on where your users are.
Home-routed roaming, which runs backwards from intuition. When a subscriber roams internationally, many operators still route the data session back to the home network before it reaches the internet. A traveller in Madrid on a UK SIM therefore appears to originate in the UK. Then they connect to the hotel wifi and appear in Madrid. Then a background sync fires over cellular again and they are back in the UK. This produces textbook impossible travel, at high frequency, for a user who is genuinely travelling — precisely the population the control claims to protect, and precisely the moment they can least afford an authentication problem. The direction is the interesting part: the alert fires because the user's traffic failed to move with them.
Dual-stack disagreement. A device with both IPv6 and IPv4 may reach one service natively over v6 with a prefix delegated by the local ISP, and another over v4 through a CGNAT pool anchored elsewhere. Same device, same second, two families, two locations. Happy Eyeballs means which one gets used is a race condition. If your event pipeline stores whichever address the connection happened to use, your velocity calculation is now sampling from two different populations of address.
Cloud and hosting addresses generated by legitimate software. A user's own automation, a scripted export, a personal VPS, a mobile app whose backend makes some calls server-side, an integration refreshing a token from a worker — all authenticate as the user, from a datacentre, on a schedule. A growing category, because most identities in a modern estate are not people and many of the machine ones still use human credentials.
Things that follow links without a human. Mail security products detonate URLs in a sandbox before delivery, corporate proxies prefetch, chat clients unfurl previews, scanners crawl. If any of your authentication or verification flows is reachable by GET — a magic link, an email confirmation, a device-approval URL — some fraction of your "user activity" is a scanner in a datacentre, arriving before the user, from a location the user has never been. Teams tend to discover this only when someone asks why one customer's alerts always come in pairs.
The unifying observation: none of these are edge cases in the sense of being rare. They are the ordinary architecture of enterprise networking and mobile telephony. A detector whose false-positive generators are "corporate VPN", "mobile carrier", "cloud", and "email security" is a detector whose false-positive rate is a function of how modern your users' environments are.
Inference three: two locations and a clock do not describe a body
The third link is the one people don't examine at all, because it feels like arithmetic rather than assumption. It isn't. Converting two placed events into a claim about human movement requires three further things to be true.
That both events belong to the same person. Shared accounts exist, service accounts exist, and accounts with delegated access exist. An assistant with the credentials, a contractor sharing a login the org has not got round to splitting, a team mailbox, a break-glass account with three legitimate holders — all generate genuine simultaneous access from genuinely different places. This is a real security problem, but it is a different one, and it will not be fixed by challenging whoever happened to authenticate second.
That both events are logins. This is where most self-inflicted noise comes from. If your detector consumes an event stream that includes token refreshes, session resumptions, background syncs and silent re-authentications, you are computing velocities over events that involve no human at all. A phone in a pocket, waking on a schedule, moving between wifi and cellular, produces a stream of authenticated requests from alternating egress points while the human is asleep. The account "moves" all night. Your own infrastructure is a false-positive generator, and it is the one you control. Decide explicitly which event types are eligible and write it down, because the default — "everything in the auth log" — is a choice that nobody made on purpose. This is the same distinction as authentication state versus session state: a refresh is a statement about a session's continuity, not about a person arriving.
That the interval is meaningful. Here is the piece of arithmetic that should change how you think about the threshold. Implied velocity is distance over elapsed time, and both terms are uncertain — the distance by tens or hundreds of kilometres from inference one, the time by clock skew and by whatever queueing sits between the event and the timestamp. As elapsed time shrinks, the location error is divided by an ever-smaller number and the implied velocity grows without bound. Two events forty seconds apart with a modest thirty-kilometre jitter in the resolved coordinates imply something in the region of 2,700 km/h — supersonic, from a user who did not stand up.
The consequence is not a detail. The detector is most likely to fire on the shortest intervals, and the shortest intervals are exactly where the automated, human-free events live. Token refreshes, retries after a network flap, an app resuming as the train leaves the tunnel, a page load that authenticates twice. Your worst-precision band and your highest-volume band are the same band, and a threshold expressed in kilometres per hour cannot separate them because the unit itself is the problem. A minimum-elapsed-time floor below which you simply do not compute a velocity is the single highest-value change most deployments can make, and it costs nothing.
Three inferences, each individually reasonable, composed into a verdict. Compose enough reasonable inferences and you get a confident-looking number that no single step in the chain would have supported on its own.
Then the base rate arrives
Everything above concerns the numerator. The arithmetic that finishes the argument concerns the denominator, and it is where engineers' intuitions reliably fail.
The full derivation lives in the adaptive MFA piece and I won't repeat the worked example. The shape of it is what matters here: when the prior probability of an event is very small, a detector with an excellent true-positive rate and a low-sounding false-positive rate still produces an alert population that is overwhelmingly false, because the false positives are drawn from an enormously larger pool. Account takeover in any given login attempt is a rare event. A detector that flags a small percentage of a very large number of legitimate logins will bury the handful of real ones no matter how good its recall is.
Two things about impossible travel make its position on that curve worse than the generic case, and both are structural.
Its false-positive sources are correlated with your user population, not random. The noise is not white. It clusters on the enterprise customers with global VPN egress, the users in countries with heavy CGNAT, the mobile-first users, the people who travel. So the alert queue is not a thin uniform drizzle across the estate; it is a torrent from a subset of accounts, which means the same names recur, which means the analyst learns to recognise the names, which means the queue is triaged by pattern-matching on the account rather than by examining the event. That is a rational adaptation to the data, and it is also precisely the failure mode by which the one real alert gets closed with a saved comment.
Its true positives are bounded from above, because evasion is trivial and cheap. For the detector to fire on an attack, the attacker must be far from the victim and the victim must be active in a nearby window. An attacker who cares about not firing it buys egress in the victim's city. Residential proxy networks make that a small, elastic operating cost — the economics of which the credential stuffing piece covers from the attacker's side. So the signal is loudest against the attacker who wasn't trying to hide, and silent against the one who was. You are running a detector with an attacker-selectable off switch and a defender-mandatory alert queue.
Now the part that makes this a security argument rather than an operations complaint. An alert queue that everyone has learned to ignore is not a neutral cost; it is a detection capability you believe you have and do not. The postmortem sentence is always some version of the alert did fire, on the ninth of the month, and it was closed as a known egress pattern. Nothing about that is a human failing. A control that produces mostly noise trains the humans attached to it, exactly as an over-firing MFA prompt trains users to approve reflexively — the mechanism MFA fatigue is built on, applied to your responders instead of your customers. The organisational chart says you have detection. The behaviour says you have a queue.
The mirror image, if you auto-challenge
Many teams, aware the queue is unmanageable, wire the detector directly to a step-up challenge. This does not remove the cost; it moves it onto users, and it does so with unusually bad timing.
Consider who trips this control. Disproportionately, people who are actually travelling. And a traveller is at their least capable of completing a challenge at precisely the moment they are asked: roaming or on a foreign SIM so SMS is unreliable or unroutable, in a hotel on a network that mangles things, in a different timezone from their support desk, possibly without the laptop that holds the other factor, possibly having just landed with a nearly dead phone. The challenge is a fair request that the user cannot satisfy.
What happens next is the actual security event. They call the help desk, which — faced with a stranded executive on a bad line who is late for a meeting — does what help desks do, and the organisation quietly discovers what its real authentication policy is. Do this often enough and you have not added a control; you have built a training programme for a bypass path, pushing traffic from your strongest authentication surface onto your weakest. Recovery and assisted flows are already the hardest and most attacked part of any identity system. Routing your travelling users into them, at volume, on a signal with this precision, is a poor trade before you even count the support cost.
There is a fairness dimension too, and it is the same one authentication is becoming invisible raises about behavioural signals in general: a control whose false positives concentrate on people with unusual networks, unusual geography or unusual travel patterns degrades specifically for them, and never shows up in an aggregate challenge-rate dashboard that looks perfectly healthy.
The constructive half: what the signal is actually good for
None of this makes geo-velocity worthless. It makes it the wrong shape for a gate. Used differently, it earns its keep. Six changes, roughly in order of value.
1. Make it a continuous feature, not a boolean. "Impossible: yes/no" discards everything. What you want in the risk vector is a small set of numbers and categories: implied velocity, elapsed time, the geolocation confidence radii of both endpoints, whether either endpoint is a country this account has ever been seen in, and how the pair compares to this account's own history. A feature can be weighted, calibrated and overruled. A boolean can only be believed.
2. Prefer network identity over distance. The highest-value substitution here. ASN, network type and reputation are more informative and far better grounded than coordinates, because they describe what the address is rather than guessing where it is. "This account has authenticated four hundred times from two residential ASNs and one corporate range, and this attempt is from a hosting provider" is a defensible statement about an administrative fact. "The user moved 1,200 km in nine minutes" is a physics claim resting on three inferences. Same input data, wildly different epistemic quality. The signal is also asymmetric in the useful direction: a hosting or anonymiser ASN should add risk, while a residential one should subtract very little, because residential proxies are cheap. If you consume one thing derived from the IP, consume this and not the distance.
3. Compare against the account's own history, never a population norm. A per-account country set, a per-account ASN set, a per-account distribution of hour-of-day: cheap to maintain, easy to explain to a user and an auditor, and dramatically better behaved than any absolute threshold. A user whose logins have come from three countries for two years is not anomalous when they appear in a fourth in the way a user with one country for two years is. Absolute rules cannot express that; a per-account baseline expresses it for free, and degrades gracefully — a new account has no baseline, and "no baseline" is an honest reason to treat something as unknown rather than as normal.
4. Require corroboration from a device signal before acting. Geo-velocity on its own should never move a decision. Geo-velocity plus an unrecognised device is a different proposition, and the reason is that the device signal is the one that isn't a chain of inferences. A cryptographically bound device credential — a registered passkey, a key in the secure element, a client certificate — answers a question with a yes or a no rather than an estimate.
This has a strong corollary that deserves saying plainly: for accounts with phishing-resistant, device-bound credentials, location largely stops mattering. A WebAuthn assertion is origin-bound and non-replayable; where the packets appear to come from adds very little to a signature produced by hardware you enrolled. The same holds for sender-constrained tokens. Every unit of effort spent on device binding reduces how much you need location to do, and location is the signal you cannot make better. Effort spent tuning velocity thresholds does not compound; effort spent on binding does.
5. Choose asymmetric responses. The binary "challenge or allow" is not the only available action, and for a signal this weak it is the worst one. Better options, in increasing order of intrusiveness: enrich the event and move on; annotate the session so any sensitive action later requires step-up; restrict the session to read-only; notify the user out-of-band; challenge. The middle two are the sweet spot for a weak signal, because a legitimate traveller reading their email never notices, while the same session attempting to change a payout account or an MFA enrolment meets a wall at the moment the intent becomes clear. That is step-up authentication doing what it exists for, and it turns a low-precision signal into a proportionate control instead of an interruption.
6. Know what you will do with an alert before you enable it. This sounds procedural and is actually the crux. A detector with no defined action produces a queue by default, and a queue with no owner is a queue that decays. Write the sentence first: when this fires, X happens, and if nobody is going to do X, we are not turning it on. Who owns the incident when every layer worked is about what happens when that question goes unanswered across teams; the single-detector version is smaller and just as consequential.
Where impossible travel genuinely pays
Two places, and the distinction between them and the login gate is the most useful thing in this article.
Fleet-level correlation. One account showing an impossible-travel pattern is noise. Fifty accounts in one tenant showing the same pattern within an hour — same origin ASN, same destination range, same shape — is a real signal, and nothing else in your stack can see it, because only the identity layer observes across accounts. The aggregation works because the errors described above are largely independent per account and correlated per network: one user's VPN noise does not replicate across fifty unrelated users, so a pattern that does replicate is telling you about infrastructure rather than geography. Same structural argument the adaptive MFA piece makes about population signals catching credential stuffing that per-login scoring is blind to — as an aggregate feature, at tenant scope, feeding an investigation rather than a gate.
Post-hoc investigation. During an incident, "show me every account with anomalous geo-velocity in this window, ranked" is a genuinely good query. It works there and not at login because the base rate has changed: you are no longer sampling from all logins but from a set already narrowed by a known compromise, so the prior is enormously higher and precision follows. Same detector, same threshold, radically different usefulness — worth sitting with, because it means "is this signal any good?" has no answer independent of what you have already conditioned on. Retrospective work is also where a human analyst supplies the context the detector lacks, and where being wrong costs a query rather than an account lockout. This is the pre-hoc/post-hoc split from Identity Isn't Your Fraud Engine seen from the other side: the signal belongs to the retrospective discipline and gets weak when forced into a synchronous decision.
There is also a third case, which is a close cousin worth separating out because it is much stronger than either.
Session-artifact impossible travel — that is, the same credential in two places at once. If a single session cookie, refresh token or device-bound session is presented from two materially different networks in an overlapping window, that is not a claim about a human's travel. It is a claim that one artifact exists in two places, which is a much cleaner proposition, and it is the classic signature of token theft: the attacker replays the stolen session while the legitimate user's session continues undisturbed. The concurrency is the signal, not the distance — and it is far more robust to the errors in inferences one and two, because you no longer need to know where anything is, only that the two egress points are structurally different (different ASN, different country, different network class) while both are active.
Post-authentication session theft is where attackers have been migrating as credentials have got harder to phish. If you are going to spend engineering effort in this general area, spend it here rather than on the login gate: the precision is better, the action is clear — invalidate the session and force reauthentication — and the cost of being wrong is one extra login rather than a stranded user in a foreign hotel. It also composes properly with sender-constrained tokens, since binding the session stops the replay working at all.
How to decide whether to keep it on
The honest answer is that you cannot know from first principles whether this control is worth its cost in your environment, because the cost depends on your users' networks and the benefit depends on your attackers. So measure it. Here is a protocol that takes a few weeks and settles the question.
Sample and actually resolve. Take a hundred alerts at random — at random, not from the top of the queue — and drive each to a real root cause. Not "known egress pattern": which VPN, which carrier, which scanner, which refresh. It is tedious and it is the only step that produces a number you can defend. You will end up with a false-positive taxonomy specific to your estate, almost always dominated by three or four causes, several of which you can eliminate outright once named.
Count what it caught first. Go through every confirmed account takeover in the last year and ask, for each one, whether this detector was the first thing to flag it, a corroborating signal, or absent. That three-way split is the number that decides the control's fate, and it is different from any detection rate a vendor can quote you, because it is measured against your confirmed incidents rather than against a labelled set. If the answer is that it was never first and rarely corroborating, you have learned something worth acting on.
Measure the challenge cost on the users it hits. If you are auto-challenging, look at completion and abandonment for challenges triggered by this signal specifically, and at help-desk contacts within an hour of one. Then look at where those users were. The pattern to expect: challenges concentrated on travelling and roaming users, completion materially worse than baseline, a visible bump in assisted recovery. That bump is your bypass path being trained.
Instrument the elapsed-time distribution. Plot alert volume against the interval between the two events. If it spikes hard at the short end, most of your alerts are infrastructure rather than geography, and a minimum-interval floor plus an event-type filter will remove a large fraction of the queue in an afternoon. Do this before anything else; it is the cheapest possible improvement.
Then make one of three decisions, explicitly.
Retire it as a trigger, keep it as an enrichment field. This is the right answer more often than not, and it is not a defeat. The velocity, the two ASNs, the two country codes and the confidence radii stay on the authentication event, where an analyst and a downstream fraud platform can use them. Nothing fires; nothing is challenged; the data remains available for the two cases where it genuinely pays.
Keep it, narrowed. Country-set membership rather than distance, a minimum interval, logins only, corroboration from an unknown-device signal required before any action, and an asymmetric response — restrict the session, don't block it. This is a defensible control and it will fire perhaps a tenth as often. Note that what survives the narrowing is essentially "this account appeared in a country it has never been in, on a device we don't know" — which is a statement about the account's own history and its device, with the geography demoted to a coarse categorical. That is what the signal was worth all along.
Promote it to where it works. Move the computation into fleet-level anomaly detection and into your investigation tooling, and separately build the concurrent-session-artifact check, which is a different and better detector wearing a similar name.
What you should not do is the thing most organisations do, which is to leave it running as a per-login trigger because turning off a security control requires someone to sign their name to the decision and nobody wants to be that person if a breach follows. That reasoning is understandable and it is how estates accumulate controls that nobody believes and everybody maintains. If you can show that the detector has never been first to a confirmed incident, that its alerts are closed unread, and that its challenges are pushing users onto the help-desk path, you have made the security case for retiring it. The signature you need is on that evidence, not on the switch.
The broader lesson generalises past this one signal, and it is why the measurement exercise is worth the fortnight. Every risk signal is a chain of inferences from something you observed to something you want to know, and its usefulness is set by the length of that chain and by the base rate of what you're hunting — not by how intuitive the story sounds when you explain it to an executive. Impossible travel tells a wonderful story: a stolen account, a distant attacker, physics catching them out. It is memorable, it demos well, and it appears on every slide about behavioural risk. The chain behind it is three links long and the base rate is brutal, and both facts were knowable before the detector was ever switched on.
Ask the same two questions of the next signal someone proposes. How many inferences deep is it, and what happens to its precision when it meets your actual traffic? Most controls will survive the question. The ones that don't were never doing the work you thought they were.