The Composability Test: Five Questions Before Adding a Feature

Arguments about scope are hard to win because both sides are reasonable.

Arguments about scope are hard to win because both sides are reasonable.

Someone proposes adding bot detection to the identity platform. The case is genuine: customers are getting hit by credential stuffing, the auth endpoints are where it lands, and we're the ones who can see it. The counter-case is equally genuine: specialist vendors do this better, it drags us toward being a security platform, and we'll maintain it forever.

Both people are right about what they're describing. What's missing is a shared way to decide, so the argument gets resolved by whoever is more senior or more persistent — which is a bad way to make architectural decisions, and a worse way to make them repeatedly.

What follows is the test I'd apply. Five questions, in order. It won't make the decision for you, but it converts an argument about instinct into an argument about specifics, and it makes the cost of a "yes" visible before you pay it.

Question 1: Does this require knowledge only we have?

The first and most useful filter.

Some decisions are impossible to make anywhere else. Whether an account has seen forty failed logins from forty different IPs in ten minutes requires knowing what an account is and having a view across its history. No upstream layer can compute that — every individual request looks unremarkable. That's ours by necessity.

Other decisions require nothing from us at all. Whether an IP appears in a threat feed, whether a TLS fingerprint suggests automation, whether a request is part of a volumetric flood — all decidable from the request itself, by a layer that sees vastly more traffic than we do.

If the decision can be made without our data, we're not the right place for it. We might be a convenient place, which is how these things get proposed, but convenience isn't a claim to ownership.

Question 2: Is someone else better positioned?

If the answer to Q1 was "no, this doesn't need our knowledge," this question asks whether we'd even be good at it.

Better positioned usually means one of: more data (a CDN sees traffic patterns across millions of sites), more specialization (a fraud vendor does nothing else), or better placement (a gateway sits in front of every API, not just auth).

The uncomfortable version of this question is: if we build this, will it be as good as what the customer could buy? Often the honest answer is no — we'd build a serviceable version that's worse than the specialist's, and then maintain it forever, and then have to defend it in evaluations against products that do only that.

There's a real counterargument worth taking seriously: an adequate built-in feature that's already integrated can beat a superior external one that requires configuration and a separate contract. Bundling has genuine value. But that's an argument about product convenience, and it should be made as one — not disguised as an architectural argument.

Question 3: Does this add a dependency to the hot path?

The one with the sharpest consequences.

If the feature means the authentication path now calls something — a third-party API, another internal service, a customer's system — then every property of that dependency becomes a property of authentication. Its latency is added to your latency budget. Its availability multiplies against yours. Its maintenance window becomes your outage.

This is where most damage happens, because the request usually arrives phrased as a small addition. Just check this one API before issuing the token. The check is small. The coupling is permanent.

Three responses, in order of preference:

  • Move it off the path entirely — publish an event, let the consumer handle it asynchronously.
  • Pre-compute it — if the data comes from an external system, synchronize it ahead of time rather than fetching at login. Same data available, completely different failure characteristics.
  • If it must be synchronous, budget it explicitly — a hard timeout, a defined behavior on failure, and a number in the latency budget that someone signed off on.

And answer the fail-open/fail-closed question before building it, not during the first incident. If this dependency is unavailable, does authentication proceed or stop? Both answers have consequences and one of them has to be chosen deliberately.

Question 4: Does this survive a deployment we don't control?

If the platform can be self-hosted, air-gapped, or deployed into a customer's own environment, this question has teeth.

A feature that assumes an external service — a threat feed, a cloud API, a hosted ML model — either doesn't work in those deployments or quietly makes them non-functional in a way nobody notices until a customer's security review. You end up with two products: the one that works and the one you ship to regulated customers.

The resolution is usually reasonable defaults that a better external layer can supersede. Basic account-scoped throttling that works everywhere, and clean integration points for the customer who already runs a specialist product in front. Not a WAF; not nothing either.

If the honest answer is "this only works in our cloud," that's not automatically a no — but it should be a deliberate decision about which deployments you're serving, made once, not discovered per-feature.

Question 5: Can we ever remove it?

The question nobody asks, and the one that determines the ten-year cost.

Some features can be deprecated. A UI can change, a default can flip, an internal implementation can be rewritten. Others are permanent the moment they ship:

  • Anything a customer writes code against becomes a frozen public API — including behaviors you never documented.
  • Anything in a compliance narrative can't be removed without a customer's auditor being involved.
  • Anything a customer's own automation depends on will break loudly when touched.

Before shipping, write down how it would be removed. Not because you plan to, but because the exercise reveals which features are decisions and which are commitments. A feature you can't describe removing is one you're agreeing to carry indefinitely — and that's fine, as long as everyone knows that's the deal being made.

Running the test

Take the bot detection proposal through it.

Does it require knowledge only we have? No — bot signals are computed from the request. Points against.

Is someone better positioned? Yes, substantially. A specialist sees orders of magnitude more traffic. Points against.

Does it add a hot-path dependency? If it calls an external scoring service on every login, yes. Points against.

Does it survive an air-gapped deployment? Only if the model runs locally, which makes it much weaker. Points against.

Can we remove it? Once customers rely on it for their security posture, no. Points against.

Five for five, which makes this an unusually clear case — most aren't. The useful output isn't the verdict; it's that the conversation is now about five specific things rather than about whether the feature is "in scope," which is a word that means whatever the speaker needs it to mean.

And the test frequently produces a better yes. Bot detection fails on all five. But consuming a bot score computed upstream passes all five: it uses our account knowledge to contextualize an external signal, adds no dependency we own, degrades gracefully when the signal is absent, and can be removed. Same customer need, different architecture.

That's usually where the value is — not in refusing, but in finding the version of yes that doesn't cost you the properties you were selling.

Where the test misleads

Three honest limits.

It's biased toward saying no. Five questions each capable of producing an objection will reject things that should ship. Some features are worth building despite failing several — because a customer segment genuinely needs them bundled, or because the integration quality of building it yourself is materially better. The test tells you what you're spending, not whether the purchase is wrong.

It undervalues bundling. Customers pay for coherence. "It works out of the box" has real worth, and this framework doesn't price it. If you apply the test mechanically you'll build something architecturally admirable that loses evaluations to a product people can actually turn on.

Q1 is fuzzier than it looks. "Requires knowledge only we have" is clear at the extremes and genuinely arguable in the middle. Device reputation, for instance, benefits from both a global view and account history. The test surfaces the ambiguity rather than resolving it — which is still progress, since the alternative was not noticing.

Why it's worth having

Feature-by-feature, every addition is defensible. That's exactly the problem — it's how platforms accrete their way into being something nobody chose to build, one reasonable yes at a time.

A written test doesn't stop that. What it does is make the cost of each yes explicit at the moment it's paid, so that if you end up owning bot detection, threat intelligence, and a fraud engine, it's because you decided to — not because five separate conversations each ended with "sure, that's small."

The point isn't to say no more often. It's to know what you're buying.