Why Every Identity Team Eventually Builds an Event Bus

Nobody sets out to run Kafka because they wanted to run Kafka.

Nobody sets out to run Kafka because they wanted to run Kafka.

It starts with a single line of code in the login handler. Someone needs an audit record — reasonable, obviously required — so you write the insert right there, inline, where the login succeeds. It's five lines. It works. You move on.

That line is the seed of a distributed streaming platform, and the path from one to the other is so predictable I can describe it without knowing anything about your system.

The four-step slide

Step one: audit. The insert goes in the login handler. Fine.

Step two: notifications. Security wants an email when someone logs in from a new device. That's a call to the notification service, in the login handler, next to the audit insert. Still fine, though the handler is starting to do things that aren't logging anyone in.

Step three: provisioning and analytics. SCIM needs to know about user changes. The product team wants login events in the data warehouse. Two more calls. The login handler now has four downstream dependencies, and someone in code review says "should this really all be inline?" and everyone agrees it shouldn't and nobody has time to fix it.

Step four: the incident. The analytics endpoint gets slow. Not down — just slow, 3 seconds instead of 30 milliseconds. And because it's called synchronously from the login path, every login on the platform now takes 3 seconds.

Your analytics pipeline just caused an authentication outage.

That's the moment the event bus stops being an architectural preference and becomes a requirement. Because you've discovered the actual problem: login availability had become the product of every downstream consumer's availability. Four consumers at 99.9% each means your login path is 99.6%, and every new consumer makes it worse.

flowchart LR
    subgraph Before["Synchronous"]
        L1["Login"] --> A1["Audit"]
        L1 --> N1["Notify"]
        L1 --> S1["SCIM"]
        L1 --> An1["Analytics"]
    end
    subgraph After["Event-driven"]
        L2["Login"] --> Bus["Event Bus"]
        Bus --> A2["Audit"]
        Bus --> N2["Notify"]
        Bus --> S2["SCIM"]
        Bus --> An2["Analytics"]
    end

The difference between those two diagrams isn't aesthetic. In the first, the login path fails if any consumer fails. In the second, a consumer failing means that consumer falls behind. That's the entire argument.

Why identity in particular

Plenty of systems have this fan-out problem. Identity has it worse, for three reasons.

Identity events are inherently interesting to everyone. A user logging in, a permission changing, an account being disabled — these are facts that security, compliance, product, support, and half a dozen internal services all legitimately need. In most domains the number of parties interested in any given event is small. In identity it's structurally large, because identity is the thing everything else is built on top of.

Identity is the most latency-sensitive path in the system. Login latency is directly user-visible and sets the first impression of your entire product. It's the worst possible place to accumulate synchronous side effects — and yet it's where all the interesting events originate.

The consumers have wildly different reliability characteristics. Audit must never lose a record. Analytics can drop a percent and nobody cares. Notifications should be timely but not at the cost of login. SCIM propagation to a customer's system depends on their availability. Coupling all of these to one synchronous path means every consumer inherits the strictest requirement and imposes its own weakest availability.

What actually gets hard

Adopting a bus solves the coupling problem and hands you a new set of problems that are genuinely harder than the one you started with. Worth knowing them going in.

At-least-once means duplicates are normal. Almost every practical message system delivers at-least-once. Your consumers will see the same event twice — after a rebalance, a retry, a redeploy. For audit, a duplicate is a confusing double entry in an investigation. For provisioning, a duplicate might mean a second "create user" call.

Consumers must be idempotent. That's not optional and it's not something you retrofit easily. Give every event a stable ID and make processing a no-op the second time.

Ordering is only guaranteed per partition. "User created" and "user deleted" arriving out of order produces a deleted user who then exists. The standard fix is partitioning by a key — user ID, tenant ID — so that all events for one entity land in the same partition and stay ordered relative to each other.

Choose that key deliberately. Partitioning by user ID gives per-user ordering but can hot-spot on a busy tenant. Partitioning by tenant gives per-tenant ordering and can hot-spot on your largest customer. There's no key that avoids both.

The dual-write problem is the one that actually loses data. Here's the trap:

db.save(user)          // succeeds
bus.publish(event)     // fails

The user exists. No event was published. Every downstream consumer is now permanently wrong, and nothing will ever correct it. Reverse the order and you get the opposite failure — an event for a user who doesn't exist.

There's no ordering of two independent systems that makes this safe. The standard answer is the transactional outbox: write the event into a table in the same database transaction as the state change, then have a separate process read that table and publish. Now the state change and the intent to publish are atomically consistent, and publishing becomes a retryable operation rather than a coin flip.

flowchart LR
    W["Write user + outbox row\n(one transaction)"] --> DB[("Database")]
    DB --> Relay["Outbox relay"]
    Relay --> Bus["Event Bus"]

Almost every team that runs an event-driven identity system arrives at the outbox pattern eventually. Usually after a data-consistency incident they can't explain.

Consumer lag becomes an operational metric that matters. With synchronous calls, "did it happen" is knowable from the stack trace. With events, a consumer can be silently behind for hours. Audit lag means your compliance record is incomplete. SCIM lag means a deprovisioned employee still has access at a customer's site — which is, notably, a security incident with a compliance dimension, not just a delayed job.

Lag per consumer needs a dashboard and an alert. It's the single most important operational signal in an event-driven identity system, and it's invisible unless you build it.

Schema evolution is forever. Once events are consumed by systems you don't own, the event schema is a public API. Adding a field is fine. Removing one, renaming one, or changing a type breaks consumers that may deploy on a different schedule than you. Additive-only changes, versioned schemas, and a compatibility policy — decided early, not after the first break.

What to do about it

Don't build it on day one. One consumer is a function call and should stay one. Premature Kafka is a real failure mode — you get the operational burden of a distributed log to serve a single audit table. The signal isn't "we might need this someday," it's the arrival of the second or third consumer of the same event.

Do design the seam on day one. Even with a direct call, put an publishIdentityEvent() abstraction between the login handler and the consumer. It costs nothing now and it means the migration to a real bus later is a change in one place instead of a change in every handler.

Not everything needs the bus. Some things genuinely must be synchronous — if a policy check gates whether login proceeds, that's not a side effect, it's part of the flow. The test is simple: if this fails, should the login fail? If yes, keep it inline. If no, it belongs on the bus. Most things that end up inline never passed that test; they just got written there first.

Consider whether you need a full streaming platform. Kafka is the right answer for high-volume, multi-consumer, replay-required workloads. It's also a serious operational commitment. Plenty of identity deployments do fine with a simpler queue, or even a polling-based outbox relay with no broker at all. The architectural property you need is decoupling, and the outbox gives you most of that. The broker is an implementation detail you can upgrade into.

The honest summary

The event bus isn't a design choice most identity teams make. It's a conclusion they're driven to, by the specific combination of a latency-critical hot path and an unusually large number of legitimate consumers of what happens on it.

Knowing that in advance doesn't mean building it in advance. It means recognizing the moment it arrives — the second consumer, the third, the first time a downstream slowdown shows up in your login latency graph — and not treating that moment as a surprise.

The teams that struggle aren't the ones that built the bus late. They're the ones who kept adding synchronous calls to the login handler because each individual one was small, and never noticed they were assembling a distributed system with none of the properties a distributed system needs.