Kafka came out of a genuine problem at LinkedIn: moving very large volumes of event data into systems that processed it in different ways, without each producer having to know about each consumer. The log abstraction that came out of that work is one of the more useful ideas in distributed systems.

What travels badly is the conclusion that an event log should be the default way services talk to each other.

The question that decides it

There is one test, and it is not about scale or fashion.

Can the caller continue without the answer?

If it can, an event is usually right. The producer emits, the consumer gets to it, and the asynchrony costs nothing because nobody was waiting.

If it cannot, the wait exists either way. Expressing it as an event does not remove the wait. It removes your ability to see it: instead of a timeout on one call with a stack trace, you have a request that produced no response, a consumer that may or may not have received it, and a correlation id to chase across four services and two retention windows.

What people are actually buying

Teams reach for events expecting decoupling, resilience and scale. What arrives alongside is a set of obligations that are easy to miss at design time.

Expected Also arrives
Producers do not know consumers Nobody knows who consumes what, six months in
A slow consumer cannot block a producer A slow consumer accumulates lag nobody is watching
Retry is free At-least-once delivery, so every consumer must be idempotent
Scale by adding consumers Ordering is only guaranteed within a partition
Replay is possible Retention has to be sized, and it is a storage bill
Services are independent A schema change now coordinates across every consumer

None of these is an argument against events. They are the work that comes with them, and a design that has not accounted for them produces a system that is harder to reason about than the one it replaced.

The shape that usually wins

The pattern most estates converge on, and the one Zalando described arriving at, is not a choice between the two. It is events to propagate change and an API to answer a question.

Events behind, request in front
An event log behind a synchronous read path, with a serving store between themA caller that needs an answer now makes one synchronous hop to a serving store shaped for that read, which is visible in a trace and has one dependency. Behind it, producers emit to a partitioned log and a projector keeps the serving store current. Nobody waits on that path. The trade is staleness, which becomes one number on one graph with an alert on it.BEHIND THE READ PATHnobody is waiting heresynchronouskeeps currentA caller that needs an answer nowServing store, shaped for this readone hop, one dependency, visible in a traceProducersPartitioned logat-least-once, ordered per partitionProjectoridempotent, with a dead letter path
Read path
synchronous, debuggable, one dependency
Write path
asynchronous, replayable
The risk
staleness, which is now a number you can alert on
The test
can the caller continue without the answer?
This diagram as text
  • A caller that needs an answer now
  • Serving store, shaped for this read — one hop, one dependency, visible in a trace
  • Behind the read path nobody is waiting here
    • Producers
    • Partitioned log — at-least-once, ordered per partition
    • Projector — idempotent, with a dead letter path

Relationships

  • A caller that needs an answer now → Serving store, shaped for this read — synchronous
  • Producers → Partitioned log
  • Partitioned log → Projector
  • Projector → Serving store, shaped for this read — keeps current
Open full size · after LinkedIn on the log and Zalando on moving a read path back to an API

The trade is explicit: the serving store is eventually consistent, so the design has to state how stale it may be and alert when it exceeds that. That is a far better position than an unbounded asynchronous chain, because staleness is one number on one graph rather than a property you discover by investigation.

If you do run a log, run it properly

Idempotent consumers, without exception. At-least-once means a consumer will see the same event twice, usually during a rebalance or a deploy. Design for it rather than hoping.

A dead letter path from day one. A message that cannot be processed will block its partition and everything behind it. The first time this happens should not be the first time anybody thinks about it.

Consumer lag as a first-class alert. Not broker health, which is usually fine while consumers fall further behind. Lag per consumer group, graphed, with a threshold.

Schemas with a registry and a compatibility rule. The moment two teams consume the same topic, the payload is an interface. Treating it as one from the start is cheaper than discovering it when a field changes.

Retention sized deliberately. It is a storage decision and a recovery decision: it determines how far back you can replay, which determines what a consumer bug costs you.

Partition count thought about once, carefully. It bounds consumer parallelism and it is awkward to change later.

Where it is genuinely the right answer

To be clear, there are cases where an event log is the correct and obvious choice:

  • Several unrelated consumers need the same stream, and adding a new one should not require changing the producer.
  • Replay has real value: rebuilding a projection, backfilling a new consumer, or recovering from a processing bug without asking producers to resend.
  • Volume genuinely exceeds what a synchronous path can absorb, measured rather than assumed.
  • Producer and consumer have legitimately different availability requirements, and buffering between them is the point.

If one of those is true, the operational weight is worth carrying. If none is, you are taking on the obligations without the benefit.

How STP approaches this

We start with whether the caller can continue without the answer, because that settles most of these conversations in a sentence. Where a log is justified we design the dead letter path, the idempotency and the lag alert as part of the build rather than as hardening afterwards, since those three are what separate a system that degrades visibly from one that fails silently. Where an estate is already event-driven and hard to debug, the usual finding is a synchronous question expressed asynchronously.

More on Kubernetes platform engineering and DevOps and platform automation, or start a conversation.