Loslegen
Javaneer
Zurück zur Stufe
Stufe 7·Praktisches System Design

Queues & Lastglättung

Message-Queues als Stoßdämpfer eines Systems: Produzenten von Konsumenten entkoppeln, stoßweise Last glätten, Back-Pressure und At-least-once-Zustellung mit idempotenten Konsumenten.

14 Min. LesezeitExperte

Deutsche Übersetzung in Arbeit

Diese Lektion ist noch nicht ins Deutsche übersetzt und wird daher auf Englisch angezeigt. Der Rest der Seite ist vollständig lokalisiert.

Auf dieser Seite

The last building block of scalable design is the message queue - the shock absorber that sits between parts of a system so they don't have to move in lockstep. You met queues in the Spring messaging module; here we see why they're a load-bearing tool in system design. A queue decouples a producer from a consumer, absorbs spiky load so a burst doesn't topple a downstream service, and lets slow work happen asynchronously off the request path. It's how a system survives a traffic spike gracefully instead of falling over.

Decoupling: producers and consumers move independently

Without a queue, a producer calls a consumer directly and synchronously - so the producer's speed is chained to the consumer's, and if the consumer is down, the producer fails too. A queue breaks that chain: the producer drops a message on the queue and moves on; the consumer picks it up when it's ready.

Synchronous (coupled):   Producer ──calls──► Consumer   (producer waits; consumer down = producer fails)

Queued (decoupled):      Producer ──►[ Queue ]──► Consumer
                         (producer returns immediately; consumer processes at its own pace)

This decoupling (the same benefit as domain events, now at infrastructure scale) means the two sides can be scaled, deployed, and fail independently. Add more consumers to process faster; the producer neither knows nor cares.

Load-leveling: the queue as a buffer

The killer feature for scale is load-leveling (a.k.a. the queue-based load-leveling pattern). Real traffic is spiky - a flash sale, a viral moment, a nightly batch - and a synchronous system must be provisioned for the peak or it collapses under it. A queue turns spikes into a steady stream:

Bursty arrivals:   ▁▁█████▁▁▁███████▁▁   (producers, wildly variable)
                          │
                     [  Queue  ]  ← absorbs the burst; depth grows then drains
                          │
Steady processing: ▃▃▃▃▃▃▃▃▃▃▃▃▃▃▃▃▃▃   (consumers, provisioned for AVERAGE, not peak)

The consumer processes at its own steady rate; when arrivals exceed that rate, the queue simply gets deeper and drains later. You provision consumers for the average load, not the terrifying peak - a huge cost saving and a guarantee that a spike causes delay, not failure. The queue depth is your buffer, and it's also a signal: a persistently growing queue means consumers can't keep up and you need to scale them.

Back-pressure: when the buffer isn't infinite

Queues aren't magic - if producers permanently outpace consumers, the queue grows without bound and eventually you're out of memory or storage. Back-pressure is the system's way of signaling "slow down": bounded queues that reject or block producers when full, so overload is handled deliberately (shed load, return a 429) rather than by silently exhausting resources and crashing. A queue smooths temporary spikes; back-pressure protects against a sustained overload the system genuinely can't handle. (Recall reactive streams - back-pressure was their whole point.)

Delivery guarantees and idempotency

Queues introduce the delivery-semantics question from the messaging and batch modules:

  • At-most-once - a message might be lost, never duplicated. Rarely acceptable.
  • At-least-once - a message is never lost but may be delivered more than once (a consumer crashes after processing but before acknowledging, so it's redelivered). The common, practical default.
  • Exactly-once - the ideal, but very hard and expensive in a distributed system; usually simulated by at-least-once + idempotent consumers.

Because at-least-once means duplicates, consumers must be idempotent - processing the same message twice has the same effect as once (the idempotency-key idea from the API module, applied to messages). This is the single most important rule of queue-based systems: design consumers to tolerate duplicates. Related patterns: a dead-letter queue for messages that repeatedly fail (so one poison message doesn't block the queue).

Async isn't free - you trade immediacy and simplicity

A queue makes work asynchronous, which means the result isn't ready when the producer returns - the caller must handle 'accepted, will be done later' (a 202, a status to poll, a later notification). You also add infrastructure to run and monitor, eventual consistency to reason about, and duplicate-handling to build. Use a queue when decoupling and load-leveling are worth that cost - spiky load, slow work, cross-service communication - not for a simple, fast, synchronous request that a direct call handles fine. Async is a powerful tool, not a default.

A ticket-number system at a busy deli

Imagine a deli with no system: at the lunch rush, a crowd mobs the single counter all at once, orders get lost, and the overwhelmed staff grind to a halt (a synchronous system collapsing under a spike). Add a take-a-number dispenser and a queue: customers grab a ticket and step back (the producer drops a message and moves on), and the staff serve numbers steadily at whatever pace they can sustain (consumers processing at a steady rate). The lunch rush no longer topples the counter - it just makes the wait longer (the queue gets deeper and drains), so the deli is staffed for the average, not the mob. If the line grows out the door for hours (sustained overload), they stop handing out tickets (back-pressure) rather than promise service they can't deliver. And because a distracted server might call the same number twice, each order is marked done when filled (idempotency) so no one gets two sandwiches.

Introduce a queue where it helps

An e-commerce checkout currently does everything synchronously in one request: charge the card, then generate a PDF invoice, send a confirmation email, update the analytics warehouse, and notify the shipping partner. During a flash sale, the endpoint times out and fails because the downstream services (email, shipping API) are slow and overwhelmed. Redesign with a queue, say what stays synchronous, and name one consumer property you must ensure.

What are the two primary benefits a message queue provides in a scalable system?

Key takeaways

  • A message queue decouples producers from consumers so they scale, deploy, and fail independently - infrastructure-scale version of event-driven decoupling.
  • Load-leveling is the key benefit: the queue buffers spiky traffic so consumers run at a steady rate, provisioned for average not peak - a spike causes delay, not failure.
  • Back-pressure (bounded queues that reject/block when full) protects against sustained overload a queue alone can't absorb.
  • Queues typically give at-least-once delivery (possible duplicates), so consumers MUST be idempotent; use a dead-letter queue for repeatedly-failing messages.
  • Async trades immediacy and simplicity for decoupling and resilience - use queues for spiky load, slow work, and cross-service comms, not for simple fast synchronous calls.
War diese Lektion hilfreich?
Diese Seite auf GitHub bearbeiten