A Queue Is a Tool for Buying Time
Summary
A queue does not make processing faster. It separates the rate at which requests are received from the rate at which they are processed, absorbing momentary bursts over time.
Why this was needed
Suppose you have an image upload API. It saves the original, creates three thumbnails, extracts metadata, and adds the image to the search index. If you do all of it synchronously, the response takes 4 seconds. Normally that is tolerable. But when uploads increase tenfold on a promotion day, 4-second requests pile up tenfold, all the worker threads get tied up, and even the health checks fail.
With a queue, the API saves only the original and replies "received". That is 200ms. Workers handle the rest at their own pace. Even when a burst arrives, the API keeps replying in 200ms and only the queue depth grows. The increased depth becomes processing delay — what took 4 seconds may take 3 minutes. In exchange, the system does not die.
This exchange is all there is to a queue. You accept delay and buy availability.
The stability condition is one inequality
There are three conditions under which a queue makes sense. First, is it acceptable for the work not to finish immediately? Second, can the request rate momentarily exceed the processing rate? Third, is it acceptable to try again later if it fails? If all three are yes, a queue is right.
Conversely, there are also clear cases where you should not add a queue. Lookups where the user must see the result immediately, and cases where the average inflow is greater than the average processing capacity. Many people misunderstand the latter. A queue absorbs bursts but does not solve chronic overload.
λ = 초당 유입 건수 μ = 워커 하나의 초당 처리 건수 × 워커 수
λ < μ → 큐 깊이가 0 근처로 돌아온다 (안정)
λ = μ → 깊이가 무작위로 떠돈다 (경계 — 운영하면 안 되는 지점)
λ > μ → 깊이가 선형으로 자란다 (터진다)
The numbers make it clear. Suppose μ = 100/s and λ is normally 60/s, but rises to 90/s for 30 minutes. The margin that disappears per second in this period is 10 items. Over 30 minutes nothing is left in the queue and it is absorbed well. But if λ rises to 110/s, 10 items pile up every second, reaching 18,000 items in 30 minutes. If this state does not end, it eats all the memory or disk and blows up.
If λ > μ persists, what you need is not a queue but more workers or an inflow limit.
Delay is not proportional to utilization
Here is something people often miss. Wait time does not grow in proportion to utilization (ρ = λ/μ); it explodes nonlinearly. As a rough feel, it goes like this.
| Utilization ρ | Wait time multiplier (1/(1−ρ)) | Meaning |
|---|---|---|
| 0.5 | 2x | Comfortable |
| 0.8 | 5x | It starts to show |
| 0.9 | 10x | A little more and it gets sharply worse |
| 0.95 | 20x | Hard to operate |
| 0.99 | 100x | Effectively an outage |
So you must not size worker capacity "exactly to the average inflow". Aim for a utilization of 70–80%, and let autoscaling or an inflow limit take what is above that.
What you meet in the field
Once you add a queue, new operational metrics appear. Queue depth and consumer lag. If you do not put these two on a dashboard, the queue becomes a silent black hole. Users report "I uploaded it but it doesn't show up", and the API dashboard is all green. The problem is piled up behind the API.
Looking at depth alone is not enough. Whether a depth of 1,000 is serious depends on the processing rate. You can only judge by also looking at depth ÷ processing rate = drain time. If you process 100 per second, it is a 10-second backlog, and at 1 per second it is a 17-minute backlog. Set alerts on this value too.
One more alert. Depth staying flat at 0 is also an abnormal signal. If all consumers die and nobody takes items out, only inflow should pile up, but if the producers have died too, the depth is flat at 0. If the count of completed items is 0 and the depth is 0, the whole pipeline has stopped.
Also, adding a queue requires designing the user experience along with it. How to show the "received" state, how long the user has to wait, how to tell them when it fails. If you add only the queue without deciding these, users just see a request that vanished. You must decide at least three things.
- Return a job ID immediately — the user must be able to ask about the status later.
- Build a path to query the status — look at pending, running, and failed through
GET /jobs/{id}. - Notify about failures — send jobs that have used up their retries to the dead-letter queue and have a person look at them.
Four things to answer before choosing a queue
Choosing the tool comes last. Before that, you must decide four properties first, and if you choose Redis or Kafka without deciding them, you will end up rebuilding everything later.
Delivery guarantee. At-least-once is the default. Exactly-once is not something the queue gives you; it is something you get by making the consumer idempotent. You record the IDs of jobs already processed and skip them if the same one arrives again.
Ordering. A queue that preserves global order cannot process in parallel. In practice, per-key ordering is usually enough — only the jobs of the same user need to be in order, and the order relative to other users does not matter. If you set the partition key to the user ID, you get this.
Retry and giving up. Decide how many times to retry, how to lengthen the interval, and when to give up. If you do not mix jitter into exponential backoff, failed jobs retry clustered at the same moment.
Dead-letter queue (DLQ). This is where jobs that have used up their retries go. Without a DLQ, such a job either circulates forever or silently disappears. Always set an alert on the count piled up in the DLQ — what piles up here is what a person must look at.
What to check in the next quiz
This module covers only concepts. From the next module on, you will build queues yourself with Redis lists and streams, and reproduce one by one where ordering, duplicates, and loss arise.