TT Lab
Get started
Learn Learning paths Courses

Queues and Asynchronous APIs

λ, μ and Queue Depth

Continue in TT Lab

Summary

The stability condition is the single inequality λ < μ. The moment the inflow rate exceeds the processing rate, the queue depth grows explosively rather than linearly.

Why this was needed

If you look at a queue depth graph, there is something odd. The load rises gradually, yet the queue depth stays stuck near 0 for a while, and then past a certain point it suddenly shoots up vertically. This is not a bug but a property of queueing.

When the inflow rate λ is lower than the processing rate μ, the queue is mostly empty. As the two get closer, the average depth grows because it is absorbing momentary fluctuations, and when λ exceeds μ, the queue grows without limit. The difference in wait time between a utilization ρ = λ/μ of 0.7 and of 0.95 is not double but far greater.

So you must not size worker capacity at exactly 100% of the inflow. The practical convention is to aim for a utilization of 70–80% and leave headroom.

How it works

The tool that translates queue depth into delay is Little's Law. L = λ x W. The average number of items in the system is the inflow rate times the average time spent in the system. Inverted, it is W = L / λ. If the queue depth is 3,000 and the processing rate is 50 per second, a message that just arrived will be processed 60 seconds later. Once you can do this calculation, "the queue has piled up a bit" turns into "a user submitting now waits a minute".

Backpressure is how the system protects itself in this situation. There are three layers.

First, a queue cap. You set a maximum length on the queue and return 429 to the producer when it is exceeded. It is more honest to say you cannot accept it now than to accept without limit and fail to do it later.

Second, load shedding. You drop the lowest-priority requests first. Rejecting some quickly and keeping the rest alive is better for overall availability than holding on to every request, processing them slowly, and dying together.

Third, adjusting consumer concurrency. You raise μ. However, you cannot raise it without limit — when you add workers, the DB or external API behind them becomes the new bottleneck. You must check each time whether you moved the bottleneck or removed it.

What you meet in the field

When you return a 429, it helps to also give a Retry-After header. A perceptive client respects this signal and refrains from unnecessary early retries. If you do not give it, clients retry with their own backoff, and those overlap.

And queue-depth alerts are better set on trends than on absolute values. Some systems are normal at a depth of 1,000 and others abnormal at 10. "Increasing continuously for 5 minutes" is a far more reliable signal.

Where to apply backpressure

When a queue starts to grow there are only three options, and if you choose none, a fourth happens — it uses up all its memory and dies.

Method What you give up Where it fits
Block the producer (blocking) Response time Internal pipelines, batches
Reject new requests (429) Some requests Public APIs
Drop the oldest first Old data Metrics, logs, real-time quotes

The third is surprisingly often right. Processing a quote that updates every second 5 minutes later has no value. If it is data for which old means worthless, dropping is the right answer. Conversely, payment events must never be dropped, so use the first or second row.

A bounded queue is the default

An unbounded queue only delays the problem and does not remove it. If you set a cap, the problem shows up in the producer, not the queue — and that is seen much sooner and can be acted on.

# ❌ 무제한 — 메모리가 다 찰 때까지 아무 신호가 없다
q = asyncio.Queue()

# ✅ 상한을 두면 put 이 기다리고, 그 대기가 곧 신호다
q = asyncio.Queue(maxsize=1000)
await asyncio.wait_for(q.put(item), timeout=0.5)   # 못 넣으면 거절한다

You decide the cap by drain time. If the processing rate is 100 per second and you allow up to 10 seconds of delay, 1,000 is the cap. Decided this way, the queue depth becomes the delay budget.

When a concurrency limit is more accurate

In many cases it is better to limit the number of in-flight jobs rather than the queue depth. This is especially so when the downstream is fragile.

sem = asyncio.Semaphore(20)      # 하류에 동시에 20개까지만

async def handle(item):
    async with sem:
        await downstream.call(item)

However large you make the queue, if the downstream can withstand only 20 at a time, sending more than that is bringing the downstream down. The queue is a buffer and the semaphore is a shield. You use the two together.

What to put on the dashboard

큐 깊이                       ← 지금 얼마나 밀렸나
소진 시간 = 깊이 ÷ 처리율      ← 이것이 경보 기준
생산율 λ 와 소비율 μ           ← 둘의 차이가 추세를 만든다
거절·폐기 건수                ← 역압이 실제로 작동했는가
가장 오래된 항목의 나이        ← 지연의 최대치

The last line is especially useful. Even at the same depth, it tells you whether old items keep staying backed up (whether first-in-first-out has broken).

What you will do in the next lab

You fill the queue with a load generator while recording the depth as a time series, estimate the wait time with Little's Law, attach a queue cap and 429, raise the worker concurrency and measure the improvement in processing rate, and then write up the stability condition as a document.