TT Lab
Get started
Learn Learning paths Courses

LLM Serving

You Cannot Control LLM Cost by Request Count

Continue in TT Lab

In one line

For an ordinary API, the cost of one request is roughly the same, but for an LLM, one request can be 100 tokens or 100,000 tokens. So a gateway must count not requests but tokens.

Why this was needed

Say you set a rate limit of 60 requests per minute. If a tenant puts a 200,000-token context in every request, the number of requests is within the limit, but the cost and GPU occupancy are hundreds of times those of other tenants. And those requests monopolize the KV cache, so other requests get preempted.

A limit on the number of requests does not prevent this situation at all. What you need is accounting in units of tokens.

How it works

There are five things a gateway must do.

First, authentication and tenant identification. Map API keys to tenants.

Second, token accounting. Count and record the request's input tokens and the response's output tokens separately. With streaming, the output tokens are finalized only when the stream ends, so even when it is cut off midway, you must account for the tokens up to that point.

Third, limits at two levels. Apply RPM (requests per minute) and TPM (tokens per minute) together. RPM prevents abuse and TPM controls cost and resources. A token bucket suits TPM well — it allows a momentary burst from a tenant that is usually quiet while keeping the long-term average.

Fourth, 429 and Retry-After. When rejecting, tell when it is all right to come back. Without this, clients each retry on their own and create waves.

Fifth, cost calculation. Each model has different input and output unit prices, so you accumulate by multiplying the token count by the unit price. Output tokens are usually more expensive than input tokens, so they must be counted separately.

A budget is attached to this. Set a daily or monthly cap per tenant and, when it is exceeded, block or downgrade to a cheaper model. The latter is called a fallback, and it is often better than the service stopping entirely.

What to limit on

If you apply only RPM (requests per minute), it is almost useless in an LLM service. That is because one request can be 100 tokens or 100,000 tokens. So you apply two axes together.

Axis What it prevents Unit
RPM Request floods, bots Requests per minute
TPM GPU time consumption Tokens per minute (input + output)
Concurrency Queue explosion, memory Number of in-progress requests

Of the three, concurrency reflects the GPU most accurately. vLLM processes in batches, so the number of concurrent requests is the KV cache usage, and if that overflows, it is an OOM or preemption. RPM and TPM are for cost accounting, and concurrency is for stability.

Implementing with a token bucket lets you allow bursts while keeping the average.

용량 = 60,000 토큰,  채우는 속도 = 1,000 토큰/초 (= 60k TPM)

요청이 오면 max_tokens 만큼 버킷에서 뺀다(예약)
완료되면 실제 사용량과의 차이를 돌려준다(정산)
버킷이 비면 429 + Retry-After: <채워질 때까지의 초>

Always send Retry-After. Without it, clients retry immediately and the situation gets worse.

Issuing a good 429 is also design

What you tell the client when rejecting decides the quality of the client code.

HTTP/1.1 429 Too Many Requests
Retry-After: 12
X-RateLimit-Limit-Tokens: 60000
X-RateLimit-Remaining-Tokens: 0
X-RateLimit-Reset-Tokens: 12s

And write which limit was hit in the body. What the client should do differs depending on whether it was RPM, TPM or the budget — for the first two it can just wait, but for exceeding the budget waiting does not help.

Caching is the biggest saving

Bigger in effect than limits is caching. There are three layers.

What it looks like in the field

The fact that you cannot know the token count in advance is what makes this design hard. Input tokens can be counted at request time, but output tokens are known only when generation ends. So in practice you reserve max_tokens as an upper bound and, after completion, settle with the actual usage. If you do not reserve, requests that arrive at the same time all judge themselves to be within budget, pass through, and then exceed it together.

Cost explosion incidents usually occur in retry loops. If a client retries on every timeout while the server is still generating, a single user causes the same answer to be produced five times. If the gateway accepts an idempotency key and merges identical in-progress requests, this waste disappears.

What you will do in the next lab

You set up a gateway in front of the mock server you built earlier. You add API key authentication, an RPM limit, token accounting, a TPM token bucket, Retry-After, per-model cost calculation, and finally blocking when a tenant budget is exceeded.