You Cannot Control LLM Cost by Request Count
In one line
For an ordinary API, the cost of one request is roughly the same, but for an LLM, one request can be 100 tokens or 100,000 tokens. So a gateway must count not requests but tokens.
Why this was needed
Say you set a rate limit of 60 requests per minute. If a tenant puts a 200,000-token context in every request, the number of requests is within the limit, but the cost and GPU occupancy are hundreds of times those of other tenants. And those requests monopolize the KV cache, so other requests get preempted.
A limit on the number of requests does not prevent this situation at all. What you need is accounting in units of tokens.
How it works
There are five things a gateway must do.
First, authentication and tenant identification. Map API keys to tenants.
Second, token accounting. Count and record the request's input tokens and the response's output tokens separately. With streaming, the output tokens are finalized only when the stream ends, so even when it is cut off midway, you must account for the tokens up to that point.
Third, limits at two levels. Apply RPM (requests per minute) and TPM (tokens per minute) together. RPM prevents abuse and TPM controls cost and resources. A token bucket suits TPM well — it allows a momentary burst from a tenant that is usually quiet while keeping the long-term average.
Fourth, 429 and Retry-After. When rejecting, tell when it is all right to come back. Without this, clients each retry on their own and create waves.
Fifth, cost calculation. Each model has different input and output unit prices, so you accumulate by multiplying the token count by the unit price. Output tokens are usually more expensive than input tokens, so they must be counted separately.
A budget is attached to this. Set a daily or monthly cap per tenant and, when it is exceeded, block or downgrade to a cheaper model. The latter is called a fallback, and it is often better than the service stopping entirely.
What to limit on
If you apply only RPM (requests per minute), it is almost useless in an LLM service. That is because one request can be 100 tokens or 100,000 tokens. So you apply two axes together.
| Axis | What it prevents | Unit |
|---|---|---|
| RPM | Request floods, bots | Requests per minute |
| TPM | GPU time consumption | Tokens per minute (input + output) |
| Concurrency | Queue explosion, memory | Number of in-progress requests |
Of the three, concurrency reflects the GPU most accurately. vLLM processes in batches, so the number of concurrent requests is the KV cache usage, and if that overflows, it is an OOM or preemption. RPM and TPM are for cost accounting, and concurrency is for stability.
Implementing with a token bucket lets you allow bursts while keeping the average.
용량 = 60,000 토큰, 채우는 속도 = 1,000 토큰/초 (= 60k TPM)
요청이 오면 max_tokens 만큼 버킷에서 뺀다(예약)
완료되면 실제 사용량과의 차이를 돌려준다(정산)
버킷이 비면 429 + Retry-After: <채워질 때까지의 초>
Always send Retry-After. Without it, clients retry immediately and the situation gets
worse.
Issuing a good 429 is also design
What you tell the client when rejecting decides the quality of the client code.
HTTP/1.1 429 Too Many Requests
Retry-After: 12
X-RateLimit-Limit-Tokens: 60000
X-RateLimit-Remaining-Tokens: 0
X-RateLimit-Reset-Tokens: 12s
And write which limit was hit in the body. What the client should do differs depending on whether it was RPM, TPM or the budget — for the first two it can just wait, but for exceeding the budget waiting does not help.
Caching is the biggest saving
Bigger in effect than limits is caching. There are three layers.
- Exact-match cache — if the prompt and parameters are the same, it returns the stored answer. In FAQ-type traffic the hit rate sometimes exceeds 30%. If the temperature is not 0, the product must decide whether giving the same answer to the same input is right.
- Semantic cache — it finds similar questions with embeddings and reuses them. The hit rate is high, but there is a risk of wrong reuse, so the threshold is set conservatively.
- Prefix cache — it reuses the KV cache of requests whose front part is the same, like a system prompt.
This is vLLM's
enable_prefix_caching, and it greatly reduces TTFT with no effect on accuracy. The first thing to turn on.
What it looks like in the field
The fact that you cannot know the token count in advance is what makes this design hard. Input tokens can be counted at request time, but output tokens are known only when generation ends. So in practice you reserve max_tokens as an upper bound and, after completion, settle with the actual usage. If you do not reserve, requests that arrive at the same time all judge themselves to be within budget, pass through, and then exceed it together.
Cost explosion incidents usually occur in retry loops. If a client retries on every timeout while the server is still generating, a single user causes the same answer to be produced five times. If the gateway accepts an idempotency key and merges identical in-progress requests, this waste disappears.
What you will do in the next lab
You set up a gateway in front of the mock server you built earlier. You add API key authentication, an RPM limit, token accounting, a TPM token bucket, Retry-After, per-model cost calculation, and finally blocking when a tenant budget is exceeded.