Building a Rate-Limiting, Token-Accounting Gateway
Goal
Build an LLM gateway yourself to implement limits and accounting in units of tokens rather than requests, and complete it with per-tenant budget control.
Why it matters
A limit of 60 requests per minute is almost meaningless in front of an LLM. If a tenant puts 200,000 tokens into every request, the number of requests is within the limit but the cost and GPU occupancy are hundreds of times those of other tenants, and those requests monopolize the KV cache and get other requests preempted. So the first design principle of an LLM gateway is to count tokens. The especially tricky parts of this lab are steps 4 and 8. Output tokens can be known only when generation ends, so if you check the budget only after completion, requests that arrive at the same time all judge themselves to be within budget, pass through, and then exceed it together. Reserving max_tokens in advance as an upper bound and settling after completion solves this problem. This is exactly the point where cost incidents happen in real services.
Steps
- Start
/root/gw/gateway.pyon 127.0.0.1:8171 and forwardPOST /v1/generateto the backend on 8170. The response must be identical to the backend's. Save the API key of tenantt1in/root/gw/key.txtand the key oft2in/root/gw/key2.txt, one line each — the grading of the later steps uses these two files. - Identify the tenant by the
X-API-Keyheader. The key-to-tenant mapping is in/opt/fixtures/llms/apikeys.json. If there is no key or the key is unknown, it is 401. - Apply a per-tenant limit of 10 requests per minute. The 11th request must be 429. Write
allowed=10 rejected=1in/root/gw/rpm.txt. GET /v1/usagereturns{"tenant":"...","prompt_tokens":<n>,"completion_tokens":<n>,"requests":<n>}. Both token values must be greater than 0.- Apply a limit of 2000 tokens per minute with a token bucket. A request that exceeds the limit must receive 429. Write
limit=2000 consumed=<n> rejected=<n>in/root/gw/tpm.txt, and rejected must be at least 1. - Every 429 response must carry a
Retry-Afterheader as an integer from 1 to 60. Writeretry_after=<정수>in/root/gw/retry.txt(the placeholder is an integer). - Calculate the cost with the per-model unit prices in
/opt/fixtures/llms/pricing.jsonand write the headertenant,model,prompt_tokens,completion_tokens,cost_usdand at least 2 rows in/root/gw/cost.csv. - Set the daily budget of tenant
t1to 0.01 dollars and return 402 when it is exceeded. Reserve in advance on amax_tokensbasis and settle with the actual usage after completion. Writebudget_usd=0.01 spent_usd=<수> blocked=truein/root/gw/budget.txt(the placeholder is a number).
Notes
- The backend at 127.0.0.1:8170 is not pre-started in this Pod. It only needs to take
{"prompt":..., "max_tokens":n}atPOST /generateand return{"text":..., "usage":{"prompt_tokens":n,"completion_tokens":n}}, so start by using the server you built in the earlier token streaming lab as it is, or by starting a minimal server with the same contract yourself. - Token bucket: defined by two values, capacity and refill rate. The capacity is the burst maximum, and the refill is the sustainable average.
- For streaming, a request cut off midway must also have the output tokens up to that point accounted for.
- Common mistake 1: checking the budget only after completion — concurrent requests all pass through and then exceed it together.
- Common mistake 2: counting input and output tokens together — the unit prices differ, so the cost is wrong.
Proxy the backend through the gateway
Start /root/gw/gateway.py on 127.0.0.1:8171 and forward POST /v1/generate to the backend on 8170. The response must be identical to the backend's. Save the API key of tenant t1 in /root/gw/key.txt and the key of t2 in /root/gw/key2.txt, one line each — the grading of the later steps uses these two files.
First build a proxy that passes everything through as it is, and then add the features one at a time.
Identify the tenant by API key
Identify the tenant by the X-API-Key header. The key-to-tenant mapping is in /opt/fixtures/llms/apikeys.json. If there is no key or the key is unknown, it is 401.
If there is no key or the key is unknown, it is 401. Map the keys to tenants.
Apply a per-minute request limit
Apply a per-tenant limit of 10 requests per minute. The 11th request must be 429. Write allowed=10 rejected=1 in /root/gw/rpm.txt.
A fixed window is the simplest. Think about which status code it is when the limit is exceeded.
Account for input and output tokens
GET /v1/usage returns {"tenant":"...","prompt_tokens":<n>,"completion_tokens":<n>,"requests":<n>}. Both token values must be greater than 0.
You must count the two separately, because the unit prices differ.
Apply a per-minute token limit
Apply a limit of 2000 tokens per minute with a token bucket. A request that exceeds the limit must receive 429. Write limit=2000 consumed=<n> rejected=<n> in /root/gw/tpm.txt, and rejected must be at least 1.
A token bucket suits it well. It allows a burst from a tenant that was quiet while keeping the long-term average.
Put the retry time in the 429
Every 429 response must carry a Retry-After header as an integer from 1 to 60. Write retry_after=<정수> in /root/gw/retry.txt (the placeholder is an integer).
Calculate the remaining time and give it as integer seconds. Without it, clients create waves.
Calculate the cost with per-model unit prices
Calculate the cost with the per-model unit prices in /opt/fixtures/llms/pricing.json and write the header tenant,model,prompt_tokens,completion_tokens,cost_usd and at least 2 rows in /root/gw/cost.csv.
It is usual for output to be more expensive than input. The price table is in the fixtures.
Block budget overruns and report usage
Set the daily budget of tenant t1 to 0.01 dollars and return 402 when it is exceeded. Reserve in advance on a max_tokens basis and settle with the actual usage after completion. Write budget_usd=0.01 spent_usd=<수> blocked=true in /root/gw/budget.txt (the placeholder is a number).
There is also the option of downgrading to a cheaper model instead of blocking. Here you implement blocking.