TT Lab
Get started
Learn Learning paths Courses

Kubernetes Operations

We Decide Who Gets the 429

Continue in TT Lab

Summary

The API server classifies every request with a FlowSchema, puts it into a slot called a PriorityLevelConfiguration, and divides concurrency among the slots. You can use these two objects to design ahead of time who starves when a runaway client appears.

Why this mechanism was needed

The old kube-apiserver had only two knobs, --max-requests-inflight and --max-mutating-requests-inflight. They limit the total number of requests handled at the same time. The problem with this approach is clear — it does not distinguish who took the spot. If one badly made controller lists resources thousands of times per second, it eats up all the spots, and the node status updates from the kubelets that came after it, and an operator's kubectl get pods, are pushed back equally. There was no way to separate important requests from noisy ones.

API Priority and Fairness splits this situation into two steps. First it classifies requests, and then it gives each classified slot its own width. In addition, a brief surge is not rejected but queued for a moment. It became a stable feature in 1.29, and the API group is flowcontrol.apiserver.k8s.io/v1.

How it works

A FlowSchema is the classification table. Incoming requests are matched in order starting from the one with the numerically smallest matchingPrecedence, and stop at the first match. The point that a smaller value comes first, unlike what the name suggests, is a spot that always confuses people. If two schemas have exactly the same value, the names are compared alphabetically and the smaller one wins, but the official documentation recommends not relying on such a situation and keeping the values from overlapping.

A rule is written as a pair of subjects (who) and resourceRules/nonResourceRules (what). Using * in a field means that item is not considered at all. And distinguisherMethod divides requests within one slot into finer flows. With ByUser, one user cannot starve another user's share, and with ByNamespace, the same protection is applied per namespace. If you leave it empty, all requests that match that schema are treated as a single flow.

A PriorityLevelConfiguration is the slot. If type is Exempt, it is not examined at all and is processed immediately. If it is Limited, it receives a division of the concurrency, written not as an absolute seat count but as a share called nominalConcurrencyShares. The design intends that when you change the server's total concurrency, all slots grow and shrink together at the same ratio. On top of that, lendablePercent (the ratio of spare seats that can be lent out) and borrowingLimitPercent (the upper limit of what can be borrowed) are added, so that when one slot is idle, its seats flow to other slots.

The behavior on overflow is set by limitResponse. With Reject, it is an immediate 429, and with Queue, it queues. The shape of the queue is set by the three values queues, handSize, and queueLengthLimit, and these three adjust "the probability of an elephant trampling a mouse" with a technique called shuffle sharding. The maximum number of requests a single flow can pile up is handSize times queueLengthLimit.

The defaults are also worth knowing. Of the four mandatory objects, the exempt schema takes requests from the system:masters group out of examination, and catch-all guarantees a place for requests that match nothing. In addition, schemas such as system-leader-election, system-nodes, kube-controller-manager, and service-accounts come in as suggested configuration.

Observation is handled by the metrics that start with apiserver_flowcontrol_. The rejection count is apiserver_flowcontrol_rejected_requests_total, and the reason label gives the reason as one of queue-full, concurrency-limit, time-out, and cancelled. How many seats each slot actually received shows up as it is in apiserver_flowcontrol_nominal_limit_seats.

What it looks like in the field

The most common is a runaway from one operator. When a controller that does not use a cache and lists everything each time is deployed, the other service accounts that use the same slot slow down with it. What you should do then is add a FlowSchema that sends only that account to a narrow slot. At that moment the damage is confined inside that slot.

The second is abuse of exempt. A 429 appears, so in a hurry you send that workload to exempt. The symptom disappears, but now that client bypasses the entire mechanism that protects the server. At the next surge, the API server goes down with it. That is why the list of schemas that point to exempt is a short list that a person must watch.

The third is a silently dead schema. If you get one character wrong in the name of the priority level it points to, no error appears and only the Dangling condition is left as True in the status. That schema behaves as if it does not exist, and the workload you meant to isolate keeps starving others in the default slot.

Limits of this lab environment

This cluster has no workload that can actually run away — a kwok Pod is not a container, so you cannot run a client inside it. So we do not do an experiment that really produces a 429. Instead, we measure how classification comes out. kubectl --v=8 prints a header, X-Kubernetes-Pf-Flowschema-Uid, that contains the UID of the schema the request actually matched, and since classification happens before authorization, the header is still attached even if a request imitating an account with no permissions ends in a 403. The seat counts can also be confirmed from real metrics.

What to do in the next lab

You start by reading and saving the default classification table and the list of slots in priority order. Then you impersonate service accounts and send requests, and measure with the response header "which slot does it go to right now." You create a narrow slot for a runaway operator and a rule that sends it there, and prove with the same method that the classification changed. Next you measure who wins when two rules match one request, build a script that watches for schemas that point to exempt, and then check in the metrics how many seats the new slot actually received.

References: