One Operator Slowed Down the Whole Cluster
Goal
You read the default FlowSchemas and PriorityLevelConfigurations in order, create an isolation slot for a runaway service account and prove with the response header that classification really changes, and then check the rule for when priority values overlap, a watch script that prevents exempt abuse, and the seat count metrics.
Why it matters
The API server classifies and isolates every incoming request. A FlowSchema is the classification table, and a request is decided by matching from the smallest matchingPrecedence and taking the first match. The PriorityLevelConfiguration that the schema points to sets the width of that slot — how many can run at the same time, and whether to queue or reject immediately on overflow. This structure matters for one reason. When there is a runaway client, you can design who starves. If you cannot read the classification table, the only thing you can do on the day 429s pile up is restart, and that means meeting the same surge again.
Steps
- Save all of the FlowSchemas the cluster puts in by default (to distinguish them from the ones you will create in this lab, they are defined as those whose name does not start with
ops-orzz-ops-) to/root/ops-apf/flowschemas.tsv— one line each of<matchingPrecedence>TAB<이름>TAB<우선순위 등급 이름>(name, priority level name), sorted in the order in which requests are examined (numeric sort). - Save all of the PriorityLevelConfigurations the cluster puts in by default (those whose names do not start with
ops-) to/root/ops-apf/levels.tsv— one line each of<이름>TAB<type>TAB<nominalConcurrencyShares>TAB<limitResponse.type>, in ascending name order. For slots that have no value, such as theExemptlevel, write-. - Create the namespace
ops-apfand create three service accounts —widget-operator,plain-reader,bulk-writer. Then create/root/ops-apf/classify.sh— it impersonates the service account given as an argument, sends one request, and prints only the name of the FlowSchema matched, determined from theX-Kubernetes-Pf-Flowschema-Uidresponse header, to standard output. Run it withplain-readerand save the result as one line to/root/ops-apf/baseline.tsv—plain-readerTAB<스키마 이름>(schema name). - In
/root/ops-apf/ops-noisy.yaml, write the PriorityLevelConfigurationops-noisy—type: Limited,nominalConcurrencyShares: 5,lendablePercent: 50,borrowingLimitPercent: 20,limitResponseistype: Queue, andqueuingisqueues: 16,handSize: 4,queueLengthLimit: 50. Apply it. - In
/root/ops-apf/ops-noisy-operator.yaml, write the FlowSchemaops-noisy-operator—matchingPrecedence: 950,priorityLevelConfiguration.nameisops-noisy,distinguisherMethod.typeisByUser, and the rule's subject iskind: ServiceAccount, theops-apfnamespace'swidget-operator, and resourceRules setsverbs,apiGroups,resources, andnamespacesall to*. Apply it. - Run
classify.shwithwidget-operatorandplain-readerrespectively and save the results as two lines to/root/ops-apf/classified.tsv—<어카운트 이름>TAB<스키마 이름>(account name, schema name), in ascending account name order. - In
/root/ops-apf/ops-bulk-a.yamland/root/ops-apf/zz-ops-bulk-b.yaml, write two FlowSchemas. Both targetops-apf's service accountbulk-writer, point toops-noisy, and have resourceRules all*.ops-bulk-ahasmatchingPrecedence: 960anddistinguisherMethod.type: ByUser, andzz-ops-bulk-bhasmatchingPrecedence: 940anddistinguisherMethod.type: ByNamespace. After applying both, save the result ofclassify.sh bulk-writeras two lines to/root/ops-apf/precedence.tsv—ops-bulk-aTAB960TAB<win 또는 lose>(win or lose),zz-ops-bulk-bTAB940TAB<win 또는 lose>(win or lose). And if the values of the two schemas had been the same 960, write on one line to/root/ops-apf/tie.txtonly the name of the one that would have won. - In
/root/ops-apf/exempt-allow.txt, write, one per line, the names of FlowSchemas that are allowed to point to theexemptpriority level (first check which such schemas actually exist in the cluster right now, and write only those names). And create/root/ops-apf/exempt-guard.sh— find all FlowSchemas that point toexempt, and printOK <이름>(name) if it is on the allow list andVIOLATION <이름>(name) if not, in name order, to standard output only, and it must exit with a non-zero code if there is even one violation. Running it in the current state must show no violations. - Read
apiserver_flowcontrol_nominal_limit_seatsfrom the API server's metrics and save it to/root/ops-apf/seats.tsv— one line each of<우선순위 등급 이름>TAB<좌석 수>(priority level name, seat count), in ascending name order.ops-noisymust be in it.
Reference
matchingPrecedenceis examined earlier the smaller the value.kubectl --v=8prints the response headers — the classification result is in there as a UID.- APF classification comes before authorization, so the header is attached even if a 403 comes back.
- If Dangling is True in a FlowSchema's status, that schema is being ignored.
- Common mistake: ordering priorities with a string sort so that 1000 comes before 2.
- Common mistake: making it point to exempt because you are in a hurry, removing the mechanism that protects the server entirely.
- Reference: https://kubernetes.io/docs/concepts/cluster-administration/flow-control/
Read the classification table in priority order
Save all of the FlowSchemas the cluster puts in by default (to distinguish them from the ones you will create in this lab, they are defined as those whose name does not start with ops- or zz-ops-) to /root/ops-apf/flowschemas.tsv — one line each of <matchingPrecedence> TAB <이름> TAB <우선순위 등급 이름> (name, priority level name), sorted in the order in which requests are examined (numeric sort).
Every incoming request is matched against the FlowSchemas one by one going down, and stops at the first match. What sets the matching order is matchingPrecedence, and unlike what the name suggests, the smaller the value, the earlier it is looked at. With a string sort, 1000 comes before 2, so use a numeric sort.
Read the width of each slot
Save all of the PriorityLevelConfigurations the cluster puts in by default (those whose names do not start with ops-) to /root/ops-apf/levels.tsv — one line each of <이름> TAB <type> TAB <nominalConcurrencyShares> TAB <limitResponse.type>, in ascending name order. For slots that have no value, such as the Exempt level, write -.
Only Limited levels get a share of the concurrency. Exempt is not examined at all, so it has no related fields. When you handle slots with no value with jq, // treats false as empty, so here it is safer to compare only null.
Measure which slot it goes to right now
Create the namespace ops-apf and create three service accounts — widget-operator, plain-reader, bulk-writer. Then create /root/ops-apf/classify.sh — it impersonates the service account given as an argument, sends one request, and prints only the name of the FlowSchema matched, determined from the X-Kubernetes-Pf-Flowschema-Uid response header, to standard output. Run it with plain-reader and save the result as one line to /root/ops-apf/baseline.tsv — plain-reader TAB <스키마 이름> (schema name).
APF classification happens before authorization — so even if a 403 comes back because there is no permission, the header is attached. To see the header you use kubectl --v=8 and must capture standard error too. Impersonation is --as=system:serviceaccount:<네임스페이스>:<이름> (namespace and name).
Create a slot for the runaway operator
In /root/ops-apf/ops-noisy.yaml, write the PriorityLevelConfiguration ops-noisy — type: Limited, nominalConcurrencyShares: 5, lendablePercent: 50, borrowingLimitPercent: 20, limitResponse is type: Queue, and queuing is queues: 16, handSize: 4, queueLengthLimit: 50. Apply it.
There is a reason the seat count is written as a share rather than directly — even if the server's total concurrency changes, all levels grow and shrink together at the same ratio. lendablePercent is the ratio of spare seats that can be lent out, and borrowingLimitPercent is the upper limit of what can be borrowed from others.
Write the rule that sends to that slot
In /root/ops-apf/ops-noisy-operator.yaml, write the FlowSchema ops-noisy-operator — matchingPrecedence: 950, priorityLevelConfiguration.name is ops-noisy, distinguisherMethod.type is ByUser, and the rule's subject is kind: ServiceAccount, the ops-apf namespace's widget-operator, and resourceRules sets verbs, apiGroups, resources, and namespaces all to *. Apply it.
A FlowSchema's status has a Dangling condition — if the priority level it points to does not exist, it becomes True and that schema is ignored. If you build the habit of checking after applying that this condition is False, you can catch right away the incident where a typo kills a schema entirely.
Prove with the header that classification really changed
Run classify.sh with widget-operator and plain-reader respectively and save the results as two lines to /root/ops-apf/classified.tsv — <어카운트 이름> TAB <스키마 이름> (account name, schema name), in ascending account name order.
Creating the rule does not mean the classification has changed. An account that is not a target must still match the old schema, and only the target account must go to the new schema. Leaving both lines is the evidence that the rule was not applied too broadly.
When two rules match the same request
In /root/ops-apf/ops-bulk-a.yaml and /root/ops-apf/zz-ops-bulk-b.yaml, write two FlowSchemas. Both target ops-apf's service account bulk-writer, point to ops-noisy, and have resourceRules all *. ops-bulk-a has matchingPrecedence: 960 and distinguisherMethod.type: ByUser, and zz-ops-bulk-b has matchingPrecedence: 940 and distinguisherMethod.type: ByNamespace. After applying both, save the result of classify.sh bulk-writer as two lines to /root/ops-apf/precedence.tsv — ops-bulk-a TAB 960 TAB <win 또는 lose> (win or lose), zz-ops-bulk-b TAB 940 TAB <win 또는 lose> (win or lose). And if the values of the two schemas had been the same 960, write on one line to /root/ops-apf/tie.txt only the name of the one that would have won.
When two schemas match the same request, the one examined first wins and the matching ends there. The rule for when the values are exactly the same is written in the official documentation — the names are compared alphabetically and the smaller one wins. However, the documentation recommends not setting the same value.
Watch the rules that point to exempt
In /root/ops-apf/exempt-allow.txt, write, one per line, the names of FlowSchemas that are allowed to point to the exempt priority level (first check which such schemas actually exist in the cluster right now, and write only those names). And create /root/ops-apf/exempt-guard.sh — find all FlowSchemas that point to exempt, and print OK <이름> (name) if it is on the allow list and VIOLATION <이름> (name) if not, in name order, to standard output only, and it must exit with a non-zero code if there is even one violation. Running it in the current state must show no violations.
A request sent to the exempt level is processed immediately with no concurrency limit and no queue. It looks convenient, but the mechanism that protects the server disappears by that much, so a person needs to know when a new schema starts pointing to this level. This kind of watch is good to run in CI.
Check with metrics whether the new slot actually received seats
Read apiserver_flowcontrol_nominal_limit_seats from the API server's metrics and save it to /root/ops-apf/seats.tsv — one line each of <우선순위 등급 이름> TAB <좌석 수> (priority level name, seat count), in ascending name order. ops-noisy must be in it.
You get the metrics with kubectl get --raw /metrics. Creating one more level does not increase the total seats; it divides the same total again according to the shares — so when you create a new level, the seat counts of the existing levels shrink together. If you look at those numbers directly, the meaning of a share becomes clear.