One webhook was switched to Fail and the cluster stopped
Goal
You narrow the scope of an admission webhook yourself at three layers, truly create a state where the policy server is dead and measure time to confirm what failurePolicy and timeoutSeconds trade off, and build a report and a checker that answer "if this dies, what stops?" for every webhook in the cluster.
Why it matters
If you collect records of policy adoption incidents, incidents caused by a wrong rule are rare. Most are one of two things — it was applied too broadly, or the behavior on failure was not decided. The cost of applying broadly is not the execution time of one policy run but a latency tax attached to every incoming write request, and a cluster has far more requests sent by controllers than by people. The second is scarier. If you set a policy to look even at kube-system and the policy server dies, under failurePolicy: Fail the cluster can no longer repair itself — because even the deployment meant to bring the server back must pass that policy. Scope and failure mode are not values chosen separately but values chosen as a pair, and this lab has you run that pair by hand. The blast radius report you build at the end is something to build not before you turn a policy on but right now, in a cluster where it is already on.
Steps
- Work in
/root/polscope(export KUBECONFIG=/root/.kube/config,kubectl config use-context kwok-lab). First create the namespacesscope-app,scope-ops, andscope-builtin, and in all three, create thedefaultservice account yourself withkubectl create sa default -n <네임스페이스>(replacing the placeholder with the namespace; this cluster has no controller manager, so it does not appear on its own). In/root/polscope/webhook-wide.yaml, write a ValidatingWebhookConfigurationscope-wide-audit— the webhook name iswide.polscope.local,admissionReviewVersionsis["v1"],sideEffectsisNone,failurePolicyisIgnore,timeoutSecondsis2, andclientConfig.urlishttps://192.0.2.77:8443/validate. Setrulesas broadly as possible —apiGroups: ["*"],apiVersions: ["*"],operations: ["CREATE", "UPDATE"],resources: ["*"], andscope: "*"— and put in no selector at all. Also create three test manifests —/root/polscope/pod.yaml(Podprobe-app, containerweb, imageregistry.internal/app:1.0),/root/polscope/pod-guard.yaml(Podprobe-guard, labelpolscope.io/guard: "on", the same container), and/root/polscope/cm.yaml(ConfigMapprobe-cm, dataa: b). After applying, measure the elapsed time of four requests —kubectl create -n scope-app -f pod.yaml --dry-run=server,kubectl create -n scope-app -f cm.yaml --dry-run=server,kubectl create ns probe-ns --dry-run=server, andkubectl get pods -n scope-app. If it took one second or more, writeSLOW, and otherwiseFAST, in/root/polscope/01-reach.txtas four lines,pod-create,configmap-create,namespace-create, andpod-list, in the form<이름> <낱말>(the name, then the word). - In
/root/polscope/webhook-strict.yaml, write a second configurationscope-strict— the webhook name isstrict.polscope.local,rulesis exactly as broad asscope-wide-audit, and only three things differ:failurePolicyisFail,timeoutSecondsis3, andclientConfig.urlis the address where the connection is refused,https://127.0.0.1:19443/validate. Do not deletescope-wide-audit; leave it as it is — two configurations standing side by side with the same scope and only the failure mode different is the control group of this lab. After applying, try four things to see whether each is blocked, and writeBLOCKif it is blocked andPASSif it passes, in/root/polscope/02-blocked.txtas four lines in the form<이름> <낱말>(the name, then the word) —pod-create-scope-app(kubectl create -n scope-app -f pod.yaml --dry-run=server),configmap-kube-system(kubectl create -n kube-system -f cm.yaml --dry-run=server),namespace-create(kubectl create ns probe-ns --dry-run=server), andwebhookconfig-edit(kubectl apply -f webhook-strict.yamlonce more). - First try to get out with an exemption label — run
kubectl label ns kube-system polscope.io/admission=exempt --overwriteand save its output, including standard error, to/root/polscope/03-lockout.txt(it is blocked). Then add twonamespaceSelector.matchExpressionstoscope-strictin/root/polscope/webhook-strict.yamland reapply it — (1) keykubernetes.io/metadata.name, operatorNotIn, values["kube-system", "kube-node-lease", "kube-public"], and (2) keypolscope.io/admission, operatorNotIn, values["exempt"]. Now the label can be attached — hangpolscope.io/admission=exemptonkube-systemandscope-ops. Finally, send/root/polscope/pod-guard.yamlto three places with--dry-run=serverand write three lines,kube-system,scope-ops, andscope-app, in/root/polscope/03-escape.txtin the form<네임스페이스> <BLOCK|PASS>(the namespace, then BLOCK or PASS). - Narrow
scope-strictin/root/polscope/webhook-strict.yamlat two layers — makerulesonlyCREATEandUPDATEonv1podsin the core group (""), changescopetoNamespaced, and useobjectSelector.matchLabelsto make it look only at objects with the labelpolscope.io/guard: "on". Leave thenamespaceSelectorfrom step 3 as it is. After applying, send the four requests again and write four lines in/root/polscope/04-narrow.txtin the form<이름> <좁히기전> <좁힌뒤>(the name, the value before narrowing, and the value after narrowing; the value before narrowing is what you saw in step 2) —pod-guard-scope-app(pod-guard.yamltoscope-app),pod-plain-scope-app(pod.yamltoscope-app),configmap-scope-app(cm.yamltoscope-app), andnamespace-create(kubectl create ns probe-ns --dry-run=server). The decision isBLOCKorPASS. - Add two
matchConditionstoscope-strictin/root/polscope/webhook-strict.yamland reapply it —skip-kube-system-samakes it not call the webhook if the requester is a service account that starts withsystem:serviceaccount:kube-system:, andskip-break-glassmakes it not call if the requester's groups includepolscope:break-glass. After applying, send the same/root/polscope/pod-guard.yamltoscope-appas three kinds of subject and write three lines,kube-system-sa,break-glass, andnormal, in/root/polscope/05-conditions.txtin the form<이름> <BLOCK|PASS>(the name, then BLOCK or PASS) —kubectl --as=system:serviceaccount:kube-system:replicaset-controller create -n scope-app -f pod-guard.yaml --dry-run=server,kubectl --as=oncall --as-group=polscope:break-glass --as-group=system:masters create -n scope-app -f pod-guard.yaml --dry-run=server, and once as the current account. - In
/root/polscope/webhook-slow.yaml, write a third configurationscope-slow— the webhook name isslow.polscope.local,clientConfig.urlis the address from which no response ever comes,https://192.0.2.77:8443/validate,sideEffectsisNone,namespaceSelector.matchExpressionsis the keykubernetes.io/metadata.nameIn["scope-ops"],objectSelector.matchLabelsispolscope.io/slow: "on",rulesisCREATEandUPDATEonv1podsof the core group (scope: "Namespaced"), and the final state isfailurePolicy: FailandtimeoutSeconds: 3. In/root/polscope/pod-slow.yaml, create a Podprobe-slow(labelpolscope.io/slow: "on", containerweb, imageregistry.internal/app:1.0). And create/root/polscope/timeout-probe.sh <Fail|Ignore> <초>(taking the failurePolicy and the timeout in seconds as arguments) — briefly changescope-slowto those two values, sendpod-slow.yamlonce toscope-opswith--dry-run=serverand measure the elapsed time, always restore the original values, and then print only one line,BLOCK <초>orPASS <초>(the seconds as a rounded integer). Measure three times with this script and write three lines in/root/polscope/06-timeout.txtin the form<failurePolicy> <타임아웃> <BLOCK|PASS> <걸린초>(failurePolicy, timeout, BLOCK or PASS, and elapsed seconds) — once each forFail 6,Ignore 6, andFail 4. - As it is now, with no webhook server, you set up built-in policies in the same place to compare. Hang the label
pod-security.kubernetes.io/enforce=baselineon the namespacescope-builtin. In/root/polscope/vap-owner.yaml, write together and apply a ValidatingAdmissionPolicypolscope-require-owner(it catchesCREATEandUPDATEonv1podsin the core group and requires that the Pod metadata have anownerlabel) and a ValidatingAdmissionPolicyBinding of the same name (validationActionsis["Deny"], andmatchResources.namespaceSelectoriskubernetes.io/metadata.nameIn["scope-builtin"]). Create three more test Pods —/root/polscope/pod-owned.yaml(Podprobe-owned, labelowner: platform),/root/polscope/pod-host.yaml(Podprobe-host, the same label withspec.hostNetwork: true), and/root/polscope/pod-both.yaml(Podprobe-both, labelsowner: platformandpolscope.io/guard: "on"). All three have the containerweb/registry.internal/app:1.0. Send four manifests toscope-builtinwith--dry-run=serverand write which one made the decision in/root/polscope/07-builtin.txtas four lines in the form<파일이름> <낱말>(the file name, then the word) — the word is one ofVAP(rejected by the ValidatingAdmissionPolicy),PSA(rejected by PodSecurity),WEBHOOK(webhook call failed), andPASS(passed), and the targets arepod.yaml,pod-host.yaml,pod-both.yaml, andpod-owned.yaml. - Create
/root/polscope/blast-radius.sh— it sweeps all ValidatingWebhookConfigurations in the cluster and outputs to standard output a JSON array holding one object per webhook. There are exactly eight keys:config(the configuration name) ·webhook(the webhook name) ·failurePolicy("Fail"if absent) ·timeoutSeconds(10if absent) ·resources(an array of all of that webhook'srules[].resources, sorted without duplicates) ·wildcard(true ifresourcescontains"*") ·excludesKubeSystem(true ifnamespaceSelector.matchExpressionshas an entry whose key iskubernetes.io/metadata.name, whose operator isNotIn, and whose values containkube-system) ·risk.riskis"high"ifwildcardis true,failurePolicyisFail, andexcludesKubeSystemis false; otherwise"medium"iffailurePolicyisFail; and"low"for the rest. Sort the array byconfigand thenwebhook. Save that output to/root/polscope/blast-radius.json. And create/root/polscope/risky.sh— it prints, one per line in the form<config>/<webhook>, only those whoseriskishigh, and ends with exit code 1 if there is even one, and with 0 and no output if there is none. Both scripts must read the objects standing in the cluster right now, not a saved file.
Notes
- You start with
export KUBECONFIG=/root/.kube/configandkubectl config use-context kwok-lab. It is a real kube-apiserver v1.30.4 that kwok runs inside the Pod, and you keep all outputs under/root/polscope. - kwok does not actually run Pods. All this lab looks at is the admission stage, so
kubectl create -f <파일> --dry-run=server(with a file name in place of the placeholder) is enough — it runs admission as it is and leaves no object. - This cluster has no controller manager, so a
defaultservice account is not created in a new namespace. If you do not create it yourself in step 1 withkubectl create sa default -n <네임스페이스>(with the namespace in place of the placeholder), a Pod request is rejected first witherror looking up service accountwithout even reaching the webhook. - Right after you change a webhook configuration, it is reflected a round trip or two late. If the result looks like the old one, send it again a few seconds later.
- The audit webhook from step 1 stays alive until this lab ends. So every write in this cluster is about 2 seconds slow — that slowness itself is "the latency tax of a broadly applied policy." Several webhooks are called in parallel, so the elapsed time is not the sum but the larger one.
- Common mistake: setting up a broad
Failwebhook and then trying to get out with a label. Attaching a label is also a write, so that request is blocked first. Only objects ofadmissionregistration.k8s.iodo not go through admission webhooks, so the escape hatch is always the webhook configuration itself. - Common mistake: pretending to be a subject with
--asand forgetting authorization. You get only the permissions of the impersonated subject, so you must also give--as-group=system:mastersfor the request to pass authorization and reach admission. - Dynamic Admission Control · Admission Controllers Reference · Validating Admission Policy · Pod Security Admission · Kubernetes API Concepts
I set one webhook broadly and every write got 2 seconds slower
Work in /root/polscope (export KUBECONFIG=/root/.kube/config, kubectl config use-context kwok-lab). First create the namespaces scope-app, scope-ops, and scope-builtin, and in all three, create the default service account yourself with kubectl create sa default -n <네임스페이스> (replacing the placeholder with the namespace; this cluster has no controller manager, so it does not appear on its own). In /root/polscope/webhook-wide.yaml, write a ValidatingWebhookConfiguration scope-wide-audit — the webhook name is wide.polscope.local, admissionReviewVersions is ["v1"], sideEffects is None, failurePolicy is Ignore, timeoutSeconds is 2, and clientConfig.url is https://192.0.2.77:8443/validate. Set rules as broadly as possible — apiGroups: ["*"], apiVersions: ["*"], operations: ["CREATE", "UPDATE"], resources: ["*"], and scope: "*" — and put in no selector at all. Also create three test manifests — /root/polscope/pod.yaml (Pod probe-app, container web, image registry.internal/app:1.0), /root/polscope/pod-guard.yaml (Pod probe-guard, label polscope.io/guard: "on", the same container), and /root/polscope/cm.yaml (ConfigMap probe-cm, data a: b). After applying, measure the elapsed time of four requests — kubectl create -n scope-app -f pod.yaml --dry-run=server, kubectl create -n scope-app -f cm.yaml --dry-run=server, kubectl create ns probe-ns --dry-run=server, and kubectl get pods -n scope-app. If it took one second or more, write SLOW, and otherwise FAST, in /root/polscope/01-reach.txt as four lines, pod-create, configmap-create, namespace-create, and pod-list, in the form <이름> <낱말> (the name, then the word).
If the webhook's address is in 192.0.2.0/24 (TEST-NET-1), the connection is neither refused nor answered. The API server waits for timeoutSeconds and then hands over to failurePolicy, so with Ignore every request passes, but the elapsed time becomes evidence that the request was inside the scope. A request outside the scope ends with no waiting. Admission hooks only onto writes — a read request does not pass through the webhook at all. You measure time in nanoseconds by subtracting between s=$(date +%s%N) and e=$(date +%s%N). kubectl create -f <파일> --dry-run=server (with a file name in place of the placeholder) runs admission as it is but leaves no object. The webhook call and the rejection message come out just like a real request.
I changed only the failure mode to Fail and every write in that scope stopped
In /root/polscope/webhook-strict.yaml, write a second configuration scope-strict — the webhook name is strict.polscope.local, rules is exactly as broad as scope-wide-audit, and only three things differ: failurePolicy is Fail, timeoutSeconds is 3, and clientConfig.url is the address where the connection is refused, https://127.0.0.1:19443/validate. Do not delete scope-wide-audit; leave it as it is — two configurations standing side by side with the same scope and only the failure mode different is the control group of this lab. After applying, try four things to see whether each is blocked, and write BLOCK if it is blocked and PASS if it passes, in /root/polscope/02-blocked.txt as four lines in the form <이름> <낱말> (the name, then the word) — pod-create-scope-app (kubectl create -n scope-app -f pod.yaml --dry-run=server), configmap-kube-system (kubectl create -n kube-system -f cm.yaml --dry-run=server), namespace-create (kubectl create ns probe-ns --dry-run=server), and webhookconfig-edit (kubectl apply -f webhook-strict.yaml once more).
An address where the connection is refused comes back as an immediate failure with no waiting. Fail turns that failure into a rejection and Ignore turns it into a pass — same scope, same dead server, but opposite answers. This is what failurePolicy trades off. The fourth line is the key point of this step. If one resource does not go through admission webhooks, then no matter what is blocked, that one can still be fixed — because if the request to fix a webhook configuration went to ask that very webhook, nobody could bring the cluster back. Do not predict the results; run the four yourself.
I tried to exclude with a label and even attaching the label was blocked
First try to get out with an exemption label — run kubectl label ns kube-system polscope.io/admission=exempt --overwrite and save its output, including standard error, to /root/polscope/03-lockout.txt (it is blocked). Then add two namespaceSelector.matchExpressions to scope-strict in /root/polscope/webhook-strict.yaml and reapply it — (1) key kubernetes.io/metadata.name, operator NotIn, values ["kube-system", "kube-node-lease", "kube-public"], and (2) key polscope.io/admission, operator NotIn, values ["exempt"]. Now the label can be attached — hang polscope.io/admission=exempt on kube-system and scope-ops. Finally, send /root/polscope/pod-guard.yaml to three places with --dry-run=server and write three lines, kube-system, scope-ops, and scope-app, in /root/polscope/03-escape.txt in the form <네임스페이스> <BLOCK|PASS> (the namespace, then BLOCK or PASS).
The two conditions are two ways of doing the same thing. The first uses the kubernetes.io/metadata.name label that the API server automatically attaches to every namespace since 1.21, so it needs no preparation, and the second manages the exemption list with labels you attach yourself, so you do not have to edit the webhook configuration when you add namespaces later. But in a cluster that is already blocked, you cannot use the second — attaching a label is a namespace UPDATE, and that very request is blocked. So the order is fixed: first make the escape hatch with the first method, and then layer the second method on top. Multiple matchExpressions of a namespaceSelector must all be satisfied to be in scope.
Once the scope was narrowed, three of the four that used to get caught stopped going through the webhook
Narrow scope-strict in /root/polscope/webhook-strict.yaml at two layers — make rules only CREATE and UPDATE on v1 pods in the core group (""), change scope to Namespaced, and use objectSelector.matchLabels to make it look only at objects with the label polscope.io/guard: "on". Leave the namespaceSelector from step 3 as it is. After applying, send the four requests again and write four lines in /root/polscope/04-narrow.txt in the form <이름> <좁히기전> <좁힌뒤> (the name, the value before narrowing, and the value after narrowing; the value before narrowing is what you saw in step 2) — pod-guard-scope-app (pod-guard.yaml to scope-app), pod-plain-scope-app (pod.yaml to scope-app), configmap-scope-app (cm.yaml to scope-app), and namespace-create (kubectl create ns probe-ns --dry-run=server). The decision is BLOCK or PASS.
Narrowing the scope is not reducing security but reducing the number of times you ask. A request screened out at rules does not even make the API server think of this webhook, and namespaceSelector and objectSelector come next. If you take everything with a wildcard and filter inside the webhook server, the result is the same but you pay the whole cost. scope is one of Namespaced, Cluster, and *, and if you write that you want to see only namespaced resources, requests for cluster-scoped resources never arrive at all. objectSelector looks at the labels of the object in the request, and namespaceSelector looks at the labels of the namespace that holds that object — they are different layers.
Open a door that only controllers and the on-call engineer pass through
Add two matchConditions to scope-strict in /root/polscope/webhook-strict.yaml and reapply it — skip-kube-system-sa makes it not call the webhook if the requester is a service account that starts with system:serviceaccount:kube-system:, and skip-break-glass makes it not call if the requester's groups include polscope:break-glass. After applying, send the same /root/polscope/pod-guard.yaml to scope-app as three kinds of subject and write three lines, kube-system-sa, break-glass, and normal, in /root/polscope/05-conditions.txt in the form <이름> <BLOCK|PASS> (the name, then BLOCK or PASS) — kubectl --as=system:serviceaccount:kube-system:replicaset-controller create -n scope-app -f pod-guard.yaml --dry-run=server, kubectl --as=oncall --as-group=polscope:break-glass --as-group=system:masters create -n scope-app -f pod-guard.yaml --dry-run=server, and once as the current account.
The scope (rules and selectors) and the conditions (matchConditions) are different layers. A selector looks only at what the request is about, while a condition looks at the whole request with CEL — even who sent it (request.userInfo), what verb it is, and what namespace it is in. But conditions are evaluated after selectors, so if you defer to conditions what a selector could have filtered out, you only add cost. A webhook is called only when all the conditions are true, so you write "I want to skip" as a negation. You can pretend to be another subject with --as, but you get only the impersonated subject's permissions, so to pass authorization too you give --as-group=system:masters along with it (authorization comes before admission). If you block even the requests controllers send, the cluster can no longer repair itself, and without a door for the on-call engineer, there is no path during an outage other than deleting the webhook configuration.
I set the timeout to 6 seconds and the outage got 6 seconds longer for every request
In /root/polscope/webhook-slow.yaml, write a third configuration scope-slow — the webhook name is slow.polscope.local, clientConfig.url is the address from which no response ever comes, https://192.0.2.77:8443/validate, sideEffects is None, namespaceSelector.matchExpressions is the key kubernetes.io/metadata.name In ["scope-ops"], objectSelector.matchLabels is polscope.io/slow: "on", rules is CREATE and UPDATE on v1 pods of the core group (scope: "Namespaced"), and the final state is failurePolicy: Fail and timeoutSeconds: 3. In /root/polscope/pod-slow.yaml, create a Pod probe-slow (label polscope.io/slow: "on", container web, image registry.internal/app:1.0). And create /root/polscope/timeout-probe.sh <Fail|Ignore> <초> (taking the failurePolicy and the timeout in seconds as arguments) — briefly change scope-slow to those two values, send pod-slow.yaml once to scope-ops with --dry-run=server and measure the elapsed time, always restore the original values, and then print only one line, BLOCK <초> or PASS <초> (the seconds as a rounded integer). Measure three times with this script and write three lines in /root/polscope/06-timeout.txt in the form <failurePolicy> <타임아웃> <BLOCK|PASS> <걸린초> (failurePolicy, timeout, BLOCK or PASS, and elapsed seconds) — once each for Fail 6, Ignore 6, and Fail 4.
timeoutSeconds is a value you set "a little longer than the time it takes when the webhook server is healthy." If you set it generously, the API server is held up that long on every request during an outage, and if you set it short, a slow response is treated as a failure and handed to failurePolicy. So the two knobs are not values chosen separately — with Fail, the longer the timeout, the longer the outage, and with Ignore, the longer the timeout, the more it passes but gets slow. Neither is free. In the script, hang the restoration with trap ... EXIT. Even if it dies midway, the webhook must not be left at a strange value. After changing the configuration, it is reflected a round trip or two late, so wait briefly before measuring. In this cluster the 2-second audit webhook from step 1 stays alive and webhooks are called in parallel, so the elapsed time comes out as the larger of the two.
In a cluster where every webhook was dead, only the built-in policies produced decisions
As it is now, with no webhook server, you set up built-in policies in the same place to compare. Hang the label pod-security.kubernetes.io/enforce=baseline on the namespace scope-builtin. In /root/polscope/vap-owner.yaml, write together and apply a ValidatingAdmissionPolicy polscope-require-owner (it catches CREATE and UPDATE on v1 pods in the core group and requires that the Pod metadata have an owner label) and a ValidatingAdmissionPolicyBinding of the same name (validationActions is ["Deny"], and matchResources.namespaceSelector is kubernetes.io/metadata.name In ["scope-builtin"]). Create three more test Pods — /root/polscope/pod-owned.yaml (Pod probe-owned, label owner: platform), /root/polscope/pod-host.yaml (Pod probe-host, the same label with spec.hostNetwork: true), and /root/polscope/pod-both.yaml (Pod probe-both, labels owner: platform and polscope.io/guard: "on"). All three have the container web/registry.internal/app:1.0. Send four manifests to scope-builtin with --dry-run=server and write which one made the decision in /root/polscope/07-builtin.txt as four lines in the form <파일이름> <낱말> (the file name, then the word) — the word is one of VAP (rejected by the ValidatingAdmissionPolicy), PSA (rejected by PodSecurity), WEBHOOK (webhook call failed), and PASS (passed), and the targets are pod.yaml, pod-host.yaml, pod-both.yaml, and pod-owned.yaml.
VAP evaluates CEL inside the API server process and PSA is a built-in admission plugin. Neither has a network round trip, a certificate, or an engine Pod, so even now, with every webhook dead, they judge exactly as usual. Failure comes only in local forms such as an "expression evaluation error" or a "wrongly attached label." So the placement in practice is to move validation that can be expressed without an engine down into built-in mechanisms to reduce the failure surface, and to leave only what is truly needed in webhooks. Also pay attention to the decision order — built-in admission comes first, so a request already rejected by a built-in policy does not even reach the webhook. So to see a webhook's failure, you must send a Pod that passes all the built-in policies. What the baseline level blocks is told directly by the rejection message.
Answer "if this dies, what stops?" on a single page
Create /root/polscope/blast-radius.sh — it sweeps all ValidatingWebhookConfigurations in the cluster and outputs to standard output a JSON array holding one object per webhook. There are exactly eight keys: config (the configuration name) · webhook (the webhook name) · failurePolicy ("Fail" if absent) · timeoutSeconds (10 if absent) · resources (an array of all of that webhook's rules[].resources, sorted without duplicates) · wildcard (true if resources contains "*") · excludesKubeSystem (true if namespaceSelector.matchExpressions has an entry whose key is kubernetes.io/metadata.name, whose operator is NotIn, and whose values contain kube-system) · risk. risk is "high" if wildcard is true, failurePolicy is Fail, and excludesKubeSystem is false; otherwise "medium" if failurePolicy is Fail; and "low" for the rest. Sort the array by config and then webhook. Save that output to /root/polscope/blast-radius.json. And create /root/polscope/risky.sh — it prints, one per line in the form <config>/<webhook>, only those whose risk is high, and ends with exit code 1 if there is even one, and with 0 and no output if there is none. Both scripts must read the objects standing in the cluster right now, not a saved file.
This report answers one question — if this webhook server dies, what stops in the cluster? With failurePolicy: Fail, writes in that scope stop, and that scope is resources and the selectors. So the dangerous combination is always the same: broad scope + Fail + system namespaces not excluded. When all three hold at once, even the requests meant to bring the engine back are blocked and the cluster cannot repair itself. If you unfold .items[].webhooks[] of kubectl get validatingwebhookconfigurations -o json with jq, you can build it in one go. The // in jq also leaks to the right side when the value is false, so fill in defaults by first asking with has("키") (the key name in quotes) whether the key exists. The grader briefly sets up a risky configuration and calls risky.sh again — an answer with the names memorized fails then.