TT Lab
Get started
Learn Learning paths Courses

Policy as Code

Scope and failure modes: the order to follow on rollout day

Continue in TT Lab

In one sentence

The risk of a policy comes not from the content of the rule but from its scope and failure mode, and so the order on the day you turn a policy on is always narrow scope → warning → widening → blocking.

Why it was needed

If you collect records of policy adoption incidents, incidents caused by a wrong rule are rare. Most are one of two things. It was applied too broadly, or the behavior on failure was not decided.

The reason applying broadly is dangerous is clear from the nature of the cost. The cost of one policy that matches any resource is not the execution time of one policy run. It is a tax attached to every write request coming into the API server. A cluster has far more requests sent by controllers than by people. Leases, events, EndpointSlices, and Pod status updates flow constantly. A single wildcard policy catches every one of them once more.

The second is scarier. If you set a policy to look even at kube-system and the policy engine dies, under failurePolicy: Fail the cluster can no longer repair itself. This is because even the deployment meant to bring the engine back must pass that policy. It is a circular dependency.

How it works

You narrow scope in three layers.

Layer What it decides Example
Resource Which group, version, kind, and verb to look at Only CREATE and UPDATE of apps/v1 deployments
Namespace Which namespaces to apply it to Only places with a particular label, via namespaceSelector
Request Which requests among those to look at objectSelector, CEL conditions in matchConditions

The three layers are also the order of cost reduction. For a request screened out at the resource layer, the API server does not even think of the policy. The namespace layer comes next, and the request layer is last. So a policy that takes everything with a wildcard and filters it out with CEL conditions is working at the most expensive of the three layers. It produces the same result at a larger cost.

Excluding kube-system is not a matter of taste but a matter of leaving an escape hatch. The same goes for the namespace where the policy engine lives. Many organizations have a runbook step "delete the webhook configuration to bring the cluster back," but if you exclude it from the scope in the first place, the occasions to use that step shrink.

The failure mode — what failurePolicy trades off.

Value When the policy server cannot respond What you lose then
Fail Rejects the request Availability. If the engine dies, writes in that scope stop
Ignore Lets the request pass The guarantee. While the engine is dead, things come in unchecked

Neither is the right answer. Even within the same cluster, a different value fits each rule. Rules that guard a security boundary use Fail, and hygiene rules (label standards, recommended settings) use Ignore. And the more a rule chooses Fail, the more important narrowing its scope becomes. This is because when the scope is narrow, even if the engine dies, only that scope stops. Scope and failure mode are not values chosen separately but values chosen as a pair.

timeoutSeconds is the third knob. It is the time to wait for the webhook response, and if you set it long, the API server is held up that long during an outage. If you set it short, a slow response is treated as a failure and handed to failurePolicy. So you decide the timeout not as "generous" but as "a little longer than the time this rule takes when it is healthy."

Built-in policies and webhook policies have different failure modes. This difference changes design choices.

So in practice, placement usually splits like this. Move validation that can be expressed without an engine down into VAP and PSA to reduce the failure surface, and leave only what is truly needed in webhooks.

The order on the day you turn a policy on.

1) 좁은 범위로 시작한다 (한 네임스페이스, 한 종류, 한 동사)
2) 경고·감사로만 켜고 며칠 센다
3) 걸린 것을 고치거나 좁은 예외로 남긴다
4) 범위를 한 단계 넓히고 2..3 을 반복한다
5) 마지막에 차단으로 올린다 (정책 본문이 아니라 켜는 쪽만 고친다)
6) 되돌리는 절차를 미리 적어 둔다

Many teams swap the order of steps 4 and 5. They raise to blocking right away in the narrow scope and then widen the scope. Then the day you widen the scope becomes the day of the incident — because the newly included namespaces run straight into blocking without observation. Not doing the widening and the raising to blocking on the same day is the heart of this order.

What you see in the field

First, the latency tax of a wildcard policy. Cases are common where one webhook matching every resource raised the whole tail of the API server latency distribution. The cause is hard to find because the policy itself is fast — what is slow is not one request but all the requests.

Second, a policy that blocked itself. The engine's namespace was not excluded from the scope, so the request to redeploy the engine asks the dead engine and gets blocked. The only way to bring it back is to delete the webhook configuration, so the outage grows longer while people look for someone who knows that procedure.

Third, turning it on with Ignore and forgetting. If you turn everything on with Ignore to avoid incidents, nobody knows what came in unchecked while the engine was dead for hours. If you chose Ignore, an engine availability alert must be part of the policy.

Fourth, an engine crowded onto one node. Even with three replicas, if they are on the same node, they disappear together when that node drops. Under Fail, that means writes stop. Topology spreading and a PodDisruptionBudget are not options for policy adoption but prerequisites.

Fifth, what you can see in this environment. In the kwok cluster, you can create a webhook configuration with an address that cannot be reached and really see the difference between Fail and Ignore. The same request is blocked on one side and passes on the other.

References

What you will do in the next lab

You deliberately break a policy and observe the breakage. You create a webhook configuration with an address that cannot be reached and confirm with the same request that it is blocked under failurePolicy: Fail and passes under Ignore, and see the escape hatch appear when you remove system namespaces from the scope with namespaceSelector. You narrow the scope at each of the three layers, resource, namespace, and request, and count which requests pass through the policy and which do not, and finally you go around the narrow scope → warning → widening → blocking order once and build even the procedure for rolling it back.