When the Authorizer Dies, Block or Let Through?
In one line
ext_authz is an HTTP filter that, for every request, asks an authorization server outside the proxy, "may this request be let in?" There are two ways to ask (HTTP and gRPC), and whether to block or let through when the authorization server cannot answer is decided by one line of configuration (failure_mode_allow).
Why this was needed
Local rate limiting and fault injection are things Envoy can decide just by looking at its configuration. Authorization is not like that. "Can the owner of this token see this order" can be answered only by the side that knows the user DB, the permission table and the organization's policy, and if you put that judgment into the code of each service, you get as many different authorizations as there are services. If just one service is slow to fix a rule, that service becomes a hole.
So a structure arose in which the judgment is gathered in one place (the authorization server), and the proxy holds the request and asks that server. Attaching OPA next to Envoy and running an in-house authorization service both use this filter. If you use action: CUSTOM and provider in Istio's AuthorizationPolicy, istiod looks at the mesh configuration (envoyExtAuthzHttp and envoyExtAuthzGrpc of extensionProviders) and puts exactly this filter into the sidecar.
How it works
HTTP mode — Envoy sends one more request to the authorization server with the original request's method and path. The body is emptied (Content-Length: 0) and headers are loaded selectively. The only ones loaded by default are Host, Method, Path, Content-Length and Authorization, and the rest are passed only if you write them in allowed_headers. If the authorization server returns a 2xx, it is allowed; otherwise its status code and body go to the client as they are.
gRPC mode — it does not resend the request; it passes the attributes of the request (source, destination, headers, path, TLS information) in a protobuf called CheckRequest. The service name is envoy.service.auth.v3.Authorization and the method is Check. Here, if you leave allowed_headers empty, all headers are passed. The documentation says the defaults of the two modes are opposite for the same filter because of compatibility, to preserve old behavior. If you change modes while changing the authorization server, the headers your policy relied on can quietly disappear.
Headers flow in both directions.
| Setting | When | Where to |
|---|---|---|
allowed_upstream_headers |
Allowed | Attached to the upstream request (a header with the same name is overwritten) |
allowed_client_headers |
Denied | Attached to the client response |
gRPC OkHttpResponse.headers |
Allowed | Attached to the upstream request |
The overwriting property is important. If the authorization server decides x-user: alice, even if the client sends admin in the same header, alice is what reaches the upstream. If the upstream trusts that header to judge permissions, this property is the last wall against forgery.
When the authorization server cannot answer — connection failure, timeout and 5xx fall here.
failure_mode_allow: false status_on_error(기본 403)를 돌려준다 ← 막는다(fail closed)
failure_mode_allow: true 요청을 그대로 보낸다 ← 흘린다(fail open)
+ failure_mode_allow_header_add: true 업스트림에 x-envoy-auth-failure-mode-allowed: true
Statistics accumulate under http.<HCM stat_prefix>.ext_authz.<필터 stat_prefix>. (the placeholders are the HCM stat_prefix and the filter's stat_prefix) as ok, denied, error and failure_mode_allowed. If you chose the let-through side, the moment failure_mode_allowed rises is the moment requests passed without authorization.
You can turn it off per path. If you put ExtAuthzPerRoute in the route's typed_per_filter_config and set disabled: true, that path does not ask the authorization server at all. It is used for requests that have no token, such as health checks and static files.
What it looks like in the field
"While we were deploying the authorization server, every service returned 403." That is the cost of the blocking side. The authorization server sits on the path of every request, so its availability becomes the availability of the service. That is why you attach the authorization server next to the proxy as a sidecar (the common placement of OPA), or run several and set a short timeout.
"A security check found a period when authorization was off." That is the cost of the let-through side. For the few minutes the authorization server was dead, every request passed, and if there was no marker header and no alert on statistics, that fact only shows up later when you dig through logs.
"I switched to gRPC and every request passes." The authorization cluster did not have HTTP/2 turned on, so every connection fails, but because it is a let-through configuration, requests quietly passed. error in the statistics will be rising by the number of requests.
Official documentation: External Authorization filter · ext_authz API · Authorization service API · Istio External Authorization
What you will do in the next lab
You write one authorization server yourself in an HTTP version and a gRPC version, and start two Envoys that ask each. You confirm from the server logs how a forged identity header is changed at the upstream and how the headers the two modes pass to the authorization server differ. You turn off authorization only for the health check path, and finally you kill both authorization servers and measure what responses the blocking side and the let-through side actually give.