Six Lines of Proxy Config, and What Happens Without Them
In one line
In proxy configuration, incidents usually come from six lines of header forwarding and a single slash in proxy_pass, and since the screen comes up fine even when both are missing, they are discovered late.
Why put a web server in front
Tomcat can speak HTTP too, so why put nginx in front? The real reason at SI sites is operational convenience and policy rather than performance.
- Keep the WAS from handling static files (saves WAS threads)
- Terminate SSL at one point (certificate replacement is finished on one server)
- Balance load across several WASs, and keep the service up even if one dies
- Route to different systems by URL (
/api/to the new one, the rest to the legacy one) - Apply access control, request size limits, compression and cache headers in one place
The last line matters especially. If you put a policy in the application, you have to deploy to change it, but if you put it in the proxy, a reload changes it. That is why operations organizations like proxies.
Six lines of proxy configuration
location / {
proxy_pass http://127.0.0.1:8080;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Forwarded-Host $host;
proxy_set_header X-Forwarded-Port $server_port;
proxy_connect_timeout 3s;
proxy_send_timeout 30s;
proxy_read_timeout 60s;
}
You have to know what happens when each line is missing; only then do you know it rather than have memorized it.
- Without
Host $host, the Host the backend receives is127.0.0.1:8080. When the application builds absolute URLs (redirects, email links, payment callbacks),http://127.0.0.1:8080/...goes out. If you discover this after go-live, payment callbacks don't come back. - Without
X-Forwarded-For, the client IP the application sees is always the proxy IP. Access history, fraud detection and IP-based access control all become meaningless. You can't answer "are you recording the visitor's IP?" in an audit.$proxy_add_x_forwarded_foris a variable that appends to the existing header, so with several proxy tiers it becomes a list. - Without
X-Forwarded-Proto— this is the biggest cause of redirect infinite loops. The proxy terminates TLS and passes HTTP to the backend. The backend thinks "this isn't HTTPS" and redirects tohttps://. It comes to the proxy again, is passed on as HTTP again, and repeats forever. The browser showsERR_TOO_MANY_REDIRECTS. The vast majority of loops come from the TLS termination point and the application's HTTPS enforcement logic not knowing about each other.
And one important security principle: these headers must be overwritten at the front-most proxy.
That is because a client can send X-Forwarded-For: 10.0.0.1 directly.
If you overwrite with proxy_set_header, the value the client sent is ignored (the variable that appends is the exception).
The slash in proxy_pass — the most common nginx bug
location /api/ {
proxy_pass http://backend; # 슬래시 없음 → /api/users/1 이 그대로 전달
}
location /api/ {
proxy_pass http://backend/; # 슬래시 있음 → /api 부분이 잘리고 /users/1 로 전달
}
If there is a URI part (including a slash), the part that matched the location is cut off.
If the backend expects the /api context and you add the slash, 404s pour in,
and conversely if the backend expects the root and you leave the slash out, it becomes /api/api/users.
Every year I see people spend half a day on this one character. When in doubt,
looking at $upstream_addr in the access log together with the request path in the backend log tells you immediately.
Request size — what a 413 really is
nginx's client_max_body_size defaults to 1MB.
In a system with file uploads, if you don't raise this value, uploading a 3MB attachment
gives 413 Request Entity Too Large. And since nginx returns this error,
nothing is left in the application log. Developers wander for days saying "there's no log at all on the server".
client_max_body_size 20m;
client_body_buffer_size 128k;
large_client_header_buffers 4 16k;
large_client_header_buffers is for when headers are large. When you attach SSO, cookies
grow, and if they exceed the default buffer, you get 400 Bad Request. The log is ambiguous for this too.
Compression and static files
gzip on;
gzip_comp_level 6;
gzip_min_length 1000;
gzip_types text/plain text/css application/json application/javascript text/xml;
gzip_vary on;
gzip_comp_levelaround 6 is the balance point of compression ratio against CPU. Even if you raise it to 9, the size shrinks by a few percent while CPU noticeably rises.gzip_min_length 1000— below 1KB there is no gain from compressing, only overhead.- If you leave out
gzip_vary on, an intermediate cache may serve the compressed copy to a client that doesn't support compression. - In
gzip_types,text/htmlis always included, so you don't need to write it.
It is better not to send static files through the proxy at all.
location /static/ {
alias /app/static/;
expires 7d;
access_log off;
}
Confusing root and alias is also a regular. For location /static/,
root /app; gives /app/static/파일, and alias /app/static/; gives /app/static/파일.
They look the same, but for location /s/ with alias /app/static/;, /s/a.js → /app/static/a.js.
Use alias when the location path and the actual directory name differ.
Blocking admin screens
Tomcat Manager or an actuator being open to the outside leads to real incidents. Blocking at the proxy is the most reliable.
location ~ ^/(manager|host-manager)/ { return 403; }
location /actuator/ {
allow 10.0.0.0/8;
deny all;
proxy_pass http://127.0.0.1:8080;
}
The lesson here: you can block with application settings too, but if you block at the proxy, you can change the policy without deploying the application. Even if a security audit finding comes down on a Friday afternoon, a reload finishes it.
What to leave in the access log
With the default combined format, you can't analyze incidents. At the very least, include these four.
log_format labhub '$remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'ua="$upstream_addr" us=$upstream_status '
'urt=$upstream_response_time rt=$request_time';
$upstream_addr— which backend it went to (essential with load balancing)$upstream_status— the status code the backend gave (distinguishes it from a 502 that nginx generated)$upstream_response_time— the time the backend took$request_time— the total time from the client's point of view
If $request_time is large and $upstream_response_time is small, the problem is not the backend but
the client's network or the response transfer. That one difference between these two values
can break the misconception "the backend is slow" many times over.
What you see in the field
The symptoms when the six header lines are missing look unrelated to each other. The client IP in the access log is always the single proxy IP (missing X-Real-IP and X-Forwarded-For), the redirect after login jumps to an internal address or to http (missing Host and X-Forwarded-Proto), and IP-based access control and audit logs become meaningless altogether.
What they have in common is that there is no problem in ordinary times. So it is revealed only weeks after go-live, in a security audit or incident analysis, and by then you can trace nothing with those logs.
The proxy_pass slash is quieter. If there is a slash at the end, the location path is stripped before forwarding, and if not, it is kept and forwarded; in the development environment the path is the root so the difference doesn't show, and then the moment a context path is attached in production it becomes a 404.