Having Two Servers and Surviving a Failure Are Different Things
In one line
The first purpose of load balancing is not to split requests in half but to keep the failure of one server from spreading into failure of the whole service. So what you choose is not the algorithm but when to notice a failure and when to put the server back in.
Why this was needed
Starting two backends and putting them in an nginx upstream does not complete high availability. If you don't decide when to exclude a dead server, whether to reuse connections, and where to keep sessions, users will see success and failure alternate on every request.
The most common incident starts like this. One backend returns 500 for 30 seconds during a deployment. nginx doesn't know, so it keeps sending half the traffic there. In the user's eyes it becomes a state of "refreshing sometimes works and sometimes doesn't". Monitoring looks at averages, so an error rate of 50% is diluted to 25% and the alert is late too.
What do you choose the algorithm by
| Method | Distribution criterion | Fits where | Watch out for |
|---|---|---|---|
round_robin (default) |
Request order | APIs with uniform processing time | With slow requests mixed in, the deviation grows |
least_conn |
Number of connections in progress | Where response times are uneven | Distorted with SSE and WebSocket, where connections stay open long |
ip_hash |
Hash of the client IP | Legacy systems that keep sessions in server memory | When you add or remove a server, the assignments change wholesale |
hash $key consistent |
Consistent hash of an arbitrary key | Where the cache hit rate matters | If you design the key badly, traffic skews to one side |
Use weight only when server performance differs. If you give different weights to identical hardware,
the capacity calculation goes wrong and one side collapses first in an outage.
Noticing failure is passive
max_fails and fail_timeout are not an active health check. They work by observing
real requests failing and temporarily excluding the server. Three things follow from this.
- Without traffic, a failure is not detected either. A server that died at dawn is excluded only after the first request of the morning is sacrificed.
- The judgment is separate for each worker process. With
worker_processes 4, each worker has its own counter. Withmax_fails=3, in the worst case 12 failed requests go out. - After
fail_timeoutpasses, it puts the server back without any check. If the server hasn't recovered yet, it fails again, and this cycle continues.
What counts as a failure is decided by proxy_next_upstream. The default is error timeout,
so a 500 emitted by the backend is not counted as a failure. This is why requests keep going to an application
that has died and returns 500. You must specify http_500 http_502 http_503 for them to be counted. But unless you
also put non_idempotent with it, a POST can be executed twice, so always limit the retry targets to idempotent requests.
Upstream Keep-Alive needs three conditions to line up together
Reusing connections removes the cost of the TCP handshake and TLS negotiation. But all three must be done for it to turn on. If even one is missing, it silently opens a new connection every time.
upstream app {
server 127.0.0.1:8080;
server 127.0.0.1:8082;
keepalive 32; # 1) 워커당 유지할 유휴 연결 수
}
location / {
proxy_pass http://app;
proxy_http_version 1.1; # 2) 기본값 1.0 은 keep-alive 를 못 쓴다
proxy_set_header Connection "";# 3) 클라이언트가 보낸 Connection 헤더를 지운다
}
The keepalive value is not "concurrent throughput" but "the number of connections to leave idle".
The value multiplied by the number of workers must not exceed the backend's maximum concurrent connection setting. If Tomcat's
maxThreads is 200 and 4 nginx workers each keep 100, then 400 connections
pile up waiting for threads.
Session stickiness is the last resort
If you pin with ip_hash, the moment you add or remove a server the hash space changes, so a good number of
users are reassigned to other servers, and since their sessions are not in that server's memory, they are all
logged out at once. This happens every time you deploy. On top of that, users behind NAT,
such as in companies and schools, share an IP and get concentrated on one server.
The priority is this order.
- Take the session out of the server — if you keep it in Redis or a DB, any server can receive it.
- Put the state in a token — with a signed JWT the server doesn't have to remember anything.
- If that still doesn't work, pin it — even then, cookie-based sticky is better than
ip_hash.
What you see in the field
If failures show up only occasionally, first leave $upstream_addr and $upstream_status in the access log
and split the results by backend.
log_format lb '$remote_addr [$time_local] "$request" $status '
'$upstream_addr $upstream_status '
'$upstream_response_time $request_time';
How you read it matters. If $upstream_addr shows two addresses separated by a comma,
a retry happened (127.0.0.1:8080, 127.0.0.1:8082). If 5xx occurs only at a particular address,
look at deployment and configuration differences before the distribution algorithm. Conversely, if all addresses are slow
at the same time, investigate the common DB or downstream service.
The difference between $upstream_response_time and $request_time is also used in diagnosis. If the two are similar,
the backend is slow, and if only $request_time is large, the client-side network or the response transfer
is slow.
What you will do in the next lab
You start two backends yourself and compare round robin, weights, session stickiness, failure exclusion and connection reuse.
You leave $upstream_addr in the log and count which way requests went, and deliberately kill one
to confirm after how many failures max_fails actually takes effect.
At the end you organize, as operating criteria, when to use and when to avoid each algorithm.