TT Lab
Get started
Learn Learning paths Courses

Tomcat & nginx Operations

Upstream Load Balancing and Failure Isolation

Continue in TT Lab

Goal

You balance load across several backends with an nginx upstream, set weights, session stickiness, failure isolation and Keep-Alive, and become able to choose a distribution algorithm to suit the situation.

Why it matters

The reason SI sites run two or more WASs is not performance but to keep the service up even if one dies. But open-source nginx has no active health check, so if you don't understand the meaning of max_fails/fail_timeout, you create a situation where "requests keep going to a dead server". Also, upstream Keep-Alive works only when all three settings are present, and there are really many systems that leave out proxy_set_header Connection "" and so open a new connection on every request. Then TIME_WAIT sockets pile up and at some point ports are exhausted.

Steps

  1. Preparation: deploy labhub.war to Tomcat (8080), and start /opt/lab/samples/labhub-boot.jar on port 8082. nginx uses /root/ng as its prefix.
  2. In /root/ng/nginx.conf, create an upstream app block and put 127.0.0.1:8080 and 127.0.0.1:8082 in it, and make / on port 8088 be proxied to this group. The log_format must include $upstream_addr, and send the access log to /root/ng/logs/access.log. http://127.0.0.1:8088/version must return 200.
  3. After sending 20 requests, copy the access log to /root/ng/rr.log. Two or more distinct upstream addresses must be recorded.
  4. Give 127.0.0.1:8080 a weight=3 and reload, then empty the access log, send 40 requests, and copy it to /root/ng/weight.log. The number of requests that went to 8080 must be between 28 and 32.
  5. Change the distribution method to ip_hash and reload, empty the log, send 20 requests, and copy it to /root/ng/iphash.log. Only one kind of upstream address must be recorded.
  6. Remove ip_hash (it must not remain in the configuration file), and set max_fails=2 fail_timeout=10s on both servers. Then stop the 8082 process and send 10 requests. All 10 must return 200.
  7. Set up upstream Keep-Alive. In the upstream block, keepalive 32;, and proxy_http_version 1.1; and proxy_set_header Connection ""; in the location. All three must be present.
  8. In the upstream block, add least_conn; and make it pass the syntax check.
  9. Write /root/ng/lb.md. For each of the four methods 라운드로빈 (round robin), 가중치 (weight), ip_hash and least_conn, write when to use it and when not to use it. The word 세션 (session) must appear, and the total must be 400 characters or more.

Notes

Configure the upstream group

In /root/ng/nginx.conf, create an upstream app block and put 127.0.0.1:8080 and 127.0.0.1:8082 in it, and make / on port 8088 be proxied to this group. The log_format must include $upstream_addr, and send the access log to /root/ng/logs/access.log. http://127.0.0.1:8088/version must return 200.

Put the upstream block in the http context. In proxy_pass you use that name like a URL. Don't forget to start both backends.

Confirm round robin distribution

After sending 20 requests, copy the access log to /root/ng/rr.log. Two or more distinct upstream addresses must be recorded.

You can tell which backend a request went to from the upstream address variable in the access log. Send enough requests and count the distinct addresses in the log.

Weighted distribution

Give 127.0.0.1:8080 a weight=3 and reload, then empty the access log, send 40 requests, and copy it to /root/ng/weight.log. The number of requests that went to 8080 must be between 28 and 32.

nginx's weighted round robin is deterministic, not random. The weight ratio becomes the distribution ratio as it is.

Session stickiness

Change the distribution method to ip_hash and reload, empty the log, send 20 requests, and copy it to /root/ng/iphash.log. Only one kind of upstream address must be recorded.

It is a method that sends the same client to the same backend. It is often used in legacy systems that keep sessions in WAS memory. Also remember the drawback that redistribution happens when servers are added or removed.

Automatically exclude a failed server

Remove ip_hash (it must not remain in the configuration file), and set max_fails=2 fail_timeout=10s on both servers. Then stop the 8082 process and send 10 requests. All 10 must return 200.

Open-source nginx has no active health check, only a passive method based on failure counts. Actually bring one backend down and confirm whether the service stays up.

Upstream Keep-Alive

Set up upstream Keep-Alive. In the upstream block, keepalive 32;, and proxy_http_version 1.1; and proxy_set_header Connection ""; in the location. All three must be present.

It works only when all three are set together. If even one is missing, a new connection is created on every request. How to handle the Connection header in particular is the trap.

Least connections method

In the upstream block, add least_conn; and make it pass the syntax check.

Think about why it is better than round robin for a service with a large deviation in request processing time.

Summarize the criteria for choosing a distribution algorithm

Write /root/ng/lb.md. For each of the four methods 라운드로빈 (round robin), 가중치 (weight), ip_hash and least_conn, write when to use it and when not to use it. The word 세션 (session) must appear, and the total must be 400 characters or more.

A document that lists only the advantages of each method is not a document. You must also write when not to use it so that it can be used for judgment later.