TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

Same Upstreams, Change Only the Policy

Continue in TT Lab

Goal

Change the load-balancing method one at a time, count the actual distribution, and confirm for yourself two properties of hash-based methods (the same key goes to the same place, and removing one server moves only that server's share).

Why it matters

Load balancing is one line of configuration, but that one line decides the cache hit rate, session persistence and the size of the shock when you remove a server. Everything sounds plausible when you read it in documentation, but when you count the numbers yourself, "random skews this much on a small sample" and "removing one server moves only this many" stay with you. In particular, the fact that a request whose key could not be extracted is routed at random is a regular cause of session affinity outages, and you cannot see it from the configuration alone.

Steps

  1. Start three upstreams as ok (8088, 8089 and 8090). In /root/envd-lb/lb-rr.yaml, put a STATIC cluster pool with lb_policy: ROUND_ROBIN and a / route, and start it with --concurrency 1 (admin 9951, listener 127.0.0.1:10051). Request 9 times and write four lines to /root/envd-lb/01-rr.txt: p8088=, p8089=, p8090= and total=.
  2. Copy /root/envd-lb/lb-rr.yaml to /root/envd-lb/lb-weight.yaml and give each endpoint a load_balancing_weight — 2 for 8088 and 1 for the other two. Start it again with that configuration, request 12 times, and write four lines to /root/envd-lb/02-weight.txt: p8088=, p8089=, p8090= and total=.
  3. Create /root/envd-lb/lb-random.yaml — a configuration with no weights and only lb_policy changed to RANDOM. Start it with that configuration, request 40 times, and write five lines to /root/envd-lb/03-random.txt: p8088=, p8089=, p8090=, total= and max_gap= (max_gap is the largest count received minus the smallest count received).
  4. Create /root/envd-lb/lb-ring.yaml — lb_policy is RING_HASH, ring_hash_lb_config.minimum_ring_size is 1024, and the route's hash_policy is the header x-user. After you start it with that configuration, request once each as nine users, u1 through u9, and write nine lines of 사용자 포트 (the placeholders are the user and the port) to /root/envd-lb/04-map.txt.
  5. Request four times as user u1 and four times as u2, and write three lines to /root/envd-lb/05-sticky.txt: u1_ports= (the four ports received, separated by spaces), u1_distinct= (the number of different ports) and u2_distinct=.
  6. Create /root/envd-lb/lb-ring2.yaml — the step 4 configuration with only the endpoint 8090 removed. After you start it again with that configuration, request again as the same nine users and write nine lines of 사용자 포트 (the placeholders are the user and the port) to /root/envd-lb/06-remap.txt. Then compare with the table from step 4 and write three lines to /root/envd-lb/06-moved.txt: moved= (the number of users whose place changed), stayed= and total=9.
  7. Create /root/envd-lb/lb-query.yaml — there are three endpoints again, and the hash_policy is the query parameter uid instead of the header. After you start it with that configuration, request three times with ?uid=u1 and three times with ?uid=u2, and write three lines to /root/envd-lb/07-key.txt: u1_distinct=, u2_distinct= and header_ignored= (the number of different ports when you request three times with only the header x-user: u1 attached and no query).
  8. In /root/envd-lb/08-report.md, write four lines — even_spread= (the three values from step 1, separated by commas), heavy_share= (the share received by the endpoint given weight 2 in step 2, as an integer percentage), sticky= (yes if the same user stuck to one place in step 5) and moved_users= (the value from step 6) — and below them write what you learned in at least four lines.

Notes

Splitting evenly is the default

Start three upstreams as ok (8088, 8089 and 8090). In /root/envd-lb/lb-rr.yaml, put a STATIC cluster pool with lb_policy: ROUND_ROBIN and a / route, and start it with --concurrency 1 (admin 9951, listener 127.0.0.1:10051). Request 9 times and write four lines to /root/envd-lb/01-rr.txt: p8088=, p8089=, p8090= and total=.

Round robin picks one at a time, going around the list. It is the default, and it is the easiest to predict when the servers are the same size and request costs are similar. Because each worker thread remembers its own turn separately, if you count with the default concurrency (the number of cores), 9 requests do not give 3, 3, 3. That is why this lab ties the workers down to one.

Put servers of different sizes in one group

Copy /root/envd-lb/lb-rr.yaml to /root/envd-lb/lb-weight.yaml and give each endpoint a load_balancing_weight — 2 for 8088 and 1 for the other two. Start it again with that configuration, request 12 times, and write four lines to /root/envd-lb/02-weight.txt: p8088=, p8089=, p8090= and total=.

When you add servers, they do not always come in the same size. If a newly bought machine is twice as fast, it is better to have it take twice as much, and what you use for that is the endpoint weight. Round robin goes around reflecting the weights, and a weight of 2 does not mean "one more time out of every two" but "twice the share of the total". The sum is 4, so with 12 requests it is 6, 3 and 3.

It looks even, but it is not

Create /root/envd-lb/lb-random.yaml — a configuration with no weights and only lb_policy changed to RANDOM. Start it with that configuration, request 40 times, and write five lines to /root/envd-lb/03-random.txt: p8088=, p8089=, p8090=, total= and max_gap= (max_gap is the largest count received minus the smallest count received).

Random is a method that remembers no state at all. So the result is the same however many workers there are, and there is nothing to recalculate when endpoints come and go. In exchange, with a small sample it skews noticeably — compare it with round robin giving exactly 3, 3, 3 in 9 requests. Where requests run to thousands per second, this difference disappears, so it is sometimes used as the default at large scale.

The same user always goes to the same server

Create /root/envd-lb/lb-ring.yaml — lb_policy is RING_HASH, ring_hash_lb_config.minimum_ring_size is 1024, and the route's hash_policy is the header x-user. After you start it with that configuration, request once each as nine users, u1 through u9, and write nine lines of 사용자 포트 (the placeholders are the user and the port) to /root/envd-lb/04-map.txt.

The hash-based method hashes the key extracted from the request to find a place on the ring, and from that place picks the nearest endpoint clockwise. The same key always goes to the same place, so the same user sticks to the same server — the local cache hit rate goes up, and the server may hold sessions. The key is chosen by hash_policy (header, cookie, query parameter, source IP). If minimum_ring_size is small, the ring is sparse and the distribution skews.

Send the same key four times and the place does not change

Request four times as user u1 and four times as u2, and write three lines to /root/envd-lb/05-sticky.txt: u1_ports= (the four ports received, separated by spaces), u1_distinct= (the number of different ports) and u2_distinct=.

Without this property, for a service with a local cache, the same user's requests go to a different server every time and the cache almost never hits. Conversely, relying on this property brings the risk that one particular user can single-handedly make one server heavy — which is why what you choose as the key matters. Count the number of different values with sort -u | wc -l.

If you remove one server, how many users change places

Create /root/envd-lb/lb-ring2.yaml — the step 4 configuration with only the endpoint 8090 removed. After you start it again with that configuration, request again as the same nine users and write nine lines of 사용자 포트 (the placeholders are the user and the port) to /root/envd-lb/06-remap.txt. Then compare with the table from step 4 and write three lines to /root/envd-lb/06-moved.txt: moved= (the number of users whose place changed), stayed= and total=9.

This is the real reason to use a hash ring. If you simply use "the hash value modulo the number of servers", the moment the number of servers changes, almost everyone changes places. With the ring method, only the keys that were attached to the vanished server move, and the rest stay as they are. Count for yourself how many people moved — it is convenient to put the two tables side by side with join or paste and compare.

Change the key from a header to a query string

Create /root/envd-lb/lb-query.yaml — there are three endpoints again, and the hash_policy is the query parameter uid instead of the header. After you start it with that configuration, request three times with ?uid=u1 and three times with ?uid=u2, and write three lines to /root/envd-lb/07-key.txt: u1_distinct=, u2_distinct= and header_ignored= (the number of different ports when you request three times with only the header x-user: u1 attached and no query).

What you extract the key from is what you treat as "the same thing". For a user session a cookie is natural, for a tenant a header, and for a cache key a query parameter. If the key cannot be extracted, there is no hash, so Envoy sends that request at random — that is why a request missing the key is not pinned. The last line is to confirm that.

Summarize it as a table of properties by method

In /root/envd-lb/08-report.md, write four lines — even_spread= (the three values from step 1, separated by commas), heavy_share= (the share received by the endpoint given weight 2 in step 2, as an integer percentage), sticky= (yes if the same user stuck to one place in step 5) and moved_users= (the value from step 6) — and below them write what you learned in at least four lines.

The purpose of making the table is to let you choose "which method to use when" yourself next time. Take the values from the files of the earlier steps, and in the explanation lines write what each method gives up and what it gains.