TT Lab
Get started
Learn Learning paths Courses

Cloud Network Design

L4 and L7, and What Health Checks Do

Continue in TT Lab

In one line

A load balancer is not a device that splits traffic but a device that removes broken targets. Splitting is a side effect, and the value comes from health checks.

Why it was needed

As soon as you grow to two servers, two questions arise. Which one do you send users to, and what notices that one has died?

Putting two addresses in DNS looks cheapest, but it falls apart on the second question. DNS does not know whether a target is alive, so it keeps handing out the address of a dead server, and even if you fix the record, the change is not reflected as long as the TTL remains.

So you put in front of the traffic a device that keeps checking the state of the targets. Splitting the load is something that device does along the way, and the part that earns its keep is removing broken targets from the list.

L4 and L7

L4 (network) L7 (application)
What it looks at IP, port HTTP headers, path, host
What it can do Forward Path-based routing, host-based branching, redirects
TLS Pass through or terminate Usually terminate
Latency Very low Slightly more
Use TCP, gRPC, games, DB Web APIs, microservices

Sending /api/* to service A and /img/* to service B can be done only by L7. Conversely, if you need extreme throughput or the traffic is not HTTP, it is L4.

Health checks are the real job

The load balancer periodically sends requests to the targets, and when failures exceed a threshold, it removes that target from the list. Thanks to this behavior, users do not know when one instance dies.

What matters in the settings.

The liveness/readiness distinction covered in the earlier course applies the same way here — a load balancer's health check is closer to readiness. It asks "is it OK to receive traffic now?", not "does it need a restart?"

Connection draining

If you cut requests in progress when removing an instance, users get errors. Draining (deregistration delay) is the time to wait so that in-flight requests finish without sending new ones.

The order during deployment is this.

1. 대상을 목록에서 빼기 시작 (새 요청 중단)
2. 드레이닝 대기 — 진행 중 요청 완료
3. 애플리케이션에 SIGTERM
4. 정상 종료

Step 3 is the same SIGTERM covered in the container course. If the draining time is shorter than the application's graceful shutdown time, requests get cut off.

DNS and TTL

In front of the load balancer there is usually DNS. Here the TTL decides the speed of outage response.

So the principle is not to make DNS switching the main means of outage response. Removing the target inside the load balancer is far faster and more reliable.

Sticky sessions

This is a feature that sends the same user to the same server. It seems convenient, but it has a price.

If you put sessions in an external store (Redis and the like), stickiness is no longer needed. If possible, that is better.

The moments that collapse behind a load balancer

A load balancer is quiet normally and reveals problems when load is applied and when deploying. If you know what happens at those two moments, you can avoid most of it.

If the health check uses the same resources as the application, they collapse together. If you make /health query the database, the moment the DB slows down, all instances fall to unhealthy at the same time. Since all of them are removed, not just one, the service stops entirely. Separate the check for being alive (liveness) from the check for being ready to receive traffic (readiness), and make the former not touch dependencies.

Be loose on the criterion for removal and strict on the criterion for returning. If an instance is removed because of a temporary delay, the load piles onto the remaining instances and they are removed too. This is where a cascade starts. It is safer to set the failure threshold generously (3–5 times) and the recovery threshold short.

The reason for 502s during deployment is usually the order. If the application dies before the load balancer removes that instance from the list, requests that arrive in between have nowhere to go. Reverse the order: when you receive SIGTERM, first switch the health check to failing, wait for the time the load balancer needs to notice (check interval × threshold), then finish the requests in progress and shut down. In Kubernetes, the preStop hook and terminationGracePeriodSeconds create this time.

The connection draining time must be longer than the longest request. If you set it to 30 seconds but there is a report-generation request that takes 60 seconds, that request is cut off on every deployment.

Timeouts differ by layer, and the inner ones must be shorter. Set them shorter as you go inward, for example 60 seconds for the load balancer, 55 for the application and 50 for the database. If it is the other way around, the load balancer cuts first, and the application holds the connection while continuing to produce responses nobody is waiting for.

Idle timeouts break connection pools. If the load balancer cuts quiet connections at 60 seconds and the application's pool believes those connections are still alive, the next request fails with connection reset. The answer is to set the pool's idle time shorter than the load balancer's.

What it looks like in the field

Next course

Cost. Every decision so far has a price attached, and we cover how to read that price.