L4 and L7, and What Health Checks Do
In one line
A load balancer is not a device that splits traffic but a device that removes broken targets. Splitting is a side effect, and the value comes from health checks.
Why it was needed
As soon as you grow to two servers, two questions arise. Which one do you send users to, and what notices that one has died?
Putting two addresses in DNS looks cheapest, but it falls apart on the second question. DNS does not know whether a target is alive, so it keeps handing out the address of a dead server, and even if you fix the record, the change is not reflected as long as the TTL remains.
So you put in front of the traffic a device that keeps checking the state of the targets. Splitting the load is something that device does along the way, and the part that earns its keep is removing broken targets from the list.
L4 and L7
| L4 (network) | L7 (application) | |
|---|---|---|
| What it looks at | IP, port | HTTP headers, path, host |
| What it can do | Forward | Path-based routing, host-based branching, redirects |
| TLS | Pass through or terminate | Usually terminate |
| Latency | Very low | Slightly more |
| Use | TCP, gRPC, games, DB | Web APIs, microservices |
Sending /api/* to service A and /img/* to service B can be done only by L7. Conversely, if you need extreme throughput or the traffic is not HTTP, it is L4.
Health checks are the real job
The load balancer periodically sends requests to the targets, and when failures exceed a threshold, it removes that target from the list. Thanks to this behavior, users do not know when one instance dies.
What matters in the settings.
- Path — a light dedicated endpoint such as
/healthz. It should not go through real business logic. If it is heavy, the health check itself becomes a load. - Interval and threshold — if short, it is sensitive and removes a target on a momentary delay; if long, you learn of an outage late. Usually you set an interval of 10–30 seconds and 2–3 consecutive failures.
- Healthy threshold — the criterion for putting it back after removing it. If you set this short, targets go in and out (flapping).
The liveness/readiness distinction covered in the earlier course applies the same way here — a load balancer's health check is closer to readiness. It asks "is it OK to receive traffic now?", not "does it need a restart?"
Connection draining
If you cut requests in progress when removing an instance, users get errors. Draining (deregistration delay) is the time to wait so that in-flight requests finish without sending new ones.
The order during deployment is this.
1. 대상을 목록에서 빼기 시작 (새 요청 중단)
2. 드레이닝 대기 — 진행 중 요청 완료
3. 애플리케이션에 SIGTERM
4. 정상 종료
Step 3 is the same SIGTERM covered in the container course. If the draining time is shorter than the application's graceful shutdown time, requests get cut off.
DNS and TTL
In front of the load balancer there is usually DNS. Here the TTL decides the speed of outage response.
- With a TTL of 300 seconds, even if you move traffic during an outage, it goes to the old address for up to 5 minutes.
- If you set it short (60 seconds), switching is fast but lookups increase.
- Some clients ignore the TTL and cache (some JVMs are well known for this).
So the principle is not to make DNS switching the main means of outage response. Removing the target inside the load balancer is far faster and more reliable.
Sticky sessions
This is a feature that sends the same user to the same server. It seems convenient, but it has a price.
- The load is not split evenly
- If that server dies, those users' sessions vanish entirely
- It does not fit well with autoscaling
If you put sessions in an external store (Redis and the like), stickiness is no longer needed. If possible, that is better.
The moments that collapse behind a load balancer
A load balancer is quiet normally and reveals problems when load is applied and when deploying. If you know what happens at those two moments, you can avoid most of it.
If the health check uses the same resources as the application, they collapse together. If you make /health query the database, the moment the DB slows down, all instances fall to unhealthy at the same time. Since all of them are removed, not just one, the service stops entirely. Separate the check for being alive (liveness) from the check for being ready to receive traffic (readiness), and make the former not touch dependencies.
Be loose on the criterion for removal and strict on the criterion for returning. If an instance is removed because of a temporary delay, the load piles onto the remaining instances and they are removed too. This is where a cascade starts. It is safer to set the failure threshold generously (3–5 times) and the recovery threshold short.
The reason for 502s during deployment is usually the order. If the application dies before the load balancer removes that instance from the list, requests that arrive in between have nowhere to go. Reverse the order: when you receive SIGTERM, first switch the health check to failing, wait for the time the load balancer needs to notice (check interval × threshold), then finish the requests in progress and shut down. In Kubernetes, the preStop hook and terminationGracePeriodSeconds create this time.
The connection draining time must be longer than the longest request. If you set it to 30 seconds but there is a report-generation request that takes 60 seconds, that request is cut off on every deployment.
Timeouts differ by layer, and the inner ones must be shorter. Set them shorter as you go inward, for example 60 seconds for the load balancer, 55 for the application and 50 for the database. If it is the other way around, the load balancer cuts first, and the application holds the connection while continuing to produce responses nobody is waiting for.
Idle timeouts break connection pools. If the load balancer cuts quiet connections at 60 seconds and the application's pool believes those connections are still alive, the next request fails with connection reset. The answer is to set the pool's idle time shorter than the load balancer's.
What it looks like in the field
- A little 5xx on every deployment → the draining time is short or there is no SIGTERM handling.
- The health check also checks the DB → when the DB slowed down, all targets were removed and the whole service stopped.
- Changed DNS during an outage but traffic did not move → TTL and client caches.
Next course
Cost. Every decision so far has a price attached, and we cover how to read that price.