TT Lab
Get started
Learn Learning paths Courses

Cost and Architectural Decisions

Same Region, So Why the Charge

Continue in TT Lab

In one line

Incoming traffic is usually free, and charges attach to outgoing traffic and traffic that crosses a boundary. If you do not know this asymmetry, the bill cannot be explained.

Why this was needed

Multi-AZ and CDNs raise availability and performance, but they also change how many times data crosses a boundary. If you compare only compute prices, you miss the cost created by repeated communication between the app and the DB, or by origin egress, so you need to convert every major traffic path in the architecture into monthly transfer volume.

How it works

Direction and boundaries

Path Charge
Internet → cloud (ingress) Usually free
Cloud → internet (egress) Expensive
Within the same AZ Usually free
Between AZs Charged in both directions
Between regions Charged
VPC peering (same region) Similar to the cross-AZ charge

The one most often missed is the cross-AZ charge. You built it as multi-AZ for availability, but if the app server (AZ-a) and the DB (AZ-b) keep talking to each other, all of that traffic is billed. Each query is small, but at thousands per second it adds up over a month.

Designs that reduce this

1. Place things with AZs in mind Steer processing so it stays within the same AZ. On Kubernetes, topology-aware routing (a setting that prefers endpoints in the same zone) solves much of this. However, you are trading it against availability, so check the behavior when one zone dies.

2. Use endpoints This is the one covered in the earlier course. An S3 gateway endpoint has no charge and bypasses the NAT.

3. Put a CDN in front For services with large egress (images, video, downloads), CDN cache hits translate directly into cost savings. The CDN's egress unit price is lower than the origin's, and the request does not go to the origin in the first place.

4. Compress Sending API responses with gzip greatly reduces the transfer volume. But do not compress streaming responses — chunks pool in the buffer, and streaming stops being streaming.

5. Take care of the odd ones Check that you are not sending logs or metrics to another region, and that backup replication does not run more often than necessary.

Common misconceptions

The habit of calculating

Putting in numbers gives you a feel for it.

API 응답 평균 20KB, 초당 500 요청, 한 달 30일
= 20KB × 500 × 86,400 × 30
= 25,920 GB ≈ 26TB/월

이그레스 단가를 GB 당 100원으로 가정하면 약 260만 원/월
gzip 으로 5KB 까지 줄이면 약 65만 원/월 — 월 195만 원 절감

If you can do this kind of calculation in a design meeting, the debate gets shorter. Because it becomes not "I think compressing would be good" but "1.95 million won a month."

What you see in the field

Make the charges visible

What makes transfer charges frightening is not the amount but that you cannot see where they came from. Compute has a name on each instance, so you can tell who is using it, but transfer charges show up on the bill lumped into a single line such as "inter-region transfer." So to reduce them, you first have to make them visible.

Turn on flow logs. Once you have a record of how much went from which address to which address, you can actually point to the paths with large charges. The records themselves cost money, but the savings are usually far larger, and after you have found the problem you can lower the sampling rate.

Attach tags and look at it in those units. Viewed by team or by service, an odd communication pattern of one service that could not be seen in the overall bill comes to light.

Manage by transfer volume, not by charge. Unit prices change and discounts get applied, but "how many terabytes a month flow through this path" comes from the design and is stable. When discussing design, it is better to talk in this number.

And you need a mechanism to catch unexpected spikes. Billing alerts usually arrive too late. By the time the alert says you have passed half the monthly budget, several days' worth has already gone out. It is much faster to treat transfer volume itself as a metric and alert when it reaches some multiple of normal. Be especially careful about incidents where repeated calls loop endlessly. If a loop forms in which two services call each other, or a retry triggers another retry, transfer volume can reach dozens of times normal within hours. Such incidents show no functional symptoms, so they are often discovered only through the bill. That is why drawing the call relationships between services and checking that no loops form is worth the same from both the cost and the reliability points of view.

What to look at next

Now we build the procedure for actually finding waste.