Why the NAT Gateway Tops the Bill
In one line
There are three ways private resources go out, and each has a different price and security. If you use NAT without thinking, it rises to the top of the bill.
Why it was needed
If you create a private subnet, the instances in it have no way out. But they still need to receive package updates, upload logs and pull container images.
The easiest answer is to place a NAT gateway, so that is usually what people do, and after that nobody looks at it again. All outbound traffic from the private subnet funnels into that one place, and the charge grows with the amount of data that passes through.
A few months later, when you break the bill down by item, NAT is at number 1. In most cases that traffic was not unnecessary; the wrong path had simply been chosen, so changing only the path makes that amount disappear. So you divide the ways out into three and decide what to use when.
The three ways
| Method | What it is | Cost | Security |
|---|---|---|---|
| Internet gateway (IGW) | A two-way passage for a public subnet | The gateway itself is free | Exposes public IPs |
| NAT gateway | One-way, private → internet | Hourly + data throughput | Safe. Nothing can come in from outside |
| VPC endpoint | Goes straight to cloud services | Gateway type free / interface type paid | Does not pass through the internet |
How the NAT charge grows
A NAT gateway is billed in two ways.
1. 시간당 요금 — 켜 두기만 해도 나간다
2. 처리 데이터 요금 — GB 당. 나가는 것도 들어오는 것도
If you place one per AZ (which you should for availability), it is multiplied by that. With 3 AZs, you have 3 NATs.
The problem is number 2. When an instance in a private subnet reads a large volume from S3, all that traffic passes through the NAT — even though it is S3 in the same region. Log ingestion, backups and container image pulls all flow through here.
What a VPC endpoint solves
An endpoint creates a dedicated passage to a cloud service inside the VPC. Traffic does not go through the internet, and does not pass through NAT either.
- Gateway type (S3, DynamoDB) — a route is added to the route table. There is no charge. There is no reason not to use it.
- Interface type (most others) — an ENI is created in the subnet. There is an hourly plus data charge, but it is usually cheaper than NAT.
It is common for the NAT charge to drop to less than half just by creating an S3 gateway endpoint. It is one of the first things to check in cloud cost optimization.
The security benefit is also large — with an endpoint policy, you can restrict access to only buckets in our account. It becomes a means of narrowing data exfiltration paths.
What to use instead of a bastion
The setup of placing a bastion (jump server) to connect to private instances was used for a long time, but today the alternatives are better.
- Session manager family — the agent makes the connection outward, so you don't have to open a single inbound port. A record of access is kept too.
- VPN / zero-trust access control — per-user authentication and auditing.
A bastion itself is something you have to manage (patches, key management, logs), and port 22 has to be open. If you can get rid of it, it is better to do so.
Connecting to on-premises
| Method | Characteristics |
|---|---|
| Site-to-site VPN | An encrypted tunnel over the internet. Quick to set up. Bandwidth and latency depend on the internet |
| Dedicated line | Stable bandwidth and latency. Takes several weeks to several months to build, expensive |
| Transit gateway | Connects many VPCs and on-premises in a star shape. Management becomes simple |
Once you go beyond three or four VPCs, linking them with peering like a mesh soon hits a limit (N VPCs means N(N-1)/2 connections). That is the time to introduce a transit gateway.
What changes when you control the way out
Most people block the way in well. The reason the way out is left open is "going from inside to outside is safe," but every step after a compromise is outbound communication. Downloading tools, receiving commands, sending data.
Go in the order of blocking by default and opening only what is needed. But blocking all at once stops services, so first count where traffic actually goes out to. The place to start is counting the allow records in the flow logs, not the deny records.
An allow list lasts longer if you manage it by name, not by IP. The addresses of external APIs change often. If you write them by IP, one day it silently breaks, and the rule you opened broadly as a stopgap stays forever. Put in a name-based policy or a proxy and manage the list there.
Making traffic go through a proxy brings two things together. A record of where it goes out to is kept, and you can block places not on the list. In exchange, the proxy becomes a single point of failure and you have to decide how to handle TLS. To see the content, you have to swap the certificate in the middle, and then that proxy becomes the place that sees all communication in plaintext. Most organizations stop at the line of looking only at the destination name (SNI) and not at the content.
The path out to the metadata service is the most special. In the cloud, 169.254.169.254 gives out an instance's credentials. If an application has an SSRF vulnerability, an attacker can make it call this address and take the permissions wholesale. The sure way to block this path is not the firewall but the settings of the metadata service itself (enforcing IMDSv2, limiting the hop count).
DNS is also a way out. Even if you block every port, if DNS queries go out, data can be sent out bit by bit through them. The practical response is to make hosts use only the internal resolver, keep that resolver's query logs, and watch for a query volume different from usual.
What it looks like in the field
- NAT was number 1 on the bill → there was no S3 endpoint.
- Container image pull traffic went through NAT → a registry endpoint or a cache is needed.
- 6 VPCs were all linked with peering → 15 connections. Cleaned up with a transit gateway.
What to look at next
Name resolution and load balancing — the last layer that decides where traffic actually goes.