TT Lab
Get started
Learn Learning paths Courses

Cloud Network Design

You Opened Outbound, So Why Does It Fail

Continue in TT Lab

In one line

A security group remembers state (stateful), and a NACL does not (stateless). All the other differences come from this one difference.

Why it was needed

If you open only inbound 80 in a security group, the replies go out on their own. This is because it remembers that the request came in and automatically allows the reply.

A NACL is different. It does not remember what came in, so you must allow the reply going out separately. And since replies go out to an ephemeral port (1024–65535), you must open that range outbound.

If you don't know this, you stare for hours at the phenomenon "I opened inbound but no reply comes."

Side by side

Security group Network ACL
Applies to An instance (ENI) The whole subnet
State Stored — replies are allowed automatically Not stored — each direction is needed separately
Rules Allow only Allow + deny
Evaluation All rules are examined and one allow lets it through In number order, the first matching rule applies
Default All inbound blocked, all outbound allowed By default everything allowed
Count Several can apply per instance One per subnet

A trap created by the NACL's ordering rule

번호  타입      동작
100   전체      ALLOW
200   TCP 22    DENY      ← 절대 적용되지 않는다

Because it was already allowed and finished at number 100, number 200 is never looked at. It is not "deny takes priority" as with a security group. It goes from the smallest number, and ends at the first matching rule.

That is why NACL rule numbers are assigned with gaps, like 100, 200, 300 — to leave room to insert rules in between later.

A powerful feature of security groups — group references

In the source of a security group rule, you can specify another security group.

DB 보안그룹 인바운드:
  5432 ←  sg-app  (앱 서버 보안그룹)

What matters is that no IP was used. No matter how many app servers there are, whether they grow and shrink with autoscaling, or whether their IPs change, you never have to edit the rule. It is writing rules by role, and it is the most practical pattern in the cloud.

How to use them together

In practice, it usually settles like this.

If you try to do fine-grained control with a NACL, it soon gets tangled because of ephemeral ports and the ordering rule.

Five common mistakes

  1. Opening port 22 (SSH) to 0.0.0.0/0 — scanners find it within minutes.
  2. Leaving all outbound open — it becomes a data exfiltration path after a compromise.
  3. Not opening ephemeral-port outbound in a NACL, so replies are blocked.
  4. Numbering NACL rules 1, 2, 3, leaving no room to insert later.
  5. Not writing the security group description, so six months later nobody knows why.

Number 5 hurts more than you would expect. You can't tell whether it is safe to delete, so the rules keep piling up.

What it looks like in the field

How to keep rules from multiplying

Earlier we gave the case of security groups growing to 40. This is not a matter of laziness but a structural problem of never creating grounds for deletion. If you decide a few rules at the start, you never reach this state.

Enforce names and descriptions as a format. If the name shows which role of which service it is, and each rule's description says why it was opened and until when it is needed, you can decide six months later whether to delete it. It is better to block creation in a check when the description is empty.

Write by role, not by address. With group references, the number of rules becomes independent of the number of resources. Conversely, once you start writing by IP, the rules grow every time servers increase, and the rules remain even when the servers disappear. If someone else comes to use that address one day, unintended access gets opened — that is also a danger of IP rules.

Write an expiry on what you open temporarily. There are always times when you open something briefly for an investigation. If there is a rule that you write a date in the description at that time, you can later mechanically pull out the rules whose expiry has passed.

Periodically check whether they are actually used. If you turn on flow logs, you can tell which rule real traffic actually went through. An allow that has not been used even once for several months is a candidate for deletion.

And not editing rules by hand is the premise of all of this. A rule opened in a hurry in the console is not in code, so it disappears at the next deployment, or conversely remains with the code and reality out of sync. Either way, it leads to a state where nobody can answer "why is this rule here?"

What to look at next

The three ways private resources go out, and the price of each.