Network Fundamentals — Hands-on in a Linux VM
Block and Rewrite — netfilter Hooks, Conntrack, NAT
In one line
An nftables rule attaches to one of the five hooks a packet passes through (prerouting, input, forward, output, postrouting). Because connection tracking (conntrack) recognizes "packets of an already open connection", you allow replies with one line, and NAT rewrites and restores addresses on top of the same conntrack.
Why this exists
A firewall is not a list of rules but defaults and order. If the policy is drop, whatever you forgot is blocked, and if it is accept, whatever you forgot is open. And rules have direction — opening an incoming SYN to a server does not automatically open the reply the server sends out. In the old days that is why you wrote rules twice, and that is why connection tracking came about. The single line ct state established,related accept means "already open connections pass", so you only need to pick out the SYN that opens a connection.
Private addresses are not routed on the internet. Yet homes, offices, clouds, and Kubernetes Pods all use private addresses. For them to talk to the outside, the address has to be rewritten at the boundary and restored when the reply comes. To restore it, you have to remember "who this reply originally belongs to", and that memory is conntrack. This is why all connections drop when a NAT device restarts.
How it works
A hook is a position. A packet addressed to the host passes prerouting → input, a packet the host sends passes output → postrouting, and when it forwards someone else's packet it passes prerouting → forward → postrouting. Rules that lock down a server go in input, and a router's rules go in forward. If you lock down the router's own input, you block even your management access.
You create a chain like type filter hook input priority 0; policy drop. Rules are read from the top, and the verdict of the first match (accept/drop) ends the evaluation. Adding counter counts how many matched, and adding log records them in the kernel log (inside a namespace you need nf_log_all_netns=1).
Masquerade goes in postrouting — because that is the point where routing is finished and the outgoing interface is known. ip saddr 10.30.0.0/24 oifname "enp1s0" masquerade rewrites the source to the uplink address when that range goes out through the uplink. DNAT goes in prerouting — the destination must be changed before routing so that the route is found for the changed destination. What Docker's -p 8080:80 creates is this rule. A NAT chain is traversed only by the first packet of a connection, and conntrack handles the rest.
What it looks like in the field
Locking yourself out. You logged in over SSH and turned on policy drop. The session freezes. You did not put in the established allowance first, so the replies to your own session died. Order is safety — allowances first, policy afterward.
NAT works but nothing goes out. You added a masquerade rule, yet the private network cannot reach the internet. The forward chain's policy is drop. NAT only rewrites addresses and does not permit passage — the two chains do different jobs. The case where ip_forward is 0 gives the same symptom.
Discipline for handling firewall rules
What separates incidents is not the rules themselves but how you change them. The self-lockout we saw earlier can be almost eliminated by one procedure.
Set up the means of reverting first. When you change rules remotely, start by scheduling an automatic return to the old rules after a few minutes. If it goes well, you cancel it, and if you get locked out, you wait and log in again. This one habit eliminates "I couldn't get into the server, so I sent someone".
Manage rules as a file and apply them as a whole. If you put them in one line at a time by hand, nobody knows what the current state is, and it disappears on reboot. If you keep a file holding the complete rule set and swap it in atomically, the running rules and the file are always the same.
Turn on logging before you block. When you are not sure what a new rule will block, do not block; just attach counter and log and observe for a few days. If nothing matches, then change the verdict to drop. It is the same order as warn → enforce in Pod Security Admission, which we saw earlier.
And you must always be aware of where the firewall is. On modern systems a packet passes through several layers of filters: a cloud security group, the host's nftables, rules created by the container runtime, and a service mesh's policy. If you fix only one layer without knowing which layer blocked it, nothing changes, and you end up at "let's just open everything". As we saw earlier, looking at whether it is refused or a timeout, and whether each layer's counters go up, narrows down which layer it is.
Finally, leave comments on rules. nftables lets you attach a comment to a rule. A rule that does not say why it was opened is one nobody can delete, and among rules that pile up that way it gets harder and harder to find what is really needed.
What you will do in the next lab
You create an input chain for a server inside a namespace, start from default drop, and open lo, established, icmp, 80, and a source-restricted 2222 one at a time, counting what gets dropped with counter. Then you use the VM as a router to send a private namespace out to the internet with masquerade, capture the source being rewritten on both interfaces, and pass a port to the inner server with DNAT.