Reading ss Output, and the Bind Address Trap
In one line
A single ss -ltnp line shows both "is it listening" and "who is it listening for". Missing the latter is half of remote-access outages.
Why this exists
Let us look at one of the most common deployment incidents.
LISTEN 0 4096 127.0.0.1:8080 0.0.0.0:* users:(("api",pid=1,fd=7))
This process is listening only on loopback. From the same host, curl localhost:8080 succeeds. But a connection coming from another host or to the Pod IP is rejected by the kernel with an RST. That is, connection refused occurs.
This is where the diagnosis splits. If it is refused to the Pod IP but succeeds to localhost inside the Pod, the cause is not the network policy but the binding address. Conversely, if it times out even to the Pod IP, then you start looking at NetworkPolicy and the CNI. This single fork halves the scope of the investigation.
Each framework has its own spot where mistakes are easy to make.
| Framework | Safe notation | Notation that causes an incident |
|---|---|---|
| Express | app.listen(8080) |
app.listen(8080, '127.0.0.1') |
| Go net/http | ":8080" |
"localhost:8080" |
| Gunicorn | --bind 0.0.0.0:8000 |
--bind 127.0.0.1:8000 |
| Rails | -b 0.0.0.0 |
Default (local only) |
| PostgreSQL | listen_addresses = '*' |
Default localhost |
How it works
You can count the options of ss on one hand.
ss -ltnp # 리슨 중인 TCP + 프로세스
ss -tulpn # TCP/UDP 리슨 전체
ss -s # 소켓 총계 (누수 판별)
ss -tan state established # 확립된 연결만
ss -tanp state syn-sent # 나갔는데 응답 없는 연결
ss -tnio # rto/rtt/cwnd 같은 타이머까지
Meaning of the options: -t TCP, -u UDP, -l listening only, -a all, -n no name resolution, -p process.
Recv-Q and Send-Q change meaning depending on the state. This comes up often on exams.
| Socket state | Recv-Q | Send-Q |
|---|---|---|
| LISTEN | Length of the completed queue waiting to be accepted | The maximum backlog |
| ESTABLISHED | Received data not yet read | Sent data not yet acknowledged |
If the Recv-Q of a LISTEN socket is stuck to the Send-Q, it means the accept queue is overflowing. If nstat -az | grep -E 'ListenOverflows|ListenDrops' is increasing then, the cause is application throughput, not the network. However much you dig through the firewall, it will not turn up.
Let us also organize the operational judgment by TCP state.
- Too many
CLOSE_WAIT→ the application is not closing sockets (an fd leak) - Too many
SYN_RECV→ suspect a SYN flood - Too many
TIME_WAIT→ normal 2MSL waiting. Respond only when port exhaustion really occurs
It is also good to know how to read /proc/net/tcp directly. It becomes a last resort in a minimal container with no tools. The local address is little-endian hexadecimal and the port is also hexadecimal (8080 = 1F90), and 0A in the state column is LISTEN.
What it looks like in the field
The misdiagnosis of blaming the firewall for refused. Production firewalls almost always have a DROP policy, so a blocked port becomes a timeout. Unless a REJECT rule that returns an RST was used on purpose, refused means "it got there and nobody was listening".
When Kubernetes Endpoints are empty. If you connect to the service IP, you get refused immediately. If kubectl get endpoints <svc> shows <none>, the selector does not match or the readinessProbe is failing. It is a label problem, not a network problem.
Narrowing it down with one ss line
When you look at network problems, using ss instead of netstat lets you pull out only what you need
quickly. If you fix the forms you use often, investigation time drops greatly.
ss -tlnp # 듣고 있는 TCP 와 그 프로세스
ss -tanp state established '( dport = :5432 or sport = :5432 )'
ss -s # 상태별 요약
ss -tin # rtt, cwnd, 재전송 — 성능을 볼 때
The meaning of Recv-Q/Send-Q differs on a listening socket. On a connected socket they are
unread data and data that could not be sent, but on a listening socket they are the current accept queue
length and its upper limit. If the two numbers are stuck together, the application is failing to take connections,
and the kernel then discards completed connections.
nstat -az | grep -E 'ListenDrops|ListenOverflows|TCPBacklogDrop'
Look at where it is bound. 127.0.0.1:8080 cannot be entered from outside. The most common case is binding to loopback inside a container
and saying "the port won't open". It must show 0.0.0.0 or * to be reachable from outside.
Which state a connection stopped in separates the causes.
| State | Meaning |
|---|---|
SYN-SENT piling up |
The peer does not respond (a firewall quietly drops) |
SYN-RECV piling up |
A problem with our side's accept queue or SYN queue |
CLOSE-WAIT piling up |
Our application is not calling close() |
FIN-WAIT-2 piling up |
The peer is not closing |
Many TIME-WAIT |
Normal. It disappears after 60 seconds |
CLOSE-WAIT is especially important because it is the only state that points to a bug in our code, not the kernel.
It does not disappear with time.
Distinguish refusal from no response. Connection refused means it reached the peer host and
that port is not open, while a timeout means the packet was dropped somewhere. The former is a service problem, the latter a path or firewall problem.
What you will do in the next lab
You start two servers, one bound to 0.0.0.0 and one to 127.0.0.1. Then you knock on both ports with your own IP and build for yourself which one fails and why. You also read /proc/net/tcp by hand.