There Is Not One Limit but Four
In one line
When people see Too many open files, almost everyone runs ulimit -n 65536 and restarts. And they hit the same error again. That is because there are four limits, not one, and a value changed in the shell does not apply to a service that is already running.
Why this was needed
nginx: accept4() failed (24: Too many open files) or java.net.SocketException: Too many open files. Three misunderstandings overlap in this message.
- Thinking there is one limit. In fact there are four.
- Thinking a value changed in the shell applies to the service. It does not.
- Thinking raising the limit solves it. Usually it is a leak, so you only push the moment of failure back a few hours.
How it works
The containment relationship of the limits is like this.
fs.file-max 시스템 전체 합 (현대 리눅스에선 대개 병목이 아님)
fs.nr_open 한 프로세스의 RLIMIT_NOFILE 이 넘을 수 없는 커널 천장
RLIMIT_NOFILE hard 관리자가 정한 상한
RLIMIT_NOFILE soft 실제 적용값 (프로세스가 hard 까지 스스로 올릴 수 있다)
The errno separates the layers. EMFILE(24) is Too many open files and means the process limit, and ENFILE(23) is Too many open files in system and means system-wide. The two words in system are half of the diagnosis.
Why the shell's ulimit is useless. /etc/security/limits.conf is read by a PAM module, and PAM works only in login sessions (SSH, su, login). A service that systemd started at boot is not a login session, so it never reads this file at all. A service's limit is set with LimitNOFILE in the unit file. And an rlimit is applied at process creation time, so it takes effect not with a reload but with a restart.
Why systemd keeps the soft limit low. Since version 240, systemd keeps the hard limit large (524288) and the soft limit low (1024). It is for compatibility - select() cannot handle fd numbers of 1024 and above, and some old programs run a loop at startup that closes every fd from 0 up to the soft limit. If that limit is a million, startup takes several seconds. The design is "programs that need more should raise it themselves". The Go runtime actually raises the soft limit to the hard limit at startup, so you get the situation where only the Go service is fine on the same machine.
There is only one reliable way to check.
grep 'Max open files' /proc/<PID>/limits # soft / hard
ls /proc/<PID>/fd | wc -l # 실제 사용량
lsof -p PID | wc -l counts even the cwd, root, executable and mmap'ed libraries, and these are not included in RLIMIT_NOFILE, so the number comes out larger. It is not a value to compare against the limit.
What you see in the field
The typical shape of a socket leak. When you count by fd type, there are 60,000 sockets, but looking at the states with ss, 20,000 are CLOSE-WAIT, and by peer address they are all the same Redis server. CLOSE-WAIT is the state where the peer sent a FIN but our side did not call close(), and the kernel waits indefinitely, with no timeout. That means the application is not returning connections, and the scope of the code review is settled on the spot.
Conversely, TIME-WAIT is cleaned up by the kernel on its own and does not consume an fd. If you confuse the two, you end up touching the wrong kernel parameters.
How to tell a leak from load. Record the fd count and the ESTABLISHED count together at 1-minute intervals. If the connection count is flat but only the fd count increases monotonically, it is a leak. Raising the limit then only pushes the outage back a few hours.
A socket is an fd too. You miss it if you count only files, but sockets, pipes, epoll and inotify instances all consume fds. If you readlink the symlinks in /proc/PID/fd, the kinds show up, like socket:[12345], pipe:[67890] and anon_inode:[eventpoll].
Where the limit actually comes from
When you meet Too many open files, raising ulimit -n is the first reaction, and in half of the
cases it has no effect. Limits come from four places, and which one won can be
confirmed only on the running process.
cat /proc/<pid>/limits | grep -i 'open files'
This value is the truth. No matter how much you change ulimit -n in the shell, it does not apply to a
process that is already running.
| Place | What it applies to |
|---|---|
ulimit -n |
That shell and its descendants |
/etc/security/limits.conf |
Sessions that logged in through PAM |
LimitNOFILE in a systemd unit |
That service — it does not look at limits.conf |
| Container runtime settings | Inside the container |
Services started by systemd are hit the most. This is why editing limits.conf and restarting
changes nothing. You have to write it in the unit itself.
[Service]
LimitNOFILE=65535
There is a system-wide limit too. Even if you raise the per-process limit, if you hit fs.file-max
you are blocked there.
sysctl fs.file-max fs.nr_open
cat /proc/sys/fs/file-nr # 현재 사용, 미사용, 최대
Count what is eating the fds. The remedy differs depending on whether they are sockets or files.
ls -l /proc/<pid>/fd | awk '{print $NF}' | sed 's/:.*//' | sort | uniq -c | sort -rn | head
If most are socket:, connections are not being closed, and if thousands are the same file,
it is code that opens a file and doesn't close it. Raising the limit buys time; it is not a fix.
Don't put in an enormous value. If you leave LimitNOFILE=infinity, some programs
try to close everything from 0 up to that number and hang for minutes. Set it a little
larger than what is actually needed.
What you will do in the next lab
Starting by reading the real limit from /proc/1/limits, you open the same file 50 times and confirm the structure in which there are many fds but the inodes converge to one. You reproduce EMFILE yourself in a shell with a lowered limit, see with your own eyes that sockets are fds too, and then build a leak-tracking tool that picks the processes with the most fds.