TT Lab
Get started
Learn Learning paths Courses

LFCS — Linux Foundation System Administrator

Processes, Signals, and Why /proc Is a Filesystem

Continue in TT Lab

In one line

The symptom "the process won't die" comes from four different causes. It is catching the signal and ignoring it, it is in a state that cannot be interrupted (D), the signal went to the wrong target, or it was dead to begin with and the parent has not reaped it (a zombie). The ways to tell them apart are all in /proc.

Why this exists

A deployment script sent kill and the process is still there. The hand that goes straight to kill -9 from here is the problem. SIGKILL gives the application no chance to clean up. Buffers of open files are not flushed, transactions are cut off in the middle, and lock files are left behind. Sending SIGKILL to a database or a queue consumer is the same as pulling the power cord.

The correct procedure is always two steps. Send SIGTERM, give a grace period, and if it is still alive then, SIGKILL. What container runtimes and Kubernetes do is exactly this, and the grace period is a configurable value.

How it works

The one-letter state is the starting point of diagnosis.

Code Meaning Practical implication
R Running or waiting to run Using CPU, or queued to get it
S Interruptible sleep Normal waiting. Can be woken by a signal
D Uninterruptible sleep Even SIGKILL does not take effect immediately. Usually disk or NFS
Z Zombie Already dead and not yet reaped by the parent
T Stopped Received SIGSTOP or SIGTSTP

If you see D, a signal will not solve it. You have to look at what the kernel is waiting for in /proc/<PID>/wchan or /proc/<PID>/stack.

SIGKILL and SIGSTOP cannot be caught, blocked, or ignored. For every other signal, the application can change how it is handled. That is why programs on which kill -TERM has no effect exist, and that may be design, not a bug. Which signals a process is catching is written out directly in the SigCgt, SigBlk, and SigIgn bitmasks of /proc/<PID>/status.

Why /proc is a filesystem. To show the kernel's internal data structures to user space, you either create a pile of new system calls or reuse an interface that already exists. Unix chose the second. So process information can be read with cat, open file descriptors appear as symbolic links under /proc/<PID>/fd/, and if you cp such a link as it is, even a deleted file is recovered. Resource limits are in /proc/<PID>/limits, so you can check the limits that process actually received, not your shell's ulimit.

nice and ulimit are both inherited. The nice value runs from -20 to 19, and the lower it is, the higher the priority. An ordinary user can only raise the value, not lower it. The soft limit of ulimit can be adjusted freely below the hard limit, but raising the hard limit is possible only for root. And because both values are inherited by child processes as they are, a limit you lowered once in a shell follows every process started under it. The answer to "why can only this daemon not open files" is often here.

What it looks like in the field

Three of the author's collected systemd problem cases come up especially often. First, if you set only Restart=on-failure and put no restart limit, a service that crashes because of a configuration error goes into infinite restarts, burning CPU and filling the disk with logs. StartLimitIntervalSec and StartLimitBurst, which set the limit, are directives of the [Unit] section, not [Service], so if you put them in the wrong section they are silently ignored.

Second, pipes and redirection do not work in ExecStart=. That is because systemd runs the command directly without going through a shell. If you need shell features, you have to run a shell explicitly, and for logs it is better to send standard output to the journal. Third, the exit code tells you the cause — 203/EXEC means the execution path is wrong or there is no execute permission, and 217/USER means the account written in User= does not exist.

There is one container trap as well. If you write the command in shell form, the shell becomes PID 1 and the application is its child. In that case SIGTERM goes to the shell, and the shell does not pass it on to the child, so the application never gets a grace period and dies from SIGKILL.

What you will do in the next lab

In this lab environment systemctl does not work. So the service lab concentrates on writing the unit file correctly. It covers the section layout, the execution and restart directives, timer units, the format difference between /etc/cron.d and a user crontab, cron's short PATH problem, log rotation settings, and even the idiom of clearing a list-type directive in a drop-in override. In the following lab you work through the boot order and GRUB settings, read /proc directly, and send signals to a background process to confirm the state changes.