Why NFS Does Not Behave Like a Local Disk
In one line
NFS looks like a filesystem but is an RPC call over the network. So it has failure modes that a local disk does not (hangs, partial failures, cache inconsistency).
Why you need this
Whether in a homelab or at a company, the moment comes when several servers have to see the same data. On Kubernetes, for Pods on several nodes to use the same PVC, you need ReadWriteMany, and the easiest implementation is NFS. But once you attach NFS, strange things start to happen. ls hangs for 30 seconds, processes sit in the D state and will not die, and a file is visible on some nodes but not on others.
How it works
Server side — exports
# /etc/exports
/export/share 10.0.0.0/24(rw,sync,no_subtree_check,root_squash)
/export/ro 10.0.0.0/24(ro,sync,no_subtree_check)
| Option | Meaning |
|---|---|
rw / ro |
Allow writes / read-only |
sync |
Respond after the write reaches disk. Slower but safe |
async |
Write only to memory and respond immediately. Data is lost if the server dies |
root_squash |
Map the client's root to nobody (the default) |
no_root_squash |
Accept root as is. In effect, handing over the server |
no_subtree_check |
Skip subtree checking. The currently recommended default |
There is a lot of confusion from not understanding root_squash. You created a file as root on the client, but the owner shows as nobody. That is not a bug but a security feature. NFS trusts UIDs as they are, so without squash, a single sudo on a client could touch any file on the server.
And UIDs/GIDs are passed as numbers, not names. If appuser on the server is 1001 but 1001 on the client is webuser, file ownership looks wrong. That is why environments that use NFS either unify UIDs centrally (LDAP) or configure NFSv4's idmapd.
Client side — mount options
nfs01:/export/share /mnt/share nfs4 _netdev,rw,soft,timeo=600,retrans=2,noatime 0 0
The most important choice is hard or soft.
| Option | When the server does not respond |
|---|---|
hard (default) |
It retries forever. The process hangs in the D state and even kill has no effect |
soft |
After timeo×retrans, it returns an I/O error |
intr |
(Old versions) Even with hard, it could be interrupted by a signal. On recent kernels this is absorbed into the default behavior |
hard is safe from the data consistency standpoint — writes are not treated as failures, so the application never proceeds in a wrong state. In exchange, if the server dies, the client freezes altogether. soft is the opposite. If there is no response, it returns EIO, so the system stays alive, but if the application does not handle that error properly, data gets corrupted.
The practical judgment is usually this. For databases or workloads where writes matter, hard; for mostly read workloads or cache-like data that can be missing, soft.
Caching and consistency
The NFS client caches attributes for performance. The default acregmin/acregmax are 3–60 seconds. So a file changed on the server may not be visible on the client immediately. A design in which several nodes write the same file at the same time is risky on NFS. Turning off the cache with noac improves consistency but greatly reduces performance.
File locking (flock, fcntl) is supported in NFSv4, but lock recovery is not perfect when the server restarts or the network drops. It is better to avoid designs that depend on locking over NFS.
What it looks like in the field
df hangs. If one NFS server dies and it is mounted with hard, even df does not respond. In that case, work around it with df -l (local only) or timeout 5 df. It is also common for a monitoring script that uses df to hang along with it in this situation.
The boot hangs. If you leave _netdev out of fstab, it tries to mount before the network is ready. Combined with hard, the boot effectively goes into an infinite wait.
On Kubernetes, a Pod will not get out of Terminating. If the NFS volume does not respond, the kubelet cannot unmount it and the Pod never terminates. In that case you have to bring the NFS server back or force-unmount on the node (umount -f -l).
What comes next
This module covers concepts only. However, you should now be able to explain why the NFS entry you wrote in the earlier fstab lab had that combination of options.