Apache Hadoop — Stand up and run HDFS and YARN in one pod
Two ways to bring back a deleted file — trash and snapshots
In one line
There are two ways to undo a deletion in HDFS. The trash is a client-side device in which the shell moves a file aside instead of deleting it, and it is off by default. A snapshot is a device in which the NameNode holds on to one point in time of a directory as a whole, so it brings back even what was deleted with -skipTrash. In exchange, the space does not come back when you delete blocks that a snapshot holds.
Why two ways are needed
On a distributed file system, mistakes are big. A mistyped path in an rm -r erases several TB, and a wrong job overwrites yesterday's result. On a local disk you would fetch it from a backup, but HDFS data is often too large to back up.
So HDFS has two devices of different natures. One is a lightweight device aimed at a person's finger slip. It keeps a file deleted from the shell somewhere else for a while. The other is a device that keeps the past state of a directory itself. However someone deleted or overwrote it, you can go back to that point in time. The two differ in where they operate, in the range they protect, and in the price they cost.
How it works
First, the trash. The space reclamation section of the design document explains that if the trash setting is on, a file deleted with the FS shell is not deleted right away but moved to the trash directory. Each user has /user/<이름>/.Trash (the placeholder is the user name), and a file just deleted goes under .Trash/Current at its original path. At set intervals, Current is turned into a checkpoint with a date name, and checkpoints past their lifetime are deleted. Only then does the NameNode delete the name and the blocks are freed.
The lifetime of the trash is fs.trash.interval in core-default.xml, in minutes. The default is 0, and 0 means the trash is off. On a cluster that was only installed and whose settings were left untouched, rm is an immediate deletion. The lab image of this course has turned this value on at 1440 minutes, that is, one day. This value can be set on both the server and the client; if the server side is on, the server value is used and the client value is ignored. Only when the server side is off does it look at the client setting.
Here the nature of this device shows. The trash is moving instead of deleting, and the shell makes that decision. The description of rm in the shell documentation also says that if the trash is on, it moves the file to the trash directory. So with -skipTrash it is deleted immediately, and for programs that delete directly through the file system API without going through the shell, it is safer not to assume this safety net exists. It cannot stop overwrites either. The trash keeps only "deleted files".
A snapshot is held by the NameNode
The snapshot documentation defines a snapshot as a read-only point-in-time copy of the file system. You need not be afraid of the word copy. The properties of the implementation that the documentation points out are the key.
- The cost of creating one is O(1). It finishes instantly.
- DataNode blocks are not copied. A snapshot only records the file's block list and size.
- Additional memory is used only for what changed after the snapshot. It is O(M) for M changed files and directories.
- The changes are recorded in reverse chronological order, so the current data is read directly as is. When you read the snapshot side, it is computed by subtracting what changed from the current state.
So a snapshot is a NameNode metadata device. Even if you delete a file, the snapshot holds its block list, so the blocks are not freed. That is why even a file deleted with -skipTrash can be brought back.
hdfs dfsadmin -allowSnapshot /data/sales # 관리자: 스냅샷을 허용
hdfs dfs -createSnapshot /data/sales s1 # 소유자: 한 시점을 찍는다
hdfs dfs -rm -skipTrash /data/sales/2026-09.csv # 휴지통을 건너뛴 삭제
hdfs dfs -cp -ptopax /data/sales/.snapshot/s1/2026-09.csv /data/sales/
hdfs snapshotDiff /data/sales s1 . # s1 과 지금의 차이
A snapshot is read by name under a reserved path called .snapshot. Bringing it back is copying from that path. The documentation's example uses -ptopax to preserve timestamps, owner, permissions, ACLs and extended attributes together. snapshotDiff shows the difference between two snapshots, or between a snapshot and the present (.), with + (created), - (deleted), M (modified) and R (renamed). What was moved outside the snapshot directory appears as deleted, and what came in from outside appears as newly created.
If you create one without giving a snapshot name, it is named by the creation time, like s20130412-151029.033. Which directories are snapshottable is seen with hdfs lsSnapshottableDir, and the list of snapshots of a directory with hdfs lsSnapshot. If you take them regularly, putting the date in the name makes it easier to run a retention policy.
The constraints are also clear. Allowing snapshots is the superuser's job, and creating, deleting and renaming are the job of that directory's owner. You cannot allow it again on an ancestor or a descendant of a directory where snapshots are allowed (no nesting). The number of snapshots that can exist at the same time on one directory is 65,536, and the default of dfs.namenode.snapshot.max.limit in hdfs-default.xml is also 65536. And a directory that still has snapshots cannot be deleted or renamed. You have to delete all the snapshots.
Putting the two ways side by side
| Trash | Snapshot | |
|---|---|---|
| Where | The client (shell) moves it | The NameNode holds it as metadata |
| Default state | Off (fs.trash.interval 0) |
Must be allowed per directory |
| What it protects | Files deleted from the shell | Deleting, overwriting and renaming, all of them |
| What it cannot protect | -skipTrash, overwriting |
Changes made before the snapshot was taken |
| Space | Freed after the lifetime passes | Held until you delete the snapshot |
What it looks like in the field
First, check whether the trash is on. The default is 0, so a newly created cluster usually has it off. If there was no "moved to trash" message after rm, it was an immediate deletion.
Second, snapshots hold space. A snapshot holds the blocks of the deleted file, so even if you delete, the free space does not increase. As the shell documentation says, count counts even what is inside snapshots unless you give -x. A retention policy that regularly deletes old snapshots must come with them.
Third, snapshot diffs are the raw material of incremental replication. -diff in the DistCp documentation is used together with -update, and uses the diff report of two snapshots to find the difference between the source and the target and apply it to the target. It moves only what changed, instead of scanning everything each time.
Fourth, a directory where you took snapshots gets blocked at cleanup. If you finish a project and try to delete the directory and it fails, look first at the remaining snapshots.
What really matters in practice
- The trash is a shell feature and is off by default. If
fs.trash.intervalis 0,rmis an immediate deletion. - A snapshot does not copy blocks. The cost of creating one is O(1), and the cost is incurred only for what changed.
- Bringing back is a copy from
.snapshot. With-ptopax, permissions and ACLs are preserved too. - A snapshot holds the space of deleted files. Set a retention period and delete old ones.
- A directory with snapshots cannot be deleted or renamed.
What you will do in the next lab
You allow snapshots on a directory and take the first snapshot, and then in turn see a file deleted with a plain rm go to the trash (this lab image has the trash turned on for one day) and bring back a file deleted by skipping the trash by copying it from the snapshot. You append a line to the end of a file, take a second snapshot, and read what changed with snapshotDiff. You confirm that trying to delete a whole directory that still has snapshots is rejected, rename the second snapshot, and then summarize the difference between the trash and snapshots in a report.