Apache Hadoop — Stand up and run HDFS and YARN in one pod
Restore an accidentally deleted file from the trash and from a snapshot
Goal
You allow snapshots on the orders folder and take one, then bring back a file deleted with the shell's rm from the trash and a file deleted with -skipTrash from the snapshot. After modifying a file, you look at the difference between two snapshots, and also try that a directory with snapshots cannot be deleted whole and renaming a snapshot.
Why it matters
There are two ways to bring back a deleted file, and they differ in nature. The trash is something the client (the shell) does: hdfs dfs -rm moves the file to .Trash/Current in the user's home instead of deleting it. So if you give -skipTrash, or if a non-shell program (Spark, an API) deletes it, it does not go through the trash.
A snapshot is something the NameNode does. It remembers the directory tree at the moment the snapshot is taken as metadata, and whatever is deleted or changed afterward, it does not delete the blocks on the snapshot side. It does not copy data, so taking one finishes instantly, and in exchange the space for the blocks the snapshot holds is not returned.
In operations, you use both together. The trash undoes a person's mistake and the snapshot undoes a job's mistake (a wrong overwrite, a deleted partition). The diff report (snapshotDiff) answers "what changed since yesterday", and distcp's incremental copy uses it too.
Steps
- Upload
a.csv(=/data/finance/q1.csv),b.csv(=q2.csv) andc.csv(=/data/small/sensor-0000.csv) to /snap/orders and allow snapshots (hdfs dfsadmin -allowSnapshot). - Take the snapshot s1 (
hdfs dfs -createSnapshot /snap/orders s1). - Delete a.csv with
hdfs dfs -rm /snap/orders/a.csv(it goes to the trash). - Delete b.csv with
hdfs dfs -rm -skipTrash /snap/orders/b.csv, and then bring it back by copying from /snap/orders/.snapshot/s1/b.csv. - Append one line to the end of c.csv (
-appendToFile), take the snapshot s2, and save the output ofhdfs snapshotDiff /snap/orders s1 s2to /root/hdp/snap/diff.txt. - Try
hdfs dfs -rm -r -skipTrash /snap/ordersand save the error output to /root/hdp/snap/deny.txt. - Rename the snapshot s2 to before-audit (
hdfs dfs -renameSnapshot). - In /root/hdp/snap/report.md, write three sections:
## 휴지통,## 스냅샷and## 지킬 것(use exactly these Korean headings in this order; they mean "Trash", "Snapshot" and "What to protect"). Put the number of lines (excluding the header line) in the diff report of step 5 in the second section.
Notes
- The trash retention time of this image is one day (
fs.trash.interval=1440minutes). With the default 0, the trash is off andrmdeletes immediately. - A file in the trash is at
/user/<사용자>/.Trash/Current/<원래 경로>(the placeholders are the user and the original path). To bring it back, just move it back in place withhdfs dfs -mv. - A snapshot is visible read-only under
<디렉터리>/.snapshot/<이름>/(the placeholders are the directory and the snapshot name). It does not appear in-ls, but you can enter it directly by path. - Common mistakes: trying to take a snapshot on a directory where snapshots are not allowed; trusting that a snapshot will stop
-rm -r -skipTrashand writing in another directory; and trying tomvfrom a snapshot (it is read-only, so bring it back withcp). - Official documentation: HDFS Snapshots · HDFS Architecture — Space Reclamation · FileSystem Shell — rm · WebHDFS — Snapshot Operations
Allow snapshots
Create HDFS /snap/orders, upload /data/finance/q1.csv as a.csv, /data/finance/q2.csv as b.csv and /data/small/sensor-0000.csv as c.csv, and run hdfs dfsadmin -allowSnapshot /snap/orders.
You cannot take a snapshot in just any directory. It works only in a directory (snapshottable) for which an administrator has allowed "this directory may have snapshots taken". See whether snapshotEnabled appears in WebHDFS's GETFILESTATUS.
Take the snapshot s1
Take the snapshot s1 with hdfs dfs -createSnapshot /snap/orders s1.
Taking one finishes instantly. It does not copy blocks; the NameNode merely remembers the tree at that moment. See it with hdfs dfs -ls /snap/orders/.snapshot/s1.
rm moves it to the trash
Delete a.csv with hdfs dfs -rm /snap/orders/a.csv. The file should have gone to /user/root/.Trash/Current/snap/orders/a.csv.
The shell's rm moves instead of deleting. "Moved: … to trash at …" in the output means that (you will not see it if you set the log level to WARN). To bring it back, just move it back from that path with mv.
What was deleted with -skipTrash comes from the snapshot
Delete b.csv with hdfs dfs -rm -skipTrash /snap/orders/b.csv, and then bring it back with hdfs dfs -cp /snap/orders/.snapshot/s1/b.csv /snap/orders/b.csv.
-skipTrash skips the trash, so it is not in the trash. But s1 holds that file's blocks, so you can read it through the snapshot path. A snapshot is read-only, so it is cp, not mv.
The difference between two snapshots
Append one line to c.csv, like echo 2026-03-01,s-9999,99.9 | hdfs dfs -appendToFile - /snap/orders/c.csv, take the snapshot s2, and save the output of hdfs snapshotDiff /snap/orders s1 s2 to /root/hdp/snap/diff.txt.
The symbols of the diff report are M (modified), + (created), - (removed) and R (renamed). Think about why b.csv appears twice, as '-' and '+': what you brought back is a new file with the same name. The grader compares your file line by line with WebHDFS's GETSNAPSHOTDIFF.
A snapshot protects the directory
Run hdfs dfs -rm -r -skipTrash /snap/orders and save the error output to /root/hdp/snap/deny.txt. /snap/orders must remain as it is.
A snapshottable directory with even one snapshot cannot be deleted. Deleting a directory is a single request sent to the NameNode, so it is rejected as a whole, and the files inside remain as they are. To delete it, you have to delete all the snapshots (-deleteSnapshot) and withdraw the allowance.
Rename a snapshot
Rename the snapshot s2 to before-audit with hdfs dfs -renameSnapshot /snap/orders s2 before-audit.
If you name snapshots by date or event, they are easy to find later. Even if you rename it, the contents and the diff report stay the same.
Record the two ways of bringing back
In /root/hdp/snap/report.md, write three sections: ## 휴지통, ## 스냅샷 and ## 지킬 것 (use exactly these Korean headings in this order; they mean "Trash", "Snapshot" and "What to protect"). Put the number of lines in the diff report of step 5 (the number of symbol lines, excluding the first header line) in the second section.
Write whose feature the trash is and when it is useless, what a snapshot remembers and what it holds, and what you would set as a rule to protect this folder.