TT Lab
Get started
Learn Learning paths Courses

Apache Hadoop — Stand up and run HDFS and YARN in one pod

Restore an accidentally deleted file from the trash and from a snapshot

Continue in TT Lab

Goal

You allow snapshots on the orders folder and take one, then bring back a file deleted with the shell's rm from the trash and a file deleted with -skipTrash from the snapshot. After modifying a file, you look at the difference between two snapshots, and also try that a directory with snapshots cannot be deleted whole and renaming a snapshot.

Why it matters

There are two ways to bring back a deleted file, and they differ in nature. The trash is something the client (the shell) does: hdfs dfs -rm moves the file to .Trash/Current in the user's home instead of deleting it. So if you give -skipTrash, or if a non-shell program (Spark, an API) deletes it, it does not go through the trash. A snapshot is something the NameNode does. It remembers the directory tree at the moment the snapshot is taken as metadata, and whatever is deleted or changed afterward, it does not delete the blocks on the snapshot side. It does not copy data, so taking one finishes instantly, and in exchange the space for the blocks the snapshot holds is not returned. In operations, you use both together. The trash undoes a person's mistake and the snapshot undoes a job's mistake (a wrong overwrite, a deleted partition). The diff report (snapshotDiff) answers "what changed since yesterday", and distcp's incremental copy uses it too.

Steps

  1. Upload a.csv (=/data/finance/q1.csv), b.csv (=q2.csv) and c.csv (=/data/small/sensor-0000.csv) to /snap/orders and allow snapshots (hdfs dfsadmin -allowSnapshot).
  2. Take the snapshot s1 (hdfs dfs -createSnapshot /snap/orders s1).
  3. Delete a.csv with hdfs dfs -rm /snap/orders/a.csv (it goes to the trash).
  4. Delete b.csv with hdfs dfs -rm -skipTrash /snap/orders/b.csv, and then bring it back by copying from /snap/orders/.snapshot/s1/b.csv.
  5. Append one line to the end of c.csv (-appendToFile), take the snapshot s2, and save the output of hdfs snapshotDiff /snap/orders s1 s2 to /root/hdp/snap/diff.txt.
  6. Try hdfs dfs -rm -r -skipTrash /snap/orders and save the error output to /root/hdp/snap/deny.txt.
  7. Rename the snapshot s2 to before-audit (hdfs dfs -renameSnapshot).
  8. In /root/hdp/snap/report.md, write three sections: ## 휴지통, ## 스냅샷 and ## 지킬 것 (use exactly these Korean headings in this order; they mean "Trash", "Snapshot" and "What to protect"). Put the number of lines (excluding the header line) in the diff report of step 5 in the second section.

Notes

Allow snapshots

Create HDFS /snap/orders, upload /data/finance/q1.csv as a.csv, /data/finance/q2.csv as b.csv and /data/small/sensor-0000.csv as c.csv, and run hdfs dfsadmin -allowSnapshot /snap/orders.

You cannot take a snapshot in just any directory. It works only in a directory (snapshottable) for which an administrator has allowed "this directory may have snapshots taken". See whether snapshotEnabled appears in WebHDFS's GETFILESTATUS.

Take the snapshot s1

Take the snapshot s1 with hdfs dfs -createSnapshot /snap/orders s1.

Taking one finishes instantly. It does not copy blocks; the NameNode merely remembers the tree at that moment. See it with hdfs dfs -ls /snap/orders/.snapshot/s1.

rm moves it to the trash

Delete a.csv with hdfs dfs -rm /snap/orders/a.csv. The file should have gone to /user/root/.Trash/Current/snap/orders/a.csv.

The shell's rm moves instead of deleting. "Moved: … to trash at …" in the output means that (you will not see it if you set the log level to WARN). To bring it back, just move it back from that path with mv.

What was deleted with -skipTrash comes from the snapshot

Delete b.csv with hdfs dfs -rm -skipTrash /snap/orders/b.csv, and then bring it back with hdfs dfs -cp /snap/orders/.snapshot/s1/b.csv /snap/orders/b.csv.

-skipTrash skips the trash, so it is not in the trash. But s1 holds that file's blocks, so you can read it through the snapshot path. A snapshot is read-only, so it is cp, not mv.

The difference between two snapshots

Append one line to c.csv, like echo 2026-03-01,s-9999,99.9 | hdfs dfs -appendToFile - /snap/orders/c.csv, take the snapshot s2, and save the output of hdfs snapshotDiff /snap/orders s1 s2 to /root/hdp/snap/diff.txt.

The symbols of the diff report are M (modified), + (created), - (removed) and R (renamed). Think about why b.csv appears twice, as '-' and '+': what you brought back is a new file with the same name. The grader compares your file line by line with WebHDFS's GETSNAPSHOTDIFF.

A snapshot protects the directory

Run hdfs dfs -rm -r -skipTrash /snap/orders and save the error output to /root/hdp/snap/deny.txt. /snap/orders must remain as it is.

A snapshottable directory with even one snapshot cannot be deleted. Deleting a directory is a single request sent to the NameNode, so it is rejected as a whole, and the files inside remain as they are. To delete it, you have to delete all the snapshots (-deleteSnapshot) and withdraw the allowance.

Rename a snapshot

Rename the snapshot s2 to before-audit with hdfs dfs -renameSnapshot /snap/orders s2 before-audit.

If you name snapshots by date or event, they are easy to find later. Even if you rename it, the contents and the diff report stay the same.

Record the two ways of bringing back

In /root/hdp/snap/report.md, write three sections: ## 휴지통, ## 스냅샷 and ## 지킬 것 (use exactly these Korean headings in this order; they mean "Trash", "Snapshot" and "What to protect"). Put the number of lines in the diff report of step 5 (the number of symbol lines, excluding the first header line) in the second section.

Write whose feature the trash is and when it is useless, what a snapshot remembers and what it holds, and what you would set as a rule to protect this folder.