Apache Hadoop — Stand up and run HDFS and YARN in one pod
Open fsimage and edits and read the namespace
Goal
You open up with your own hands how the NameNode leaves the namespace on disk: the fsimage, a snapshot of one point in time, and the edits, the record of changes after it. You see writes rejected in safe mode, make a new fsimage with saveNamespace, and read the contents with the Offline Image Viewer (oiv) and the Edits Viewer (oev).
Why it matters
The NameNode works holding the whole file tree in memory. Memory disappears when it is turned off, so it leaves two things on disk. The fsimage is the whole namespace up to some transaction number (txid), and the edits are a record of the changes after it (creating directories, closing files, renaming …), one line each. When the NameNode comes back up, it reads the fsimage and reapplies the edits after it.
If the edits get long, startup gets slow, so they must be merged into the fsimage periodically (a checkpoint). In production the secondary NameNode or standby NameNode does it, and in this lab you do it yourself with the administrator command saveNamespace. The reason that command requires safe mode is that the namespace must not change while it is being saved.
Safe mode is a state in which the NameNode accepts only reads. Right after startup it enters automatically until enough block reports from DataNodes have gathered, and administrators also enter by hand for maintenance. It is a common cause of the outage "every write fails".
If you can read an fsimage offline, you can analyze the whole list of files, sizes and owners without putting load on a live NameNode. Finding where small files are concentrated usually starts here.
Steps
- Save the output of
hdfs dfsadmin -safemode getto /root/hdp/meta/safemode.txt. - Upload
/data/finance/q1.csvto /user/root/meta/q1.csv, then enter safe mode (-safemode enter), try to create the directory /user/root/meta/in-safemode, and save the error output to /root/hdp/meta/deny.txt. - While still in safe mode, run
hdfs dfsadmin -saveNamespace, and write the txid (the number in the file name) of the newly created fsimage as an integer to /root/hdp/meta/fsimage_txid.txt. - Leave safe mode (
-safemode leave) and create the directory /user/root/meta/after-save. - Unpack the fsimage from step 3 with
hdfs oiv -p XMLinto /root/hdp/meta/fsimage.xml. - Unpack the same fsimage with
hdfs oiv -p Delimitedinto /root/hdp/meta/fsimage.tsv. - Close the edits segment currently being written with
hdfs dfsadmin -rollEdits, unpack withhdfs oevthe closed segment (edits_<시작>-<끝>, where the placeholders are the start and end txids) that contains the directory creation of step 4 into /root/hdp/meta/edits.xml, and then write that segment name and the number of OP_MKDIR to /root/hdp/meta/ops.json as{"segment": "edits_…-…", "OP_MKDIR": 정수}(the second value is an integer). - In /root/hdp/meta/report.md, write three sections:
## 안전 모드,## fsimageand## edits(use exactly these Korean headings in this order; the first means "Safe mode"). Put the txid from step 3 in the second section and the OP_MKDIR count from step 7 in the third.
Notes
- The NameNode's metadata directory:
/var/lib/hadoop/name/current, containingfsimage_<txid>(and.md5),edits_<시작>-<끝>(a closed segment),edits_inprogress_<시작>(being written) andseen_txid. - Only the two most recent fsimages are kept (
dfs.namenode.num.checkpoints.retained). If you ran saveNamespace several times, write the txid of the last one. hdfs oiv -p XML -i <fsimage> -o <출력>,hdfs oiv -p Delimited -i <fsimage> -o <출력>,hdfs oev -i <edits 파일> -o <출력>(the default output format is XML). The placeholders mean the output path and the edits file.- Common mistakes: moving on to the next step without leaving safe mode (every write fails); and trying to open
edits_inprogress_with oev (open a closed segment). - Official documentation: HDFS Architecture — The Persistence of File System Metadata · Offline Image Viewer · Offline Edits Viewer · HDFS Commands — dfsadmin
Is it in safe mode
Save the output of hdfs dfsadmin -safemode get to /root/hdp/meta/safemode.txt.
Right after startup it is in safe mode until block reports gather and then leaves by itself. It should be off now. If it is on, the NameNode is still starting up or the disk is short.
Writes are rejected in safe mode
Upload /data/finance/q1.csv to HDFS /user/root/meta/q1.csv, enter safe mode with hdfs dfsadmin -safemode enter, and save the error output of hdfs dfs -mkdir /user/root/meta/in-safemode to /root/hdp/meta/deny.txt.
Safe mode is a state in which the namespace is frozen. Reads (ls, cat) work, but everything that changes names is rejected. The saveNamespace of the next step requires exactly this state, because the tree must not change while it is being saved.
saveNamespace: a checkpoint by hand
While still in safe mode, run hdfs dfsadmin -saveNamespace, and write the txid of the newly created fsimage_<txid> in /var/lib/hadoop/name/current (where the placeholder is the txid) as an integer to /root/hdp/meta/fsimage_txid.txt.
saveNamespace writes the in-memory namespace as a new fsimage and moves the edits to a new segment. The number in the file name is "the last transaction number reflected in this image". Drop the leading zeros and write it as an integer.
Leave safe mode and write again
Leave safe mode with hdfs dfsadmin -safemode leave and create the directory /user/root/meta/after-save.
This directory is not in the fsimage of step 3; it is recorded only in the edits. You will confirm that difference with your own eyes in steps 5 to 7.
Unpack the fsimage as XML
Unpack the fsimage file from step 3 (fsimage_<txid>, where the placeholder is the txid) with hdfs oiv -p XML -i <파일> -o /root/hdp/meta/fsimage.xml (the placeholder is the input file).
In the XML, the NameSection has this image's txid, and the INodeSection has the name, permissions and blocks of each inode (file or directory). Find q1.csv's block size and confirm that after-save is absent. The grader checks those two as well.
Unpack it as a table: this is easier for analysis
Unpack the same fsimage with hdfs oiv -p Delimited -i <파일> -o /root/hdp/meta/fsimage.tsv (the placeholder is the input file).
Delimited is one line per inode, and the columns are path, replication factor, modification time, block size, number of blocks, file size, quota, permissions, owner and group. When looking for "which directory has small files concentrated" on a cluster of millions of files, you read this table with a spreadsheet or Spark.
Open the edits and read the change records
Close the segment currently being written with hdfs dfsadmin -rollEdits, and unpack the closed segment edits_<시작>-<끝> (the placeholders are the start and end txids) that contains the directory creation of step 4 with hdfs oev -i <파일> -o /root/hdp/meta/edits.xml (the placeholder is the input file). Write that segment name and the number of OP_MKDIR in it to /root/hdp/meta/ops.json as {"segment": "edits_…-…", "OP_MKDIR": 정수} (the second value is an integer).
The edits are records in which each transaction carries an OPCODE and a TXID. The after-save OP_MKDIR is in the segment that starts right after the txid of step 3. Do not open edits_inprogress_, which is not closed.
Where and how the metadata is left
In /root/hdp/meta/report.md, write three sections: ## 안전 모드, ## fsimage and ## edits (use exactly these Korean headings in this order; the first means "Safe mode"). Put the txid from step 3 in the second section and the OP_MKDIR count from step 7 in the third section.
Write in what order the NameNode uses the fsimage and the edits when it comes back up, and where after-save existed and where it did not.