Apache Hadoop — Stand up and run HDFS and YARN in one pod
Split a 12 MiB file into blocks and follow them down to the disk
Goal
You upload a 12MiB file with a 4MiB block, see with fsck that it splits into three blocks, and find and confirm for yourself that those blocks are ordinary files on DataNode disks. You also see what happens in a cluster with only one DataNode when you raise the replication factor, how small files eat blocks, and how the block size changes the file checksum.
Why it matters
In HDFS, the block is the unit of storage, replication and parallel processing. A file is cut at the block size (default 128MB), each block is replicated independently across several DataNodes, and MapReduce or Spark usually attach one task to one block. It is the writer (the client) that decides the block size, so it can differ per file. The replication factor also differs per file. The NameNode counts "how many replicas are there now" for each block and, if short, instructs copying to other DataNodes. If there is nowhere to copy to, it remains under-replicated and fsck reports it. In operations, the under-replicated and missing block counts in fsck are the first health indicators you look at. The fact that a block is an ordinary file on a DataNode disk matters when handling failures. If one disk dies, the block files on it disappear, and the NameNode learns this from block reports and refills them from other replicas.
Steps
- Create /user/root/blocks in HDFS and upload
/data/blobs/sample-12m.binto /user/root/blocks/big.bin with a block size of 4MiB (-D dfs.blocksize=4194304). - Save the output of
hdfs fsck /user/root/blocks/big.bin -files -blocks -locationsto /root/hdp/blocks/fsck.txt. - Find the file of the first block under the DataNode data directory (
/var/lib/hadoop/data) and write its absolute path as one line to /root/hdp/blocks/block0.path. - Upload the same local file without giving a block size to /user/root/blocks/big-default.bin.
- Raise the replication factor to 3 with
hdfs dfs -setrep 3 /user/root/blocks/big.bin(without-w), and save the output ofhdfs fsck /user/root/blocks/big.binto /root/hdp/blocks/underrep.txt. - Upload the 50 files
/data/small/sensor-0000.csvtosensor-0049.csvto /user/root/blocks/tiny/, and write the number of files, the number of blocks and the sum of bytes to /root/hdp/blocks/tiny.json as{"files": 정수, "blocks": 정수, "bytes": 정수}(all three values are integers). - With
hdfs dfs -checksum, get the checksums of big.bin and big-default.bin once in the default mode and once with-D dfs.checksum.combine.mode=COMPOSITE_CRC, and save the four lines to /root/hdp/blocks/checksum.txt. - In /root/hdp/blocks/report.md, write three sections:
## 블록으로 쪼개기,## 복제and## 체크섬(use exactly these Korean headings in this order; they mean "Splitting into blocks", "Replication" and "Checksum"). Put the number of blocks from step 1 in the first section, and the number of under-replicated blocks from step 5 and the number of blocks from step 6 in the second section.
Notes
- The block file is
/var/lib/hadoop/data/current/BP-<블록 풀 ID>/current/finalized/subdir*/subdir*/blk_<ID>(where the placeholder in BP-… is the block pool ID). Theblk_<ID>_<세대>.metabeside it is the checksum file (the second placeholder is the generation stamp). Find it withfind /var/lib/hadoop/data -name 'blk_<ID>'. -setrep -wwaits until replication is complete, but because there is one DataNode, it never finishes. Use it without-w.- Besides the shell command (
hdfs fsck), you can also call fsck through the NameNode web at/fsck?ugi=root&path=...&files=1&blocks=1. - Common mistakes: giving a block size smaller than 1MiB (you hit
dfs.namenode.fs-limits.min-block-size), and writing the.metafile as the block file. - Official documentation: HDFS Architecture — Data Replication · HDFS Commands — fsck · hdfs-default.xml · FileSystem Shell — checksum
Upload with a 4MiB block
Create /user/root/blocks in HDFS and upload /data/blobs/sample-12m.bin (12MiB) to /user/root/blocks/big.bin with hdfs dfs -D dfs.blocksize=4194304 -put.
The block size is not a cluster setting but a value the client decides at the moment of writing. It is 12MiB ÷ 4MiB, so first calculate how many blocks there will be. The grader looks at the block size and the number of blocks that the NameNode remembers.
See the blocks with fsck
Save the output of hdfs fsck /user/root/blocks/big.bin -files -blocks -locations to /root/hdp/blocks/fsck.txt.
For each block, one line shows BP-…:blk_<ID>_<세대> (the ID and the generation stamp), the length, the number of live replicas, and the DataNode holding the replicas. The block pool (BP) ID is the name tag of this NameNode's namespace. The grader compares the block IDs you saved with the current fsck result.
A block is a file on disk
Find the file of the first block you saw in step 2 under /var/lib/hadoop/data, and write its absolute path as one line to /root/hdp/blocks/block0.path.
Find it with find /var/lib/hadoop/data -name 'blk_<ID>' (the side without .meta attached, where the placeholder is the block ID). Compare that file's size and the first 4MiB of the original with cmp: an HDFS block is an unprocessed piece of bytes. The grader looks at it that way too.
Upload with the default block size
Upload the same /data/blobs/sample-12m.bin to /user/root/blocks/big-default.bin without giving a block size.
The default block size is 128MB, so a 12MiB file fits entirely in one block. Even if the block is larger than the file, it does not eat 128MB of disk: a block file is written only to its actual length. But as you will see in the quota lab, the reservation is made in units of the block size.
If you raise the replication factor and there is only one DataNode
Raise the replication factor to 3 with hdfs dfs -setrep 3 /user/root/blocks/big.bin (without -w), and save the output of hdfs fsck /user/root/blocks/big.bin to /root/hdp/blocks/underrep.txt.
The NameNode changes the target replication to 3, but there is no DataNode to copy to, so the blocks remain under-replicated. See Under-replicated blocks and Missing replicas in the fsck summary. If you give -w, it waits for this replication to finish, so the command never returns.
Even a small file takes one block
Upload the 50 files /data/small/sensor-0000.csv to sensor-0049.csv to /user/root/blocks/tiny/, and write the number of files, the number of blocks and the sum of bytes of that directory to /root/hdp/blocks/tiny.json as {"files": 정수, "blocks": 정수, "bytes": 정수} (all three values are integers).
Each file is only a few hundred bytes, yet there is one block per file. The NameNode holds files and blocks each as objects in memory, so 50 small files are 100 objects. Get the number of blocks from Total blocks in the fsck summary and the bytes from hdfs dfs -du -s.
The block size changes the checksum
Save to /root/hdp/blocks/checksum.txt the four lines of output of hdfs dfs -checksum /user/root/blocks/big.bin /user/root/blocks/big-default.bin and of hdfs dfs -D dfs.checksum.combine.mode=COMPOSITE_CRC -checksum on the same two files.
The default file checksum is "the MD5 of the MD5s of the chunk CRCs", so the block boundaries are mixed into the result. Even with the same content, the value differs if the block size differs, and that is why distcp verification goes wrong between clusters with different block sizes. COMPOSITE_CRC makes CRCs that are independent of block boundaries and gives the same value.
Blocks, replication and checksums in numbers
In /root/hdp/blocks/report.md, write three sections: ## 블록으로 쪼개기, ## 복제 and ## 체크섬 (use exactly these Korean headings in this order; they mean "Splitting into blocks", "Replication" and "Checksum"). Put the number of blocks from step 1 in the first section, and the number of under-replicated blocks from step 5 and the number of blocks from step 6 in the second section.
Write one or two sentences each on who decides the block size, when under-replication arises and who fixes it, and why the checksum is tied to the block size.