TT Lab
Get started
Learn Learning paths Courses

Apache Hadoop — Stand up and run HDFS and YARN in one pod

Bundle 2,000 small files into an archive and measure NameNode object counts

Continue in TT Lab

Goal

You upload 2,000 sensor CSVs (a few hundred bytes per file) to HDFS and measure how many objects (files, directories, blocks) the NameNode has to hold, then bundle them with a Hadoop Archive (HAR) to see how much that number drops. You also try reading the original files one by one through har:// even after bundling, and moving the archive whole with distcp.

Why it matters

The NameNode holds every file, directory and block as a Java object in its heap. Whether a file is 1 byte or 1GB, the number of objects is the same. So the limit of HDFS is usually not the disk but the NameNode's heap, and millions of small files bring the NameNode down first while using almost no disk. From the MapReduce and Spark side too, small files are slow because a task or open cost is attached to each file. The remedy is to make fewer files: either gather when writing (Spark's repartition and coalesce, large container file formats), or bundle what has already been created. A HAR is the latter approach. It turns many files into a few large files (part-*) and an index (_index and _masterindex), reducing the NameNode's object count, and when reading, the har:// file system looks at the index and finds the original path for you. In exchange, an archive, once made, cannot be modified (it is read-only). This lab runs the archive and distcp with the local executor, without YARN. Both are MapReduce jobs, but where they run changes with one setting: the same job runs on a laptop and on a cluster.

Steps

  1. Upload the /data/small directory (2,000 files) to HDFS /user/root/small/raw (uploading in several streams with hdfs dfs -put -t 8 is fast).
  2. Write the number of files, directories and blocks of raw to /root/hdp/small/objects.json as {"raw_files": 정수, "raw_dirs": 정수, "raw_blocks": 정수} (all three values are integers).
  3. Create the archive with hadoop archive -D mapreduce.framework.name=local -archiveName raw.har -p /user/root/small -r 1 raw /user/root/small/archive.
  4. Count the files inside the archive with hdfs dfs -ls -R har:///user/root/small/archive/raw.har and write the count as an integer to /root/hdp/small/har_files.txt.
  5. Read raw/sensor-0042.csv through the har:// path and save it locally as /root/hdp/small/sensor-0042.csv.
  6. Write the NameNode object counts (files + directories + blocks) of raw and raw.har to /root/hdp/small/compare.json as {"raw_objects": 정수, "har_objects": 정수} (both values are integers).
  7. Copy the archive with hadoop distcp -D mapreduce.framework.name=local /user/root/small/archive /user/root/backup/small-archive.
  8. In /root/hdp/small/report.md, write three sections: ## NameNode 가 치르는 값, ## HAR and ## 옮기기 (use exactly these Korean headings in this order; they mean "What the NameNode pays", "HAR" and "Moving it"). Put raw_objects from step 6 in the first section and har_objects in the second.

Notes

Upload 2,000 small files

Upload /data/small to HDFS /user/root/small/raw. Uploading in several streams, like hdfs dfs -put -t 8 /data/small /user/root/small/raw, is fast.

If you upload in one stream, each file makes several round trips with the NameNode, so 2,000 files take more than 10 minutes. Small files pay their cost not by size but by count: this alone is the small files problem.

Count the objects the NameNode holds

Find the number of files, directories and blocks of /user/root/small/raw with hdfs dfs -count and hdfs fsck, and write them to /root/hdp/small/objects.json as {"raw_files": 정수, "raw_dirs": 정수, "raw_blocks": 정수} (all three values are integers).

One block per file (because the file is smaller than the block size), plus one directory. In the NameNode heap, all three are objects. Even with a 128MB block size, the block of a file of a few hundred bytes does not take 128MB, but it takes exactly one object all the same.

Bundle with an archive

Create /user/root/small/archive/raw.har with hadoop archive -D mapreduce.framework.name=local -archiveName raw.har -p /user/root/small -r 1 raw /user/root/small/archive.

An archive is a MapReduce job. It divides the list of input files, concatenates the contents into part-*, and writes into _index where each file went. -D mapreduce.framework.name=local means to run it inside the current JVM without YARN. The original (raw) stays as it is.

Look inside with har://

From the output of hdfs dfs -ls -R har:///user/root/small/archive/raw.har, count the files (the lines whose permissions start with -) and write the count as an integer to /root/hdp/small/har_files.txt.

The har:// file system reads _index and shows the original directory tree. The NameNode does not know these 2,000: a file inside the archive is not a NameNode object but a line in the index.

Read one file inside the archive

Save the contents read with hdfs dfs -cat har:///user/root/small/archive/raw.har/raw/sensor-0042.csv to the local /root/hdp/small/sensor-0042.csv.

The path is har://<아카이브 경로>/<-p 기준 상대 경로>, where the first placeholder is the archive path and the second is the path relative to the -p parent. The har file system finds the position inside the part file through _masterindex and _index and reads only those bytes. Compare with the original /data/small/sensor-0042.csv using cmp.

Compare the object counts

Find the NameNode object count (number of files + number of directories + number of blocks) for each of raw and raw.har, and write them to /root/hdp/small/compare.json as {"raw_objects": 정수, "har_objects": 정수} (both values are integers).

raw is 2,000 files + 1 directory + 2,000 blocks. raw.har is just one directory, a few files (_index, _masterindex, part-0 and _SUCCESS) and their blocks. Calculate how many times smaller it has become.

Move it whole with distcp

Copy the archive with hadoop distcp -D mapreduce.framework.name=local /user/root/small/archive /user/root/backup/small-archive. The files inside raw.har of the copy must have the same lengths as the original.

distcp is a MapReduce job that divides the list of files and has several maps copy them in parallel. It is the standard tool for moving and backing up between clusters, and if you bundle small files, the number of files to copy drops, so this becomes faster too. If the target path does not exist, the contents of the source directory are copied under that name.

Put the price of small files in numbers

In /root/hdp/small/report.md, write three sections: ## NameNode 가 치르는 값, ## HAR and ## 옮기기 (use exactly these Korean headings in this order; they mean "What the NameNode pays", "HAR" and "Moving it"). Put raw_objects from step 6 in the first section and har_objects in the second, as numbers.

Write what the NameNode heap counts, what a HAR reduces and what it gives up (read-only), and what the writing side should do so that small files are not created in the first place.