Apache Hadoop — Stand up and run HDFS and YARN in one pod
Upload a file to a single-pod HDFS and read it two ways
Goal
You check the configuration and state of a pseudo-distributed HDFS running in one Pod, upload one access log, and read it in two ways: the shell (hdfs dfs) and HTTP (WebHDFS). You confirm with the fileId and the block ID that moving a file (mv) only changes the NameNode's name tag.
Why it matters
HDFS is a file system with its roles split in two. The NameNode holds in memory the directory tree and the metadata saying "this file is these blocks, and those blocks are on these DataNodes", and the DataNodes hold the bytes of the blocks as files on their own disks. The client asks the NameNode for a location and exchanges the bytes directly with the DataNode.
Knowing this separation explains a lot about operations. Renaming and moving do not touch bytes, so they finish instantly regardless of size, and if the NameNode stops, nothing can be read even though the bytes are fine. Where the configuration comes from (*-site.xml overrides *-default.xml) also matters for the same reason: if clients and daemons see different configurations, the same command behaves differently.
WebHDFS is a way of asking the same NameNode over HTTP. You can get listings, status and contents with just curl and no Java client, so scripts and monitoring usually use this route.
Steps
- Ask for
fs.defaultFS,dfs.replicationanddfs.blocksizewithhdfs getconf -confKeyand write them to /root/hdp/first/conf.txt as three키=값lines (key=value). - Save the output of
hdfs dfsadmin -reportto /root/hdp/first/report.txt. - Create the directory /user/root/first/logs in HDFS.
- Upload
/data/logs/access-2026-03-01.logto /user/root/first/logs/. - Count the lines of that file with
hdfs dfs -catand write the count as an integer to /root/hdp/first/lines.txt. - Call WebHDFS
LISTSTATUSwith curl and save the listing response for /user/root/first/logs to /root/hdp/first/liststatus.json. - Create /user/root/first/archive, move the log file to /user/root/first/archive/access-0301.log, and then save the output of
hdfs fsck <새 경로> -files -blocks(where the placeholder is the new path) to /root/hdp/first/blocks.txt. - In /root/hdp/first/report.md, write three sections:
## 두 역할,## 두 길and## 이름만 바뀐다(use exactly these Korean headings in this order; they mean "Two roles", "Two routes" and "Only the name changes"). Put the line count from step 5 in the second section and the fileId that stayed the same before and after the move in the third.
Notes
- HDFS is already running when the Pod starts (
lab-hadoop status). If it is off, turn it on withlab-hadoop start. The data remains on the disks of the NameNode and DataNode, so it is intact when you turn it back on. - Each
hdfs dfscommand is one JVM, so it takes 1–2 seconds. If you reduce logging withexport HADOOP_ROOT_LOGGER=WARN,console, the output is cleaner. - WebHDFS:
curl -s "http://localhost:9870/webhdfs/v1<경로>?op=LISTSTATUS&user.name=root"(where the placeholder is the path).op=OPEN, which fetches contents, redirects you to a DataNode (307), so you needcurl -L. - Common mistakes: confusing a local path with an HDFS path (
/user/root/...is inside HDFS); printing-reportbefore the NameNode is up; and thinking mv copies the file. - Official documentation: HDFS Architecture · Pseudo-Distributed Operation · FileSystem Shell · WebHDFS REST API
Where the configuration comes from
Ask for the three values fs.defaultFS, dfs.replication and dfs.blocksize with hdfs getconf -confKey <키> (where the placeholder is the key), and write them to /root/hdp/first/conf.txt as three key=value lines, in the form 키=값, such as fs.defaultFS=값.
The first two are values this image changed in core-site.xml and hdfs-site.xml, and nobody changed the block size, so the default from hdfs-default.xml comes out. The rule that the site files override the default files shows up in all three.
The cluster as the NameNode sees it
Save the output of hdfs dfsadmin -report to /root/hdp/first/report.txt.
The upper part of the report is the whole cluster (capacity, usage, under-replicated blocks), and the lower part is the state of each DataNode. See how many DataNodes are alive and what the configured capacity is. The grader reads the same numbers from the NameNode's JMX and compares them.
Make a place inside HDFS
Create the directory /user/root/first/logs in HDFS (including the intermediate directories).
It is hdfs dfs -mkdir -p. A directory exists only in the NameNode's metadata, and nothing is created on the DataNodes: a directory has no blocks.
Upload a file
Upload the local /data/logs/access-2026-03-01.log to HDFS /user/root/first/logs/ under the same name.
hdfs dfs -put gets a block location from the NameNode and streams the bytes to the DataNode. While uploading, it exists under a name with ._COPYING_ attached, and when done it becomes its real name. The grader compares the length and contents with the local original.
Read it through the shell
Read /user/root/first/logs/access-2026-03-01.log with hdfs dfs -cat, count the lines, and write the count as an integer to /root/hdp/first/lines.txt.
hdfs dfs -cat 경로 | wc -l, where the placeholder stands for the path. The bytes come directly from the DataNode. Also compare for yourself whether the line count equals that of the local original.
Read it through HTTP: WebHDFS
Call http://localhost:9870/webhdfs/v1/user/root/first/logs?op=LISTSTATUS&user.name=root with curl and save the response JSON as is to /root/hdp/first/liststatus.json.
In the response's FileStatuses.FileStatus, each file has its length, block size, replication factor, owner and fileId (the inode number). Remember the fileId: in the next step you will see whether it is the same after the move.
Moving only changes the name tag
Create /user/root/first/archive in HDFS and hdfs dfs -mv the log file to /user/root/first/archive/access-0301.log. Then save the output of hdfs fsck /user/root/first/archive/access-0301.log -files -blocks to /root/hdp/first/blocks.txt.
mv changes only the name inside the NameNode. So the fileId stays the same and the block IDs stay the same, and it finishes instantly even if the file is 1TB. The grader compares the fileId you saved in step 6 with the current fileId, and the block ID you saved with the current block ID.
Record the two roles in numbers
In /root/hdp/first/report.md, write three sections: ## 두 역할, ## 두 길 and ## 이름만 바뀐다 (use exactly these Korean headings in this order; they mean "Two roles", "Two routes" and "Only the name changes"). Put the line count from step 5 in the second section and the fileId that stayed the same before and after the move in the third section, as numbers.
In the first section, write what the NameNode and the DataNode each hold; in the second, how the shell and WebHDFS showed the same file; and in the third, what stayed the same after the mv.