Apache Hadoop — Stand up and run HDFS and YARN in one pod
Simple authentication takes the name you give at face value
In one line
HDFS permission checking is almost the same as POSIX and even has ACLs, but HDFS does not judge "who you are". With the default simple authentication, it trusts the name the client gives as is, so permissions are a device for preventing mistakes, not for preventing attacks.
Why this distinction is needed
Permissions contain two questions: who are you (authentication) and is this allowed for that person (authorization). HDFS answers the latter carefully. The problem is the former.
The HDFS permissions guide explains that the way to identify users is selected with hadoop.security.authentication. In core-default.xml, the default of this value is simple, and the description adds in parentheses "no authentication". In simple mode, the identity of the client process is decided by the host operating system, and on Unix-like systems it is the value of whoami. The guide goes a step further and says: in any mode, the mechanism that determines user identity is outside HDFS, and HDFS itself has no feature for creating users or handling credentials.
So in this lab, if you give just the environment variable HADOOP_USER_NAME=alice, that shell becomes alice. There is no password. That this is a design, not a bug, is made clear by the secure mode documentation: in the default configuration, it is your job to block all network access so that attackers cannot reach the cluster, and if you want to restrict who accesses the data, you have to put authentication in place with Kerberos.
How it works
Permission bits. Every file and directory has an owner, a group, and permissions for three classes: the owner, the group members and all other users. For a file, r is read, and w is write and append. For a directory, r is listing, w is creating or deleting inside, and x is accessing children. There is no concept of an executable file, so there are no setuid or setgid bits. The owner of a newly created file is the client's identity, and the group follows the parent directory's group (the BSD rule). This is where people often get it wrong: the user's primary group is not attached, as in Linux.
Check order. If the user name equals the owner, only the owner permissions are looked at. Otherwise, if the file's group is in the user's group list, only the group permissions are looked at. If neither, the other permissions are looked at. Once it matches above, it does not look below. Also, every operation must be able to traverse the path. To reach /foo/bar/baz, x is needed on /, /foo and /foo/bar.
Deleting is the parent's business. In the guide's per-operation table, delete requires not the file's own permissions but w on the parent directory. This means that even if the file is read-only, you can delete it if you can write to the directory. To keep others from deleting one another's files in a shared directory that everyone writes to, you set the sticky bit. In a directory with the sticky bit set, only the superuser, the directory owner and the file owner can delete or move the files inside.
Groups are decided by the NameNode. The user name is stated by the client, but which groups that user belongs to is different. The group mapping documentation says that the mapping of users to groups in HDFS happens on the NameNode, so the system configuration of the NameNode host decides group membership. The default implementation uses the operating system's group lookup as is. Also, HDFS stores the user and group of a file as strings, not numeric IDs. So if the finance group does not exist on the NameNode machine, HDFS does not know it however many groups you create on the client side. The lookup result is cached for 300 seconds by default, so right after you put someone into a group, the check may still use the old membership.
The superuser. The identity that is the same as the NameNode process is the superuser, and permission checks never fail for the superuser. Members of dfs.permissions.superusergroup in hdfs-default.xml (default supergroup) are superusers too. In this lab, where the NameNode was started as root, root is the superuser.
ACLs: expressing exceptions that differ from the org chart
Permission bits express only "one owner, one group". If you want to open the finance folder to the finance group and give read access additionally to only bob, the audit officer, you cannot do it with bits. That is why ACLs exist. dfs.namenode.acls.enabled defaults to true in 3.5.0.
hdfs dfs -chown alice:finance /proj/finance
hdfs dfs -chmod 750 /proj/finance
hdfs dfs -setfacl -m user:bob:r-x /proj/finance # bob 에게 목록과 통과
hdfs dfs -setfacl -m default:user:bob:r-x /proj/finance # 앞으로 만들 자식에게 상속
hdfs dfs -getfacl /proj/finance
When an ACL is attached, two steps are inserted into the check order. After the owner, it looks at the named user entries, and after the group, at the named group entries. These two and the unnamed group entry are filtered once more by the mask. The mask is the upper limit of the permissions the extended entries can grant. The trap is that when you chmod a file with an ACL, it is not the group bits but the mask that changes. With one chmod 700, bob's read access is silently blocked.
The default ACL exists only on directories and is copied when new children are created. The guide says the copy happens only at the moment of creation, and if you later change the parent's default ACL, existing children do not change. So to give permissions to existing files, you have to run setfacl -R separately. The number of entries on one file is also limited to 32 access ACL entries and 32 default ACL entries, and a file with an ACL uses more NameNode memory. The approach the guide recommends is to solve most things with permission bits and add only a few exceptions with ACLs.
What it looks like in the field
First, "I blocked the permissions but someone read it." On a simple-authentication cluster, anyone can say their name is alice. WebHDFS too, if security is off, accepts the user.name query value as the user as is. Permissions only prevent a colleague's mistakes. If you need a real boundary, it is Kerberos, and until then network access control is the only wall.
Second, a new file's group is different from what you expected. It is because it follows the parent directory's group. If you set the team folder's group right first, the new files under it follow automatically.
Third, you put the user in the group but it is still rejected. Group membership is looked up and cached on the NameNode side, not the client. Check whether that group and its members exist on the NameNode machine and whether the cache is still holding the old value.
Fourth, you set a default ACL, but old files still cannot be read. Inheritance is a copy at creation time. For files that already existed, you have to set it separately.
Fifth, turning off permission checking is not a solution. If you set dfs.permissions.enabled to false, only the checking is turned off, and the mode, owner and group remain. The guide says that regardless of this value, chmod, chgrp, chown and setfacl always check permissions.
What really matters in practice
- HDFS only authorizes and leaves authentication to the outside. The default simple trusts the name it is told.
- Deleting is decided by w on the parent directory. Set the sticky bit on shared folders.
- A new file's group comes from the parent directory. Get the team folder's group right first.
- chmod on a file with an ACL changes the mask. The extended entries can all narrow at once.
- A default ACL is copied only at the moment of creation. Fix existing files separately.
What you will do in the next lab
You create the finance folder with alice as owner, finance as group and 750, and upload a file as alice. You switch the user with one environment variable and see bob rejected, then open read access to bob alone with an ACL and confirm that it actually reads. You set a default ACL and see whether a newly created file inherits those permissions, and confirm that trying to delete someone else's file in a shared folder with the sticky bit set is rejected. Finally, you write in the report what simple authentication cannot prevent and why Kerberos is needed.