getent and dig Take Different Routes
In a nutshell
getent goes through the name service switch and dig does not. If the two results differ, the problem is not DNS but the stub resolver segment.
Why this was needed
You received a report that "DNS isn't working". You log in and type dig api.example.com, and it comes out fine. Yet the application keeps saying it cannot find the name.
If you conclude here that "dig works, so DNS is fine", you lose hours. This is because dig does not take the path the application uses.
The application calls glibc's getaddrinfo(), and glibc looks into /etc/nsswitch.conf, at its hosts: line, to decide in which order to ask. Usually it is files dns, so it looks at /etc/hosts first and then DNS. dig, on the other hand, skips all of this and asks the nameserver in /etc/resolv.conf directly.
So if there is a wrong entry in /etc/hosts, dig is normal and only the app goes to the wrong place. Conversely, for a name that exists only in hosts, the app works but dig cannot find it. The very fact that the two tools can give different answers for the same name is itself a diagnostic tool.
How it works
Name resolution is not one system but a path made of five segments of different natures joined together.
- Application + stub resolver -
/etc/hosts,/etc/nsswitch.conf, the runtime's own cache - Resolver selection -
/etc/resolv.conf - Recursive resolver cache - most queries end here
- Delegation tracing - root -> TLD -> authoritative server
- Authoritative server response
The reason for dividing into segments is that you can ask about each segment separately. getent looks including segment 1, dig starts from segment 2, dig @서버 +norecurse (the server goes where the Korean word is) looks only at that resolver's cache, and dig +trace walks segment 4 directly. If "+trace is normal but the app fails", the culprit is not the authoritative server but the resolver segment.
/etc/resolv.conf decides the arithmetic of latency.
- Only up to 3
nameserverentries are used options timeout:ndefaults to 5 seconds, andoptions attempts:ndefaults to 2 timesoptions ndots:n- the threshold for how many dots a name needs for it to be queried first as an absolute name without search domains
Two calculations come out of this. If the first nameserver is dead, the user feels the 5 seconds as they are with the defaults. And if ndots is large, several failed queries go out ahead of the resolution of a single external name - a typical cause of slow name resolution in containers. The solution is surprisingly simple: put a dot at the end of the name to make it an absolute name.
TTL and negative caching. DNS has no push. A change does not spread; caches expire. And the answer "it doesn't exist" is cached too. NXDOMAIN (the name itself does not exist) and NODATA (the name exists but only that type does not) are different events, and the cause of a situation in which you fix a typo in a record and it keeps failing for half a day is usually the negative cache time.
What it looks like in the field
Read the status of dig first. If it is NOERROR but the answer section is empty, it is NODATA. SERVFAIL means not "the name doesn't exist" but "it could not produce an answer", and indicates a failure to reach upstream or a DNSSEC validation failure. REFUSED is a refusal by policy.
The single letter aa in the flags. Without it, the answer is not authoritative but came from a cache. An authoritative server always gives the original TTL, and a resolver gives the remaining time. It is normal for the two to differ.
Split horizon. In a setup where the internal DNS and the external DNS give different answers to the same name, where you asked changes the answer. That is why, when you receive an outage report, you must first check "are you connected to the VPN".
One last principle. Once the name resolves normally, end the DNS investigation at that point and move on to the connection layer. Continuing to dig at DNS even though name resolution works is a typical form of wasted time.
The actual path by which names are resolved on Linux
The situation in which dig works but the program does not is frequent. This is because the two go different paths. dig asks the DNS server directly, and programs go through getaddrinfo().
What decides the order is /etc/nsswitch.conf.
hosts: files mdns4_minimal [NOTFOUND=return] dns
files is /etc/hosts, and it comes before DNS. So a single old entry left in the hosts file hides DNS entirely. After mdns4_minimal, the [NOTFOUND=return] means if it is not found there, do not go to DNS but fail, so .local domains are not passed on to DNS.
You see the answer a program sees with getent.
getent hosts api.example.com # nsswitch 를 그대로 따른다
getent ahostsv4 api.example.com # IPv4 만
In /etc/resolv.conf, search multiplies queries. When looking for one short name, it asks once for each search domain. With ndots:5, a name with fewer than 5 dots is tried with search appended first, so even a name like api.example.com is looked up as is only after going through search. This is the reason external domain lookups are slow in Kubernetes Pods. If you put a dot at the end of the name (api.example.com.), that process is skipped.
If you use systemd-resolved, there is one more layer. /etc/resolv.conf points to 127.0.0.53, and the actual servers are behind it. In that case, look at the state separately.
resolvectl status
resolvectl query api.example.com
There are caches in several places. The resolver daemon, the application (especially the JVM's networkaddress.cache.ttl), and the browser. The situation in which you changed a record but only one side sees the new value is usually the application cache. In some cases the JVM's default is a permanent cache, so during a failover it keeps holding on to the old address.
What you will do in the next lab
You will read /etc/resolv.conf and /etc/nsswitch.conf yourself to confirm how the stub resolver works, and add an entry to /etc/hosts to create yourself a situation in which getent finds it and dig does not. You will write the same name on two lines to see which one wins, and finally write a helper that reports the name resolution result through its exit code.