TT Lab
Get started
Learn Learning paths Courses

Networking Fundamentals

DNS — Why dig Works but the Application Fails

Continue in TT Lab

In a nutshell

Because dig and applications resolve names through different paths in the first place, the fact that dig succeeded is no basis for saying the application will succeed.

Why this was needed

There is a combination that confuses people most in an outage. dig +short api.internal.example.com returns the address fine, but the application on the same host dies saying it cannot find the name. Here most people jump to the conclusion "DNS is fine, the application is weird" and start digging through libraries. This direction is almost always wasted effort.

How it works

First, the structure of DNS itself. A resolver asks a root server to find the TLD server, asks the TLD server to find the authoritative server, and gets the final answer from the authoritative server (a combination of recursive and iterative queries). Each response carries a TTL and is cached for that long.

The kinds of records are also worth knowing. A is an IPv4 address, AAAA is an IPv6 address, CNAME is an alias to another name, MX is a mail server, NS is an authoritative server, TXT is an arbitrary string, and SRV is the host and port of a service.

Now the key point. The path an application actually takes is different. C, Python, Ruby, PHP, and even Java on Linux mostly call glibc's getaddrinfo(), and glibc follows this order.

  1. It reads the hosts line of /etc/nsswitch.conf to decide which sources to use in which order.
  2. If it is files, it looks at /etc/hosts.
  3. If it is dns, it queries applying the nameserver, search, and options of /etc/resolv.conf.
  4. If it is resolve, it asks systemd-resolved.

dig, nslookup, and host skip this path entirely. They look at neither nsswitch.conf nor /etc/hosts, take only the nameserver address from resolv.conf, and throw the query out over UDP 53. So dig cannot find a name that exists only in /etc/hosts, and conversely, a situation arises in which only the application connects to an old IP because of a stale /etc/hosts entry someone left while debugging.

The tool that reproduces the same path as the application is getent hosts. If you use getent ahosts, it shows IPv4 and IPv6 in the order they are tried. To sum up: dig shows the DNS server's answer, and getent shows the answer the application will receive. If the two results differ, that difference itself is the location of the cause.

What it looks like in the field

The resolv.conf of a Kubernetes Pod contains options ndots:5 and four search domains. ndots means "if the number of dots is fewer than this value, try attaching the search list first". Thanks to this, you can use short names such as api or api.prod inside a Pod.

The problem is external domains. api.stripe.com has 2 dots, fewer than 5, so it tries attaching the four search domains first. Since it asks for A and AAAA together, to resolve one external domain, 10 queries go out, and 8 of them went out only to receive NXDOMAIN. If there are hundreds of Pods, most of the CoreDNS load is filled with queries that are certain to fail.

Here we must correct a commonly circulated prescription. The advice "just lower ndots to 1" breaks the cluster. If ndots is 1, cross-namespace calls with one dot such as payments.prod do not go through search and all fail. The safe lower bound is 2. A more reliable way is to put a trailing dot at the end of external domains to make them absolute names. One trailing dot removes the four search domains.

You must not lump response codes together either. NXDOMAIN is the definitive statement that the name does not exist, SERVFAIL means the resolver could not produce an answer, and a NOERROR with ANSWER 0 means the name exists but there is no record of that type. The last one is misdiagnosed most often. A typical case is when A exists but AAAA does not.

Seeing how far it got when a name does not resolve

"It seems to be a DNS problem" is a common remark, but about half the time it is actually DNS. Where it got stuck shows up if you ask step by step.

dig looks at the cache, and dig +trace looks at the chain.

dig example.com A +short          # 지금 내 리졸버가 아는 답
dig example.com A +trace          # 루트부터 권한 서버까지 따라간다
dig @8.8.8.8 example.com A        # 다른 리졸버는 뭐라고 하는가
dig example.com SOA +short        # 이 이름의 주인은 어느 서버인가

If it works with +short but not in the browser, it is not DNS but the next stage. If it works with @8.8.8.8 but not with the default resolver, it is your side's resolver or cache.

nslookup/dig and applications can see different answers. These two ask DNS directly, but most programs go through getaddrinfo(). On that path, /etc/hosts, /etc/nsswitch.conf, and mDNS get in. What a program actually sees is shown more accurately by getent hosts example.com.

TTL is not "from when it works" but "until when the old one lingers". Before changing a record, reduce the TTL to 300 seconds, change it after that TTL has passed, and raise it again once stable. If you do not keep this order, people go to different places not at the moment of the change but until the moment the old TTL has fully drained.

dig example.com A          # ANSWER SECTION 의 두 번째 열이 남은 TTL 이다

A name that does not exist is different from having no answer. NXDOMAIN means the name does not exist, and if it is NOERROR with ANSWER 0, the name exists but there is no record of that kind. This is where clients that try IPv6 first get slow when only AAAA is missing and A exists.

Search domains get appended silently. If /etc/resolv.conf has several domains in its search line, every lookup of a short name sends out that many queries. A well-known case is that, because of ndots:5 in a Kubernetes Pod, a lookup of an external domain succeeds after failing four times. If you write it with a trailing dot, as example.com., the search is skipped.

What to check in the quiz that follows

Check whether you can explain the difference between dig and getent, and whether you understand the mechanism of the latency that ndots creates.