TT Lab
Get started
Learn Learning paths Courses

Network Troubleshooting

The Order in Which You Read dig

Continue in TT Lab

In one line

You do not read dig output from the top; you read it in the order status → flags → each section. If you keep this order, a single response splits the cause.

Why this exists

The report "DNS isn't working" mixes at least five different situations. The name does not exist at all, only that type of record does not exist, the resolver cannot produce an answer, the server refuses to answer, and the cache is stale. The responses to these five are all different. And one dig output settles which of the five it is.

How it works

Let us look at the response code first.

status ANSWER Actual meaning Common cause What to do next
NOERROR 1 or more Both the name and the type exist Normal On to the next layer
NOERROR 0 The name exists but that type does not A exists and AAAA does not; only a CNAME exists and its target was not created Recheck with a different type
NXDOMAIN — The authoritative server confirms that the name itself does not exist Typo, not created, a name with a search domain appended Query the authoritative server directly
SERVFAIL — The resolver could not produce an answer DNSSEC validation failure, upstream not responding, zone load failure Resolver logs
REFUSED — The server has no intention of handling it Recursion not allowed, ACL, does not hold the zone Whether the query target is right

The case of NOERROR with ANSWER: 0 is the most confusing. The name exists but only that record type does not. A situation where only IPv6 is missing, or where a CNAME was set but the target A record was not created, falls here.

Next are the flags.

It is good to get combinations by purpose into your fingers.

dig +short A example.com               # 값만
dig +noall +answer example.com         # 답변 섹션만 (TTL·타입 포함)
dig +noall +authority +additional example.com
dig @1.1.1.1 example.com               # 특정 서버에 직접
dig +norecurse @10.0.0.53 example.com  # 캐시에 있는지만 확인
dig -x 8.8.8.8                         # 역방향
dig -t MX example.com
dig +tcp example.com
dig +time=2 +tries=1 example.com
dig +trace example.com                 # 루트부터 위임 추적

+trace has a trap. +trace queries each step directly starting from the root, so it bypasses your resolver. That is why "+trace is normal but the application fails" is common, and the culprit then is not the authoritative server but the resolver segment.

TTL and negative caching

If you repeat the same query to a resolver, you can see the TTL decreasing. On the other hand, if you ask an authoritative server directly, the original configured value always comes out. This difference is the easiest way to tell "whether it came from a cache".

The result for a nonexistent name is cached too. The negative cache time is the smaller of the SOA's MINIMUM field and the TTL of the SOA record itself. So it happens that NXDOMAIN continues for a while even after you created the record. The answer to "I created it so why doesn't it work" is usually this.

The numbers in resolv.conf

nameserver 10.0.0.2
search prod.svc.cluster.local svc.cluster.local cluster.local
options ndots:5 timeout:2 attempts:3 rotate

Kubernetes's ndots:5 is a famous trap. api.stripe.com has 2 dots, which is fewer than 5, so it appends all 4 search domains first and only then asks for the original name. Counting A/AAAA pairs, 8 of 10 queries go out just to receive NXDOMAIN. However, contrary to the common advice, lowering ndots to 1 breaks the cluster — cross-namespace calls such as payments.prod all fail. The safe lower bound is 2. In a hurry, putting one dot at the end of the name (api.stripe.com.) states explicitly that it is an FQDN and skips search.

What it looks like in the field

A firewall with only UDP 53 open. Usually there is no problem. Because responses are small. But the moment records grow or a DNSSEC signature is added and the response exceeds 512 bytes, only that name stops resolving. Sometimes this is the cause of the report "only one particular domain doesn't work". Reproduce it with dig +notcp +bufsize=512 and look at ;; MSG SIZE rcvd:.

Do not flush the whole cache. rndc flush feels satisfying, but afterward upstream queries surge for a few minutes. The standard way is to clear only the problem name with rndc flushname api.example.com.

What you will do in the next lab

You pull A/MX/NS/TXT with dig, observe the TTL decreasing, do a reverse lookup, and calculate the negative cache time from an NXDOMAIN response. At the end you put in a hosts entry and submit, as a table, how the answers of getent and dig diverge.