The certificate was renewed, but the browser still showed the old one
DNS has no propagation, only caches and TTLs
In one line
Turning a name into an address follows a fixed order: /etc/hosts → resolver configuration → cache → authoritative server. At any point along the way, a stale answer can live on for as long as its TTL. The answer to "I changed the DNS record, so why do clients still go to the old address?" is almost always somewhere in that order.
Why you need this
You moved a service to a new server and changed the A record. Your laptop returns the new address, but half your customers are still connected to the old server. Whoever owns the change says "propagation takes time," but that is resignation, not an explanation. DNS has no such thing as propagation. All there is is caches and TTLs, and each cache returns the answer it stored until its TTL runs out. So you can answer "when will it be over?" exactly: it is over once the TTL of the record before the change has elapsed.
There is an accident in the opposite direction too. If someone looks up a name before you create it, the "does not exist" answer is cached as well. That is why NXDOMAIN keeps coming back for a while even after you add the record, and that time is set not by the A record's TTL but by the zone's SOA. If you do not know these two things, all you can do in the face of a DNS incident is wait.
How it works
When a program calls getaddrinfo(), glibc first looks at the hosts: line of /etc/nsswitch.conf. The order is usually files dns, so if the name appears in /etc/hosts, DNS is never asked at all. Because getent hosts <이름> follows exactly this order, getent is more accurate than dig for checking "the answer the program sees" (replace the placeholder with the name you want to resolve). dig is a tool that asks a DNS server directly, without going through the resolver library.
Once resolution moves on to DNS, /etc/resolv.conf sets the rules. You can list up to MAXNS (currently 3) nameserver entries, and they are asked in the order written. The search list and options ndots:n (default 1) decide how a short name is expanded: if the number of dots in the name is less than ndots, the search domains are appended one at a time and tried first, and only if they all fail is the original name queried. As described in the official documentation, a Kubernetes Pod's resolv.conf has search <ns>.svc.cluster.local svc.cluster.local cluster.local and options ndots:5. So when a Pod looks up an external name with fewer than five dots, such as api.example.com, the queries with search domains appended go out first, and only after three NXDOMAIN answers does it ask for the real name. That is why ndots is the first suspect when an external API is unusually slow from a Pod.
The resolver (cache) follows RFC 1034: it recursively walks to the authoritative servers to get an answer, and stores that answer for the number of TTL seconds attached to each record. The TTL is set by the zone on the authoritative server, and the cache returns the answer with the remaining time counting down. If you query the same cache twice with dig, the TTL in the second answer is lower, and that is the proof.
$ dig @127.0.0.1 -p 5301 www.lab.internal A +noall +answer
www.lab.internal. 120 IN A 10.0.0.10
$ dig @127.0.0.1 -p 5301 www.lab.internal A +noall +answer # 3초 뒤
www.lab.internal. 117 IN A 10.0.0.10
A CNAME says "this name is an alias of that name." When a resolver meets a CNAME, it follows the chain to the end and puts the CNAME and the final A in a single answer. The longer the chain, the more records there are to cache, and each one has its own TTL running separately.
How is a nonexistent name cached? When an authoritative server returns NXDOMAIN (the name does not exist) or NODATA (the name exists but has no record of that type), RFC 2308 says to put the zone's SOA in the authority section and set its TTL to the smaller of the SOA's MINIMUM field and the SOA record's own TTL. The resolver remembers "does not exist" for that long. If you tried a lookup before adding a record to the zone, you keep getting the "does not exist" answer for that long even after the record is added.
$ dig @127.0.0.1 -p 5301 nope.lab.internal A +noall +comments +authority
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, ...
;; AUTHORITY SECTION:
lab.internal. 60 IN SOA ns1.lab.internal. admin.lab.internal. 2026091101 3600 600 86400 60
What it looks like in the field
When planning a migration, an experienced operator does one thing: lower the TTL a few days before the change. If you change a record with a TTL of 3600, in the worst case two servers share the traffic for an hour. If you lower it to 60 ahead of time (that change itself takes as long as the old TTL), the actual migration finishes within a minute. Do not forget to restore the TTL afterward. A low TTL reduces the cache hit rate and increases the load on the authoritative server.
On Kubernetes, the extra queries created by ndots:5 show up as CoreDNS load and latency. The standard fixes are to write external names as absolute names with a trailing dot (api.example.com.) or to lower ndots with the Pod's dnsConfig. /etc/hosts can be a trap as well. If a line you added while debugging long ago is still there, only that machine goes somewhere else — so when you get a report that "it is different only on this machine," you check getent hosts first.
What you will do in the next lab
You start an authoritative server and a cache server in Python inside a Pod. You put an A record with a TTL of 120 and an SOA (minimum 60) in the zone file, and with dig you read the TTL counting down, the CNAME chain, and the SOA TTL of an NXDOMAIN answer. Then you edit the zone to reproduce a state where the authoritative server gives the new answer but the cache gives the old one, flush the cache, calculate the worst-case propagation time, and write it in the report. Finally, with +search +ndots=5, you check in the server log which queries a short name generates.