CKA — Kubernetes Administrator
CoreDNS — How a Name Becomes an IP
Summary
When a Pod asks for a name, the query goes to the nameserver in the /etc/resolv.conf that the kubelet wrote for it, which is the ClusterIP of the kube-dns Service. Behind that is a CoreDNS Pod, and what CoreDNS answers and how is decided by the Corefile in the coredns ConfigMap. When resolution is blocked, you check this chain in order from the front.
Why this matters
A Service's ClusterIP changes if you delete and recreate it. A Pod IP changes even more often. If an application hard-codes an IP in its configuration, it breaks at every redeploy, so Kubernetes made the name the stable address and put the path from a name to an IP inside the cluster. That path is cluster DNS, and its current implementation is CoreDNS.
The problem is that this path is made of several pieces. The Pod's resolv.conf, the kube-dns Service, the EndpointSlice, the CoreDNS Pod, the Corefile, the upstream DNS — if any one piece is off, you get the same symptom: "the name can't be found." With the same symptom, trying to guess the cause wastes all your time. This is why the official document Debugging DNS Resolution fixes an order.
How it works
The Pod side — resolv.conf and the search order
According to the document DNS for Services and Pods, the kubelet writes a resolv.conf for each Pod. It looks like this.
nameserver 10.32.0.10
search <namespace>.svc.cluster.local svc.cluster.local cluster.local
options ndots:5
The search line decides the name resolution order. When a Pod in the test namespace asks for data, it first tries expanding it to data.test.svc.cluster.local. So a Service in the same namespace is found by its short name, while a Service in another namespace (prod) needs the namespace attached, like data.prod. As the document puts it, DNS queries that do not specify a namespace are limited to the Pod's namespace. ndots:5 means that if a name has fewer than five dots, the search domains are tried first, so finding one outside domain costs knocking on three cluster domains first.
Which nameserver to use is decided by the dnsPolicy in the Pod spec. The document separately notes that Default is not the default value, which shows how confusing this part is.
| dnsPolicy | Behavior |
|---|---|
ClusterFirst |
The default when not specified. Queries outside the cluster domain are forwarded upstream |
ClusterFirstWithHostNet |
A hostNetwork: true Pod must state this explicitly to use cluster DNS |
Default |
Inherits the node's name resolution configuration as is |
None |
Ignores the Kubernetes configuration and specifies everything directly with dnsConfig |
The server side — the kube-dns Service and the Corefile
The DNS server is the CoreDNS Pods in the kube-system namespace, and in front of them is a ClusterIP Service named kube-dns (53/UDP, 53/TCP). The reason the name is not coredns but kube-dns is backward compatibility with the original kube-dns, and the Pod label k8s-app=kube-dns remains for the same reason. When debugging, you find the Pods by this label.
CoreDNS is a server that combines plugins, and its configuration file is the Corefile. The default Corefile shown in the document Customizing DNS Service looks like this.
.:53 {
errors
health {
lameduck 5s
}
ready
kubernetes cluster.local in-addr.arpa ip6.arpa {
pods insecure
fallthrough in-addr.arpa ip6.arpa
ttl 30
}
prometheus :9153
forward . /etc/resolv.conf
cache 30
loop
reload
loadbalance
}
If you know the meaning of each line, logs and symptoms connect. errors records errors to stdout, health reports status on port 8080, and lameduck 5s marks it unhealthy for 5 seconds before shutdown to drain traffic. ready returns 200 on port 8181 when all plugins are ready. The kubernetes plugin answers queries for the cluster domain based on Service and Pod IPs, and ttl 30 is the response TTL (per the document, default 5 seconds and maximum 3600 seconds). forward . /etc/resolv.conf passes queries outside the cluster domain to the upstream listed in the CoreDNS Pod's resolv.conf. cache 30 is the response cache, loop is a safety mechanism that stops the process when it detects a forwarding loop, reload is a plugin that automatically rereads the Corefile when it changes, and loadbalance is a round robin that shuffles the order of A, AAAA, and MX records. After you edit the ConfigMap, the document says to allow about 2 minutes for it to take effect.
A stub domain, which sends only specific domains to a different server, also goes in the same file. As in the document's example, in a server block for consul.local:53, if you put forward . 10.150.0.1, only names ending in .consul.local go to that server.
The debugging procedure — from the front
Here are the official steps in order.
- Bring up a
dnsutilsPod (the document'sregistry.k8s.io/e2e-test-images/agnhostimage) and runnslookup kubernetes.default. If an answer comes back, DNS is healthy. - If it fails, first look at that Pod's
/etc/resolv.conf. Check that the search path and nameserver look like the form above. - In
kube-system, use-l k8s-app=kube-dnsto check that the CoreDNS Pods are Running. If there are none, the add-on was not deployed. - Look at the logs with the same label. A healthy log prints
plugin/reload: Running configuration. - Check that the
kube-dnsService exists, and that the EndpointSlice with the labelkubernetes.io/service-name=kube-dnshas addresses. If the endpoints are empty, the Service exists but there are no Pods behind it. - To see whether queries even reach CoreDNS, put the
logplugin in the Corefile and check whether query lines appear in the log. - If the log has
SERVFAIL, suspect permissions. CoreDNS must be able to list and watch Services and EndpointSlices, and if thesystem:corednsClusterRole is missing thediscovery.k8s.iogroup'sendpointslices, this error occurs. - Finally, check the namespace. A Service in another namespace must be queried as
<service>.<namespace>(service name and namespace).
The "known issues" in the document are also worth remembering. Distributions that use systemd-resolved, such as Ubuntu, replace /etc/resolv.conf with a stub file, which can create a forwarding loop, and in that case you must point the kubelet's --resolv-conf at /run/systemd/resolve/resolv.conf (kubeadm detects this automatically). Also, glibc reads only up to 3 nameservers, and Kubernetes uses one of them.
What it looks like in the field
The logs are full of SERVFAIL but the Pods are fine. This is exactly the example the document gives — if the Pods are Running and the Service exists but the response is SERVFAIL, it is often that CoreDNS has no permission to read EndpointSlices. Use kubectl describe clusterrole system:coredns to see whether list and watch exist on endpointslices.discovery.k8s.io, and if not, fix the ClusterRole.
Only hostNetwork Pods cannot find Service names. For a Pod that uses the node network, if you do not specify dnsPolicy, ClusterFirst behaves like Default and uses the node's resolv.conf. This is why the document tells you to state ClusterFirstWithHostNet explicitly, and it is often forgotten when deploying log collectors and node agents.
What to check in the next quiz
The quiz asks about the relationship between base and overlay, the two patch methods, why kubectl top fails and which endpoint metrics-server reads, the default value of dnsPolicy, and what cause a SERVFAIL log points to. Rather than memorizing commands, you just need to be able to answer "which piece is missing."