TT Lab
Get started
Learn Learning paths Courses

Network Troubleshooting

What /etc/hosts Is and When It Wins

Continue in TT Lab

In one line

/etc/hosts is a local name tag consulted before DNS, and dig does not look at that file at all. This one sentence is the exit from countless mazes.

Why this exists

A report like this comes in. "DNS seems normal, but only the application connects to the old server."

When you check, dig api.internal really does return the new IP. Yet the application keeps going to the old IP. If you start digging through the DNS server here, a day goes by.

The culprit is usually /etc/hosts. A line someone put in temporarily while debugging last week is still there. And this incident does not show the face of a DNS error. Because the name resolves fine. The symptom appears as connection refused or a timeout, so nobody suspects DNS.

How it works

First, the file format. One line is IP + canonical name + aliases.

127.0.0.1       localhost
127.0.1.1       myserver
::1             localhost ip6-localhost ip6-loopback
10.0.1.50       api.internal.mycompany.com api-internal
192.168.1.100   db-master.mycompany.com

Next is the file that decides the order, /etc/nsswitch.conf.

grep '^hosts' /etc/nsswitch.conf
# hosts: files dns myhostname
# hosts: files mdns4_minimal [NOTFOUND=return] dns myhostname
# hosts: files resolve [!UNAVAIL=return] dns myhostname

The tokens mean this.

Token What it looks at
files /etc/hosts
dns Queries the nameserver in /etc/resolv.conf directly
resolve Queries systemd-resolved over D-Bus
myhostname Its own host name and local interface addresses
mdns4_minimal Resolves .local names with mDNS

What is in square brackets is an action rule. [NOTFOUND=return] means that if the preceding module answers "no such name", stop there. It does not go on to the dns that comes after.

That is why organizations that named their internal domain .local fall into this trap. mdns4_minimal [NOTFOUND=return] intercepts .local names, and if it fails, it ends right there, so it does not pass on to DNS. But dig does not look at nsswitch, so it returns a normal response. dig works but only the application fails.

This is where the most important sentence in this course comes from.

dig shows the DNS server's answer, and getent shows the answer the application will receive. If the two results differ, that difference is the location of the cause.

dig, nslookup, and host are all tools dedicated to the DNS protocol. They do not look at /etc/hosts, nsswitch, or systemd-resolved's per-link settings. They just take the nameserver address from resolv.conf and throw a query over UDP 53. The only tools that reproduce the application path are getent hosts / getent ahosts / resolvectl query.

What it looks like in the field

Making a temporary workaround in hosts and not removing it. During incident response, "let's just pin it in hosts and move on" is a common judgment and usually right. The problem is that the line is still there three months later. That is why many teams have a rule that whenever you touch hosts, you must write the date, the reason, and the person responsible in a comment.

hosts baked into a container image. An entry put in at build time remains in the image and points to the same IP in every environment. This is how incidents where staging connects to the production DB happen.

What you will do in the next lab

You put an entry directly into /etc/hosts and check with getent, and then audit a broken hosts file fixture to find the problem lines. And you submit a final version that satisfies the rules.