It's Always DNS: A Practical Debugging Guide for the Usual Suspect

There is a particular flavour of outage that arrives with no obvious cause and no obvious culprit. The service was fine an hour ago. Nothing deployed. The error is a timeout, or a connection refused, or — most maddening of all — it works on your laptop. Someone on the call says the thing everyone is already thinking, half as a joke and half as a prayer: it's probably DNS.
It usually is. Not because DNS is badly engineered — it is, on the whole, a remarkably durable forty-year-old protocol — but because name resolution sits underneath everything and is touched by almost nothing you control directly. A typical lookup passes through a stub resolver, a name service switch, a local caching daemon, a recursive resolver somewhere on your network, and finally an authoritative server you have never heard of. Each of those layers caches. Each can be configured to lie to you, usually for a good reason. When one disagrees with the others, you get a failure that is intermittent, environment-specific, and invisible to the application logs.
The meme persists not because engineers are lazy, but because the resolution path is genuinely hard to see, and most people debug it by poking one end and hoping. The cure is to stop guessing and walk the path deliberately, layer by layer, asking each one what it actually thinks the answer is.
Walk the path, don't poke it
The single most useful habit here is to know which tool talks to which layer. This trips people up constantly, because the tools look interchangeable and are not.
dig and drill are DNS clients. They speak the protocol directly to a resolver of your choosing and show you the raw answer — record types, TTLs, authority sections, the lot. They are the right tool for asking "what does this server return for this name?". They are also, crucially, not what your application uses. dig does not read /etc/hosts, it does not consult nsswitch.conf, and it does not respect your search domains the way the C library does. A name can resolve perfectly under dig and still fail in your service, because they took different routes.
getent hosts is the tool that closes that gap. It goes through the name service switch — the same nsswitch.conf path glibc itself uses — so it consults /etc/hosts, mDNS, and whatever else is wired into the hosts: line before it touches DNS at all. If dig says yes and getent says no, your problem is not in DNS; it is in NSS, or /etc/hosts, or the order in which they are consulted. That single comparison eliminates half the search space.
nslookup is the one to be wary of. Its output format is a relic and it has its own resolution logic that matches neither the C library nor modern dig — fine for a quick sanity check, actively misleading for anything subtle. Reach for dig or drill and getent; leave nslookup to the documentation written in 2008.
A first pass on any suspected resolution failure looks roughly like this:
# What does the application's path actually return?
getent hosts api.internal.example.com
# What does DNS return, bypassing NSS and /etc/hosts?
dig +short api.internal.example.com
# Which server am I even asking, and is it answering?
dig api.internal.example.com # check the SERVER: line at the bottom
# Walk the delegation from the root, ignoring every cache in between.
dig +trace api.internal.example.com
# Is the negative answer cached, and for how long?
dig api.internal.example.com # read the SOA in the AUTHORITY section
The +trace is worth dwelling on. It starts at the root servers and follows the delegation chain down to the authoritative server, ignoring your local caches entirely. When the cached answer and the +trace answer disagree, you have found your problem: something between you and the authority is serving a stale or wrong record. That is the difference between "the data is wrong" and "the data is right but somebody is caching the wrong thing" — the single most valuable distinction in DNS debugging.
Where stale answers come from
Almost every confusing DNS failure is, underneath, a caching failure. A record changed, the change is correct, and some layer between you and the truth has not noticed yet. TTLs exist precisely to bound how long that lag can be — and they are routinely ignored, clamped, or misunderstood.
A TTL is a suggestion. The authoritative server says "you may cache this for 300 seconds", and every resolver downstream is free to honour that, cap it lower, extend it higher to save itself work, or serve it stale because the upstream went away. systemd-resolved, dnsmasq, corporate forwarders, and your cloud provider's resolver all have opinions, and they do not have to agree. You can lower a record's TTL to five seconds an hour before a migration and still be served the old value, because something in the chain decided five seconds was inconveniently short.
Negative answers cache too, which is the part people forget. Under RFC 2308, an NXDOMAIN or empty response is cached for a duration controlled by the minimum of the SOA record's TTL and its minimum field — not by anything on the record you were looking up, because there is no record. So you create the DNS entry, you confirm it is live at the authority, and your application keeps insisting the name does not exist for the next fifteen minutes. It is not lying. It cached the "no" before you created the "yes", and the "no" has not expired. When a freshly created name stubbornly refuses to resolve, negative caching is the first thing to suspect, and dig showing the SOA in the authority section is how you find the duration you are waiting out.
Here is the symptom-to-cause mapping I reach for most often:
| Symptom | Likely DNS cause |
|---|---|
Works under dig, fails in the app |
NSS / /etc/hosts / nsswitch.conf, not DNS |
| New record won't resolve for minutes | Negative (NXDOMAIN) caching — RFC 2308 |
| Old IP after a migration | Stale positive cache; TTL longer than you think |
| Resolves on one host, not another | Split-horizon, or different resolv.conf / search domains |
| Intermittent, ~half of lookups fail | One resolver in a pair is broken or out of sync |
| Slow, then works | Trying a dead resolver first, falling back on timeout |
| Fully-qualified works, short name doesn't | search domains or ndots, especially in Kubernetes |
The resolver everyone forgets to check
On most modern Linux systems /etc/resolv.conf is no longer a file you edit; it is a symlink to something systemd-resolved manages, and the 127.0.0.53 stub address you see there is a local daemon, not a real upstream. This catches people out constantly. You point dig @8.8.8.8 at a public resolver to test, it works, and you conclude DNS is fine — but the application is going through 127.0.0.53, which is forwarding to a per-link resolver pushed by DHCP or VPN, which is the one actually misbehaving. resolvectl status and resolvectl query talk to the daemon directly; querying 8.8.8.8 tells you nothing about the path your app takes.
Search domains and ndots are the other silent saboteur, and they bite hardest in Kubernetes. The default ndots:5 means the resolver will try appending your search domains to any name with fewer than five dots before it tries the name as-is. So api.example.com — three dots — gets tried first as api.example.com.svc.cluster.local, then api.example.com.cluster.local, and so on down the list, generating a fistful of NXDOMAINs and the latency to match before it finally queries the name you actually meant. Every external lookup from inside a pod pays this tax. The fix is usually a trailing dot to force the fully-qualified form, or a tuned ndots in the pod's dnsConfig, but the first step is simply knowing it is happening.
One more difference worth carrying, because it produces failures that look like sorcery: musl and glibc resolve differently. Alpine ships musl, and musl's resolver queries all nameservers in resolv.conf in parallel and historically handled search domains and EDNS differently from glibc. A name that resolves cleanly on your Debian host can fail, or resolve more slowly, inside an Alpine container on the same network — same resolv.conf, different resolver implementation reading it. If your debugging has reached the point of disbelief, check whether the failing thing is statically linked against musl before you check anything else.
None of this is exotic. It is the same protocol doing exactly what it was told, by half a dozen components that were each told something slightly different. The meme endures because the path is invisible by default — and the cure, every time, is to make it visible and ask each layer in turn. A path beats a panic.