DNS Troubleshooting
Diagnosing name resolution with dig, drill, host, and resolvectl: record types, CNAME chains, TTLs, authoritative vs recursive, DNSSEC, split-horizon, and common failure modes.
DNS Troubleshooting
Diagnosing name resolution from the operator's seat — querying, reading answers, and isolating failures with dig, drill, host, and resolvectl.
Overview
Almost every "the network is down" ticket is a DNS ticket. A name fails to resolve, resolves to the wrong place, or resolves correctly somewhere but not on the host that matters. The job is rarely to run a nameserver — it is to ask the right server the right question and read the answer carefully.
Resolution is a chain. A program calls the stub resolver (the C library, via /etc/nsswitch.conf), which forwards to a recursive resolver (your ISP, 1.1.1.1, 8.8.8.8, or a local systemd-resolved). The recursive resolver, if it has nothing cached, walks the delegation tree: root servers → TLD servers (.com) → the zone's authoritative servers, then caches the result and hands it back. The single most important distinction in troubleshooting is authoritative vs recursive: an authoritative server speaks for the zone and sets the aa flag; a recursive resolver merely relays (and may be serving you a stale or poisoned cache).
sequenceDiagram
participant App as App / stub resolver
participant Rec as Recursive resolver<br/>(1.1.1.1 / resolved)
participant Root as Root (.)
participant TLD as TLD (.com)
participant Auth as Authoritative<br/>(ns.example.com)
App->>Rec: A? example.com (rd=1)
Note over Rec: cache miss
Rec->>Root: A? example.com
Root-->>Rec: referral to .com servers
Rec->>TLD: A? example.com
TLD-->>Rec: referral to example.com NS
Rec->>Auth: A? example.com
Auth-->>Rec: ANSWER 93.184.x.x (aa=1)
Rec-->>App: ANSWER (cached, aa=0, ra=1)
When something breaks, the diagnostic instinct is to move along this chain: ask the recursive resolver, then ask the authoritative server directly (+norecurse), then walk the delegation yourself (+trace). Where the answer changes tells you where the fault is.
The Mental Model — Authoritative vs Recursive
| Authoritative server | Recursive resolver | |
|---|---|---|
| Role | Holds the zone's records | Looks answers up on your behalf |
| Examples | ns1.example.com, Route 53, Cloudflare NS |
1.1.1.1, 8.8.8.8, 127.0.0.53 |
Sets aa flag |
Yes (for its own zones) | No |
Honours rd (recursion desired) |
Ignores it / refuses | Yes |
| Answer freshness | Authoritative truth, TTL fresh | May be cached (stale possible) |
| Query it with | dig +norecurse @ns … |
dig @1.1.1.1 … |
Reading the flags line tells you which kind you reached. A reply with aa is from the horse's mouth; a reply with ra (recursion available) but no aa came from a cache. If you ask a recursive resolver and the answer looks wrong, ask the authoritative server — if that is also wrong, the zone is wrong; if only the resolver is wrong, you have a caching or split-horizon problem.
dig — The Primary Tool
dig (Domain Information Groper, from BIND's bind-utils/bind9-dnsutils) is the reference DNS client. Learn it first; everything else is a convenience wrapper.
Anatomy of the Output
; <<>> DiG 9.18 <<>> example.com A
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 53096
;; flags: qr rd ra ad; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
;; QUESTION SECTION:
;example.com. IN A
;; ANSWER SECTION:
example.com. 300 IN A 104.20.23.154
example.com. 300 IN A 172.66.147.243
;; AUTHORITY SECTION: ← which servers are authoritative (on referrals/NODATA)
;; ADDITIONAL SECTION: ← glue / OPT pseudo-section (EDNS)
;; Query time: 23 msec
;; SERVER: 8.8.8.8#53(8.8.8.8) ← which resolver answered
;; WHEN: Mon Jun 15 21:23:27 BST 2026
;; MSG SIZE rcvd: 72
Sections: QUESTION echoes what you asked; ANSWER is the data; AUTHORITY lists the NS for the zone (and carries the SOA on a negative answer); ADDITIONAL carries glue records and the EDNS OPT pseudo-section.
Status codes (rcode) — the first thing to read:
| Status | Meaning | Where to look |
|---|---|---|
NOERROR |
Query succeeded (may still be empty — see NODATA) | Check the ANSWER count |
NXDOMAIN |
The name does not exist | Typo, wrong zone, record never created, search-domain suffixing |
SERVFAIL |
Resolver failed to get an answer | DNSSEC validation failure, broken upstream, unreachable authoritative |
REFUSED |
Server declined to answer | Not authoritative + recursion disabled, ACL blocks you, wrong server |
NOERROR with ANSWER: 0 is NODATA — the name exists but has no record of that type (e.g. an A query for a name that only has AAAA). It is not the same as NXDOMAIN.
Flags — the second thing to read:
| Flag | Name | Reading |
|---|---|---|
qr |
Query Response | This is a response (always set in replies) |
aa |
Authoritative Answer | Came from an authoritative server for the zone |
rd |
Recursion Desired | You asked for recursion (set in the query) |
ra |
Recursion Available | The server offers recursion (it is a resolver) |
ad |
Authenticated Data | The resolver DNSSEC-validated the answer |
tc |
Truncated | Answer exceeded UDP size; retry over TCP |
Everyday Invocations
# Just the addresses — scriptable
dig +short example.com
# 104.20.23.154
# The answer section only, no preamble or stats — the daily driver
dig +noall +answer example.com
# example.com. 300 IN A 104.20.23.154
# Target a specific resolver (compare answers across resolvers)
dig @1.1.1.1 example.com
dig @8.8.8.8 example.com
# A specific record type
dig -t MX example.com # or: dig example.com MX
dig -t AAAA example.com
dig -t TXT example.com
# Reverse lookup (PTR) — -x builds the in-addr.arpa name for you
dig -x 8.8.8.8
dig +short -x 8.8.8.8
# dns.google.
# Force TCP (large answers, or when UDP is filtered)
dig +tcp example.com
# Multiple names / types in one shot
dig +noall +answer example.com A example.com AAAA example.com MX
Finding the Authoritative Servers
# Who is authoritative for the zone?
dig +short example.com NS
# The SOA — serial, refresh/retry/expire, and the negative-cache minimum
dig +noall +answer example.com SOA
# example.com. 3600 IN SOA elliott.ns.cloudflare.com. dns.cloudflare.com. \
# 2405749864 10000 2400 604800 1800
# serial refresh retry expire negTTL
The SOA's final field is the negative-caching TTL — how long resolvers cache an NXDOMAIN/NODATA for this zone. A high value here is why a just-created record can appear missing for an hour.
Asking the Authoritative Server Directly
This is the workhorse for "the resolver says X but is that actually true?". Point dig at one of the zone's nameservers and disable recursion so you get the raw, authoritative record:
# Ask the zone's own nameserver, no recursion — expect aa in the flags
dig +norecurse @ns.example.com example.com A
# Resulting flags: qr aa (note: aa set, ra absent)
If the authoritative answer differs from what your resolver returns, the resolver is serving stale or poisoned data — flush it. If the authoritative answer is itself wrong, the zone needs fixing.
Walking the Delegation — +trace
# Follow the delegation from the root downward, querying each level
dig +trace example.com
+trace starts at the root, follows referrals through the TLD to the authoritative servers, and prints each hop. It bypasses your recursive resolver entirely, so it is invaluable for delegation problems: a broken NS record, a lame delegation (parent points at a server that is not authoritative), or a missing glue record shows up as the walk stalling or branching wrong. (It needs outbound UDP/53 to the root and TLD servers — restrictive networks may block it.)
DNSSEC and Wire-Level Flags
# Request DNSSEC records — RRSIG appears alongside the data
dig +dnssec +noall +answer cloudflare.com A
# cloudflare.com. 300 IN A 104.16.132.229
# cloudflare.com. 300 IN RRSIG A 13 2 300 ... (signature)
# Find authoritative servers and query each for the SOA — quick zone audit
dig +nssearch example.com
Reading TTLs
The number between the name and the class is the TTL in seconds — how long this answer may be cached. Query the same name twice against a recursive resolver and watch it count down toward zero as the cache ages; ask the authoritative server and it resets to the full configured value. A stubbornly high TTL is the usual reason a change "hasn't propagated".
dig +noall +answer +ttlunits example.com # render TTL as 5m, 1h, etc.
drill, host, and nslookup
drill — the DNSSEC-friendly alternative
drill ships with ldns (ldnsutils / ldns) and is the tool of choice when chasing DNSSEC chains — its output is terser and its trace mode validates signatures.
drill example.com # basic lookup
drill -x 8.8.8.8 # reverse
drill example.com MX # by type
drill -D example.com # request + show DNSSEC records
drill -T example.com # trace AND validate the chain of trust (the killer feature)
drill -S example.com # chase and verify the full signature chain
host — quick one-liners
host is the fastest way to get a human-readable answer. Great for a sanity check; thin on detail.
host example.com # A, AAAA, and MX in plain English
host -t MX example.com # specific type
host -a example.com # "all" — verbose, ANY-style dump
host 8.8.8.8 # reverse lookup (PTR)
nslookup — legacy, avoid
nslookup still exists everywhere and you will see it in old runbooks, but prefer dig. Its interactive mode is clumsy, it hides the status/flags that matter, and its behaviour around search domains and authoritative-vs-recursive answers is muddier. Use it only when nothing else is installed:
nslookup example.com # basic
nslookup example.com 1.1.1.1 # against a specific server
nslookup -type=mx example.com # by type
If you reach for nslookup reflexively, retrain the muscle — dig +short is shorter and clearer.
systemd-resolved and the Host's View
Modern Linux desktops and many servers run systemd-resolved, which changes where "the resolver" actually lives. This is the single biggest source of "but dig works and the app doesn't" confusion.
The 127.0.0.53 Stub
/etc/resolv.conf on a resolved host usually points at the local stub:
nameserver 127.0.0.53
options edns0 trust-ad
So dig (which reads resolv.conf) talks to the stub on 127.0.0.53, not to the real upstream. The real upstreams and per-link config live in resolved, inspected with resolvectl:
# Global + per-link resolver config: upstream DNS, search domains, DNSSEC mode
resolvectl status
# Resolve a name THROUGH resolved (honours per-link DNS, split-DNS routing)
resolvectl query example.com
# Cache statistics (hits, misses, current size)
resolvectl statistics
# Flush the resolved cache (does NOT restart the service)
resolvectl flush-caches
# Which search/routing domains apply, per link
resolvectl domain
Per-link DNS matters on multi-homed and VPN hosts: resolved can send corp.example.com queries down the VPN link and everything else to the public resolver. resolvectl status shows this routing; resolvectl query obeys it, whereas a bare dig (hitting the stub or a hard-coded @server) may not.
getent vs dig — the High-Value Distinction
dig and friends speak DNS only. Applications do not — they call the C library's resolver, which obeys /etc/nsswitch.conf:
# /etc/nsswitch.conf
hosts: files mdns4_minimal [NOTFOUND=return] dns
That hosts: line means a name is resolved by trying, in order: /etc/hosts (files), multicast DNS (mdns), then dns. So:
# What the application layer actually sees (files + mdns + dns, in nsswitch order)
getent hosts example.com
getent ahosts example.com # all address families, with socket types
When getent hosts foo and dig foo disagree, the answer is almost always one of:
- an entry in
/etc/hostsshadowing DNS (filescomes first), - mDNS (
.local) answering instead of DNS, - nsswitch ordering doing something you did not expect.
dig cannot see any of that. Make getent hosts your first command when an app misbehaves but dig looks fine — it reproduces what the app does. This single check resolves a large fraction of "DNS works but the service can't connect" incidents.
Record Types Reference
| Type | Holds | Notes for troubleshooting |
|---|---|---|
A |
IPv4 address | The common case |
AAAA |
IPv6 address | App may prefer AAAA; missing AAAA causes "works on IPv4 only" |
CNAME |
Alias to another name | Resolver follows the chain; illegal at a zone apex |
MX |
Mail exchanger (priority + host) | 0 . (null MX, RFC 7505) means "this domain sends/receives no mail" |
TXT |
Free text | Home of SPF (v=spf1 …), DKIM (…._domainkey), DMARC (_dmarc) |
NS |
Authoritative nameservers | Delegation; mismatch parent↔child = lame delegation |
SOA |
Zone apex metadata | Serial, refresh/retry/expire, and negative-cache TTL |
PTR |
Reverse (IP → name) | Lives under in-addr.arpa / ip6.arpa; query with dig -x |
SRV |
Service location (host + port + priority/weight) | _service._proto.name, e.g. _sip._tcp |
CAA |
Which CAs may issue certs | ACME/Let's Encrypt checks this; wrong CAA blocks issuance |
ALIAS / ANAME |
Provider-specific apex "CNAME" | Not a real RR — DNS host flattens it to A/AAAA at query time |
SPF/DKIM/DMARC are all TXT records; debug mail auth with dig +short TXT example.com, dig +short TXT _dmarc.example.com, and dig +short TXT selector._domainkey.example.com.
CNAME Chains and Flattening
A CNAME says "this name is really that name". Resolvers follow the chain and return both the alias and the final address:
dig +noall +answer www.github.com
# www.github.com. 3600 IN CNAME github.com.
# github.com. 60 IN A 140.82.121.3 # apex A; low TTL, value rotates
The Apex Gotcha
A CNAME cannot coexist with any other record at the same name, and a zone apex (example.com itself) must carry SOA and NS records. Therefore CNAME at the apex is illegal — you cannot CNAME example.com → myapp.cloud.net. Symptoms: the zone fails to load, or the provider rejects the record. The fixes:
- Use the DNS host's
ALIAS/ANAMEpseudo-record, which flattens to A/AAAA at query time. - Or use a provider feature like Cloudflare's CNAME flattening.
- Or point the apex at static IPs with
A/AAAAand CNAME only thewwwsubdomain.
Chains Too Long or Looping
Chains should be short. Each hop is another lookup (latency), and a self-referential or mutually-referential pair (a → b → a) is a loop. Most resolvers cap the chain length and return SERVFAIL rather than spinning. Spot it by reading the answer section: if you see many CNAME hops or the chain never reaches an A/AAAA, that is the bug.
TTLs and Caching
The TTL governs how long every cache between the authoritative server and the application may keep an answer. This is why a DNS change is not instant.
"It Hasn't Propagated"
DNS does not push; it expires. After you change a record, old answers persist in caches until their TTL elapses. Check what the authoritative server now serves versus what a public resolver still has:
# Truth, fresh from the source
dig +norecurse @ns.example.com example.com A
# What a public cache is still handing out (watch the TTL count down)
dig @1.1.1.1 example.com A
dig @8.8.8.8 example.com A
If the authoritative answer is new but resolvers still return the old one, you are simply waiting out the TTL. There is no "force propagation" — only patience or a cache flush on resolvers you control.
Lower the TTL Before a Migration
The standard play: at least one full TTL before a cutover, drop the record's TTL to 60–300s. Wait for the old (high) TTL to expire everywhere, do the migration on the low TTL so rollback is fast, then raise it again afterwards. Plan changes around the old TTL, not the new one.
Negative Caching
NXDOMAIN and NODATA answers are cached too, for the duration of the zone's SOA minimum (last SOA field). Create a record that "should exist" and still get NXDOMAIN? A resolver may be holding a negative answer from before the record existed. Check the SOA minimum and wait it out, or flush.
Flushing OS and Browser Caches
# Linux — systemd-resolved
resolvectl flush-caches
# Linux — nscd (if running)
sudo nscd -i hosts
# macOS — flush the system resolver cache
sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder
Browsers keep their own in-process DNS caches (and may use DNS-over-HTTPS, bypassing the OS entirely): Chrome at chrome://net-internals/#dns, Firefox via a restart or about:networking#dns. If dig is correct but only the browser is wrong, suspect the browser cache or its DoH setting.
DNSSEC
DNSSEC signs records so a validating resolver can prove the answer was not tampered with. As an operator you mostly read its signals rather than manage keys.
The ad Flag and RRSIG
# A signed zone, validated by the resolver → ad in the flags
dig cloudflare.com A | grep flags
# ;; flags: qr rd ra ad; ...
# An unsigned zone → no ad flag (this is normal, not an error)
dig github.com A | grep flags
# ;; flags: qr rd ra; ...
# See the signatures themselves
dig +dnssec +noall +answer cloudflare.com A
# ... A 104.16.132.229
# ... RRSIG A 13 2 300 20260616212320 ... (signer, validity window, signature)
ad present means the resolver fetched the signatures and verified them up the chain of trust. ad absent on a zone that is not signed is expected — most zones are unsigned.
The Classic Symptom — SERVFAIL on Validation Failure
When a signed zone's signatures are broken (expired RRSIG, missing DS at the parent, key rollover botched), a validating resolver returns SERVFAIL — not a wrong answer. So if a name gives SERVFAIL from 8.8.8.8/1.1.1.1 (both validate) but resolves fine from a non-validating resolver, suspect DNSSEC:
# Validating resolver — SERVFAIL if DNSSEC is broken
dig dnssec-failed.org @8.8.8.8 | grep status
# ;; ->>HEADER<<- ... status: SERVFAIL
# Ask without validation (+cd = checking disabled) — if THIS succeeds, it's DNSSEC
dig +cd dnssec-failed.org @8.8.8.8 | grep status
# ;; ->>HEADER<<- ... status: NOERROR
+cd (checking disabled) tells the resolver to skip validation. SERVFAIL that clears under +cd is a DNSSEC validation failure, full stop.
Validating Locally — delv
delv (ships with BIND tools) performs full validation client-side and tells you the verdict in plain language:
delv cloudflare.com A
# ; fully validated
# cloudflare.com. 300 IN A 104.16.132.229
# cloudflare.com. 300 IN RRSIG A ...
delv dnssec-failed.org A
# ;; resolution failed: failure
# Current BIND (9.20+) prints the terse "failure"; older versions said
# "no valid signature found". Either way: validation failed, chain of trust broken.
drill -T / drill -S do the same from the ldns side. Use whichever is installed to confirm whether the break is signing-side (zone) or just your resolver.
Split-Horizon (Split-Brain) DNS
The same name resolves to different answers depending on who is asking — typically a private RFC 1918 address for clients inside the network or on the VPN, and a public address for everyone else. It is a deliberate design, and a frequent source of "works for me, not for them".
Detecting It
Compare an internal resolver against a public one:
# The internal / corporate resolver's answer
dig @10.0.0.53 app.corp.example.com
# A public resolver's answer for the same name
dig @8.8.8.8 app.corp.example.com
Different addresses (or NXDOMAIN from the public side) confirm split-horizon. On a host using systemd-resolved, the deciding factor is which link's resolver the query is routed to:
resolvectl status # per-link DNS + which domains route where
resolvectl domain # routing/search domains per link
resolvectl query app.corp.example.com # resolve the way the system actually will
VPN and Search-Domain Effects
When a VPN comes up it usually pushes resolver and search-domain settings. Symptoms of getting this wrong: internal names fail because the VPN's resolver was not installed for that domain, or external names break because the VPN grabbed all DNS. resolvectl status shows whether the VPN link is DNS Domain ~corp.example.com (route only that domain) or has the catch-all ~. (route everything). The ~. catch-all on the wrong link is a classic "VPN broke my internet" cause.
Common Failure Modes
The fastest path to root cause is reading the status code and which server answered, then moving along the resolution chain.
NXDOMAIN vs SERVFAIL vs REFUSED vs Timeout
| Symptom | Likely cause | Next step |
|---|---|---|
NXDOMAIN |
Name genuinely absent; typo; wrong zone; search-domain suffixed a bad name | Check spelling and the trailing dot; dig +norecurse @authoritative; check negative cache |
SERVFAIL |
DNSSEC validation failure; authoritative unreachable; broken upstream | dig +cd (rules in/out DNSSEC); dig @authoritative directly; dig +trace |
REFUSED |
Server not authoritative and recursion off; ACL blocks your IP | You are asking the wrong server, or from a disallowed source — use a resolver that serves you |
| Timeout / "no servers could be reached" | Resolver unreachable; UDP/53 filtered; wrong nameserver in resolv.conf |
resolvectl status; dig @1.1.1.1 to bypass local config; check firewall, ss -u |
Resolver Not Reachable
# Is the configured resolver even responding? Bypass it with a known-good one.
dig @1.1.1.1 example.com
dig @8.8.8.8 example.com
# What does the host think its resolver is?
cat /etc/resolv.conf # often the 127.0.0.53 stub
resolvectl status # the real upstreams behind the stub
If @1.1.1.1 works but the default does not, the fault is in local resolver config (resolv.conf, resolved upstream, or a dead local cache), not in DNS itself.
Wrong Search Domain and ndots
/etc/resolv.conf carries search domains and an options ndots:N. A name with fewer than ndots dots is tried with each search domain appended first, before being tried as-is. Kubernetes sets ndots:5 by default, so dig api (and many in-cluster lookups) generate a flurry of failed suffixed queries before the bare name — a real latency and NXDOMAIN-noise source.
cat /etc/resolv.conf
# search corp.example.com
# options ndots:5
# 'api' becomes api.corp.example.com. first. To force the literal name,
# add the trailing dot — it means "fully qualified, do not append search":
dig api.example.com. # the trailing dot disables search suffixing
A surprising NXDOMAIN for a name you know exists is very often the search domain being appended (or not appended when you expected it). The trailing dot is the cure: it marks the name as fully qualified.
/etc/hosts Shadowing
getent hosts example.com # if this differs from dig, check /etc/hosts first
grep example.com /etc/hosts
Because files precedes dns in nsswitch.conf, a stray /etc/hosts line (often a leftover 127.0.0.1 someservice) silently overrides DNS for applications while dig — which never reads /etc/hosts — shows the "correct" DNS answer. This mismatch is exactly why getent is your first reproduction command.
EDNS, Truncation, and TCP Fallback
Large answers (many records, DNSSEC signatures) can exceed the UDP size and come back truncated (tc flag set), at which point the resolver retries over TCP. Middleboxes that block DNS-over-TCP, or break EDNS, cause intermittent failures on exactly the big responses.
# See truncation, then confirm TCP works
dig +dnssec example.com # look for 'tc' in the flags on a big answer
dig +tcp example.com # force TCP — if this works but UDP fails, suspect EDNS/MTU
dig +bufsize=512 example.com # shrink the advertised EDNS buffer to dodge fragmentation
A name that resolves with +tcp but fails or truncates over plain UDP points at an EDNS/MTU or middlebox problem on the path.
The Missing Trailing Dot
In zone files and many tools, a name without a trailing dot is relative and gets the origin appended; with the dot it is absolute. The classic bug: a CNAME target written www.example.com (no dot) in a zone for example.com silently becomes www.example.com.example.com.. From the operator side, the same logic explains surprise suffixing in lookups — when in doubt, append the dot and re-query.
Quick Reference
dig Flags and Options
| Option | Effect |
|---|---|
+short |
Answer data only, one value per line |
+noall +answer |
Just the ANSWER section, no preamble or stats |
@server |
Query a specific resolver/authoritative server |
-t TYPE |
Query a record type (A, AAAA, MX, TXT, NS, SOA, …) |
-x ADDR |
Reverse (PTR) lookup |
+trace |
Walk the delegation from the root downward |
+norecurse |
Do not request recursion (ask an authoritative server directly) |
+dnssec |
Request DNSSEC records (RRSIG, etc.) |
+cd |
Checking disabled — skip DNSSEC validation (diagnose SERVFAIL) |
+tcp |
Force the query over TCP |
+nssearch |
Find authoritative servers and query each for the SOA |
+ttlunits |
Render TTLs as 5m, 1h, etc. |
+bufsize=N |
Set the advertised EDNS UDP buffer size |
Record Types
| Type | Purpose |
|---|---|
A / AAAA |
IPv4 / IPv6 address |
CNAME |
Alias to another name (never at a zone apex) |
MX |
Mail exchanger (priority + host; 0 . = null MX) |
TXT |
Free text — SPF, DKIM, DMARC live here |
NS |
Authoritative nameservers (delegation) |
SOA |
Zone apex metadata + negative-cache TTL |
PTR |
Reverse mapping (IP → name) |
SRV |
Service location (host, port, priority, weight) |
CAA |
Which CAs may issue certs for the name |
ALIAS/ANAME |
Provider apex "CNAME" flattened to A/AAAA |
Resolver / Host Commands
| Task | Command |
|---|---|
| Resolve as the app does (NSS) | getent hosts NAME |
| Resolver config behind the stub | resolvectl status |
| Resolve through resolved | resolvectl query NAME |
| Flush resolved cache | resolvectl flush-caches |
| Flush macOS cache | sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder |
| Quick human-readable lookup | host NAME |
| Validate DNSSEC locally | delv NAME / drill -T NAME |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Change "hasn't propagated" | Old answer still within its TTL in caches | Compare dig +norecurse @authoritative vs dig @1.1.1.1; wait out the TTL; lower TTL before the next change |
| New record returns NXDOMAIN | Negative answer cached (SOA minimum) | Check the SOA minimum; flush resolvers you control; wait it out |
getent and dig disagree |
/etc/hosts or mDNS shadowing DNS via nsswitch ordering |
Inspect /etc/hosts and /etc/nsswitch.conf; dig cannot see either |
| SERVFAIL on a signed zone | DNSSEC validation failure (expired RRSIG, bad DS) | dig +cd to confirm; delv/drill -T to locate the break; fix signing or DS |
dig works, app doesn't |
App hits the stub/NSS, not your @server |
Reproduce with getent hosts and resolvectl query, not bare dig |
| Internal name fails on VPN | Per-link DNS / search-domain routing wrong | resolvectl status; check ~domain vs ~. catch-all on each link |
| Same name, different answers | Split-horizon DNS | dig @internal vs dig @8.8.8.8; expected by design — query the right resolver |
| Surprise NXDOMAIN for a real name | Search domain appended (or not) per ndots |
Append a trailing dot to force the literal name; check resolv.conf search/ndots |
| Big lookups fail, small ones fine | EDNS/truncation, TCP fallback blocked | dig +tcp and dig +bufsize=512; fix the middlebox/MTU on the path |
| REFUSED from a server | Not authoritative + recursion off, or ACL | Query a resolver that serves you, or the zone's authoritative NS |
CNAME rejected at apex |
CNAME illegal alongside SOA/NS at the zone root | Use ALIAS/ANAME/flattening, or A/AAAA at the apex |
Related Topics
The following topics complement this DNS troubleshooting cheatsheet:
- Linux Network Tools —
getent,resolvectl, andip route getfor the host-side view of resolution and routing - tcpdump / Wireshark — capture port-53 traffic to see the exact queries and answers on the wire (
udp port 53) - Nginx —
resolverdirectives, upstream name resolution, and DNS-driven service discovery for proxied backends - Cert-Manager (ACME) — DNS-01 challenges,
CAArecords, and propagation checks that gate certificate issuance - Consul — service discovery that answers DNS for
.consul, and the split-horizon patterns it introduces - Kubernetes Networking (CoreDNS) — in-cluster resolution, the
ndots:5default, andClusterFirstvsDefaultdnsPolicy