Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

DNS Troubleshooting

Diagnosing name resolution with dig, drill, host, and resolvectl: record types, CNAME chains, TTLs, authoritative vs recursive, DNSSEC, split-horizon, and common failure modes.

DNS Troubleshooting

Diagnosing name resolution from the operator's seat — querying, reading answers, and isolating failures with dig, drill, host, and resolvectl.

Overview

Almost every "the network is down" ticket is a DNS ticket. A name fails to resolve, resolves to the wrong place, or resolves correctly somewhere but not on the host that matters. The job is rarely to run a nameserver — it is to ask the right server the right question and read the answer carefully.

Resolution is a chain. A program calls the stub resolver (the C library, via /etc/nsswitch.conf), which forwards to a recursive resolver (your ISP, 1.1.1.1, 8.8.8.8, or a local systemd-resolved). The recursive resolver, if it has nothing cached, walks the delegation tree: root servers → TLD servers (.com) → the zone's authoritative servers, then caches the result and hands it back. The single most important distinction in troubleshooting is authoritative vs recursive: an authoritative server speaks for the zone and sets the aa flag; a recursive resolver merely relays (and may be serving you a stale or poisoned cache).

Authoritative(ns.example.com)TLD (.com)Root (.)Recursive resolver(1.1.1.1 / resolved)App / stub resolverAuthoritative(ns.example.com)TLD (.com)Root (.)Recursive resolver(1.1.1.1 / resolved)App / stub resolvercache missA? example.com (rd=1)A? example.comreferral to .com serversA? example.comreferral to example.com NSA? example.comANSWER 93.184.x.x (aa=1)ANSWER (cached, aa=0, ra=1)Authoritative(ns.example.com)TLD (.com)Root (.)Recursive resolver(1.1.1.1 / resolved)App / stub resolverAuthoritative(ns.example.com)TLD (.com)Root (.)Recursive resolver(1.1.1.1 / resolved)App / stub resolvercache missA? example.com (rd=1)A? example.comreferral to .com serversA? example.comreferral to example.com NSA? example.comANSWER 93.184.x.x (aa=1)ANSWER (cached, aa=0, ra=1)

When something breaks, the diagnostic instinct is to move along this chain: ask the recursive resolver, then ask the authoritative server directly (+norecurse), then walk the delegation yourself (+trace). Where the answer changes tells you where the fault is.

The Mental Model — Authoritative vs Recursive

Authoritative server Recursive resolver
Role Holds the zone's records Looks answers up on your behalf
Examples ns1.example.com, Route 53, Cloudflare NS 1.1.1.1, 8.8.8.8, 127.0.0.53
Sets aa flag Yes (for its own zones) No
Honours rd (recursion desired) Ignores it / refuses Yes
Answer freshness Authoritative truth, TTL fresh May be cached (stale possible)
Query it with dig +norecurse @ns … dig @1.1.1.1 …

Reading the flags line tells you which kind you reached. A reply with aa is from the horse's mouth; a reply with ra (recursion available) but no aa came from a cache. If you ask a recursive resolver and the answer looks wrong, ask the authoritative server — if that is also wrong, the zone is wrong; if only the resolver is wrong, you have a caching or split-horizon problem.

dig — The Primary Tool

dig (Domain Information Groper, from BIND's bind-utils/bind9-dnsutils) is the reference DNS client. Learn it first; everything else is a convenience wrapper.

Anatomy of the Output

; <<>> DiG 9.18 <<>> example.com A
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 53096
;; flags: qr rd ra ad; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1

;; QUESTION SECTION:
;example.com.            IN  A

;; ANSWER SECTION:
example.com.    300  IN  A  104.20.23.154
example.com.    300  IN  A  172.66.147.243

;; AUTHORITY SECTION:        ← which servers are authoritative (on referrals/NODATA)
;; ADDITIONAL SECTION:       ← glue / OPT pseudo-section (EDNS)

;; Query time: 23 msec
;; SERVER: 8.8.8.8#53(8.8.8.8)   ← which resolver answered
;; WHEN: Mon Jun 15 21:23:27 BST 2026
;; MSG SIZE  rcvd: 72

Sections: QUESTION echoes what you asked; ANSWER is the data; AUTHORITY lists the NS for the zone (and carries the SOA on a negative answer); ADDITIONAL carries glue records and the EDNS OPT pseudo-section.

Status codes (rcode) — the first thing to read:

Status Meaning Where to look
NOERROR Query succeeded (may still be empty — see NODATA) Check the ANSWER count
NXDOMAIN The name does not exist Typo, wrong zone, record never created, search-domain suffixing
SERVFAIL Resolver failed to get an answer DNSSEC validation failure, broken upstream, unreachable authoritative
REFUSED Server declined to answer Not authoritative + recursion disabled, ACL blocks you, wrong server

NOERROR with ANSWER: 0 is NODATA — the name exists but has no record of that type (e.g. an A query for a name that only has AAAA). It is not the same as NXDOMAIN.

Flags — the second thing to read:

Flag Name Reading
qr Query Response This is a response (always set in replies)
aa Authoritative Answer Came from an authoritative server for the zone
rd Recursion Desired You asked for recursion (set in the query)
ra Recursion Available The server offers recursion (it is a resolver)
ad Authenticated Data The resolver DNSSEC-validated the answer
tc Truncated Answer exceeded UDP size; retry over TCP

Everyday Invocations

# Just the addresses — scriptable
dig +short example.com
# 104.20.23.154

# The answer section only, no preamble or stats — the daily driver
dig +noall +answer example.com
# example.com.    300  IN  A  104.20.23.154

# Target a specific resolver (compare answers across resolvers)
dig @1.1.1.1 example.com
dig @8.8.8.8 example.com

# A specific record type
dig -t MX example.com          # or: dig example.com MX
dig -t AAAA example.com
dig -t TXT example.com

# Reverse lookup (PTR) — -x builds the in-addr.arpa name for you
dig -x 8.8.8.8
dig +short -x 8.8.8.8
# dns.google.

# Force TCP (large answers, or when UDP is filtered)
dig +tcp example.com

# Multiple names / types in one shot
dig +noall +answer example.com A example.com AAAA example.com MX

Finding the Authoritative Servers

# Who is authoritative for the zone?
dig +short example.com NS

# The SOA — serial, refresh/retry/expire, and the negative-cache minimum
dig +noall +answer example.com SOA
# example.com.  3600  IN  SOA  elliott.ns.cloudflare.com. dns.cloudflare.com. \
#   2405749864 10000 2400 604800 1800
#   serial     refresh retry expire negTTL

The SOA's final field is the negative-caching TTL — how long resolvers cache an NXDOMAIN/NODATA for this zone. A high value here is why a just-created record can appear missing for an hour.

Asking the Authoritative Server Directly

This is the workhorse for "the resolver says X but is that actually true?". Point dig at one of the zone's nameservers and disable recursion so you get the raw, authoritative record:

# Ask the zone's own nameserver, no recursion — expect aa in the flags
dig +norecurse @ns.example.com example.com A

# Resulting flags: qr aa   (note: aa set, ra absent)

If the authoritative answer differs from what your resolver returns, the resolver is serving stale or poisoned data — flush it. If the authoritative answer is itself wrong, the zone needs fixing.

Walking the Delegation — +trace

# Follow the delegation from the root downward, querying each level
dig +trace example.com

+trace starts at the root, follows referrals through the TLD to the authoritative servers, and prints each hop. It bypasses your recursive resolver entirely, so it is invaluable for delegation problems: a broken NS record, a lame delegation (parent points at a server that is not authoritative), or a missing glue record shows up as the walk stalling or branching wrong. (It needs outbound UDP/53 to the root and TLD servers — restrictive networks may block it.)

DNSSEC and Wire-Level Flags

# Request DNSSEC records — RRSIG appears alongside the data
dig +dnssec +noall +answer cloudflare.com A
# cloudflare.com.  300  IN  A      104.16.132.229
# cloudflare.com.  300  IN  RRSIG  A 13 2 300 ...  (signature)

# Find authoritative servers and query each for the SOA — quick zone audit
dig +nssearch example.com

Reading TTLs

The number between the name and the class is the TTL in seconds — how long this answer may be cached. Query the same name twice against a recursive resolver and watch it count down toward zero as the cache ages; ask the authoritative server and it resets to the full configured value. A stubbornly high TTL is the usual reason a change "hasn't propagated".

dig +noall +answer +ttlunits example.com     # render TTL as 5m, 1h, etc.

drill, host, and nslookup

drill — the DNSSEC-friendly alternative

drill ships with ldns (ldnsutils / ldns) and is the tool of choice when chasing DNSSEC chains — its output is terser and its trace mode validates signatures.

drill example.com               # basic lookup
drill -x 8.8.8.8                # reverse
drill example.com MX            # by type
drill -D example.com            # request + show DNSSEC records
drill -T example.com            # trace AND validate the chain of trust (the killer feature)
drill -S example.com            # chase and verify the full signature chain

host — quick one-liners

host is the fastest way to get a human-readable answer. Great for a sanity check; thin on detail.

host example.com               # A, AAAA, and MX in plain English
host -t MX example.com         # specific type
host -a example.com            # "all" — verbose, ANY-style dump
host 8.8.8.8                   # reverse lookup (PTR)

nslookup — legacy, avoid

nslookup still exists everywhere and you will see it in old runbooks, but prefer dig. Its interactive mode is clumsy, it hides the status/flags that matter, and its behaviour around search domains and authoritative-vs-recursive answers is muddier. Use it only when nothing else is installed:

nslookup example.com           # basic
nslookup example.com 1.1.1.1   # against a specific server
nslookup -type=mx example.com  # by type

If you reach for nslookup reflexively, retrain the muscle — dig +short is shorter and clearer.

systemd-resolved and the Host's View

Modern Linux desktops and many servers run systemd-resolved, which changes where "the resolver" actually lives. This is the single biggest source of "but dig works and the app doesn't" confusion.

The 127.0.0.53 Stub

/etc/resolv.conf on a resolved host usually points at the local stub:

nameserver 127.0.0.53
options edns0 trust-ad

So dig (which reads resolv.conf) talks to the stub on 127.0.0.53, not to the real upstream. The real upstreams and per-link config live in resolved, inspected with resolvectl:

# Global + per-link resolver config: upstream DNS, search domains, DNSSEC mode
resolvectl status

# Resolve a name THROUGH resolved (honours per-link DNS, split-DNS routing)
resolvectl query example.com

# Cache statistics (hits, misses, current size)
resolvectl statistics

# Flush the resolved cache (does NOT restart the service)
resolvectl flush-caches

# Which search/routing domains apply, per link
resolvectl domain

Per-link DNS matters on multi-homed and VPN hosts: resolved can send corp.example.com queries down the VPN link and everything else to the public resolver. resolvectl status shows this routing; resolvectl query obeys it, whereas a bare dig (hitting the stub or a hard-coded @server) may not.

getent vs dig — the High-Value Distinction

dig and friends speak DNS only. Applications do not — they call the C library's resolver, which obeys /etc/nsswitch.conf:

# /etc/nsswitch.conf
hosts:  files mdns4_minimal [NOTFOUND=return] dns

That hosts: line means a name is resolved by trying, in order: /etc/hosts (files), multicast DNS (mdns), then dns. So:

# What the application layer actually sees (files + mdns + dns, in nsswitch order)
getent hosts example.com
getent ahosts example.com      # all address families, with socket types

When getent hosts foo and dig foo disagree, the answer is almost always one of:

  • an entry in /etc/hosts shadowing DNS (files comes first),
  • mDNS (.local) answering instead of DNS,
  • nsswitch ordering doing something you did not expect.

dig cannot see any of that. Make getent hosts your first command when an app misbehaves but dig looks fine — it reproduces what the app does. This single check resolves a large fraction of "DNS works but the service can't connect" incidents.

Record Types Reference

Type Holds Notes for troubleshooting
A IPv4 address The common case
AAAA IPv6 address App may prefer AAAA; missing AAAA causes "works on IPv4 only"
CNAME Alias to another name Resolver follows the chain; illegal at a zone apex
MX Mail exchanger (priority + host) 0 . (null MX, RFC 7505) means "this domain sends/receives no mail"
TXT Free text Home of SPF (v=spf1 …), DKIM (…._domainkey), DMARC (_dmarc)
NS Authoritative nameservers Delegation; mismatch parent↔child = lame delegation
SOA Zone apex metadata Serial, refresh/retry/expire, and negative-cache TTL
PTR Reverse (IP → name) Lives under in-addr.arpa / ip6.arpa; query with dig -x
SRV Service location (host + port + priority/weight) _service._proto.name, e.g. _sip._tcp
CAA Which CAs may issue certs ACME/Let's Encrypt checks this; wrong CAA blocks issuance
ALIAS / ANAME Provider-specific apex "CNAME" Not a real RR — DNS host flattens it to A/AAAA at query time

SPF/DKIM/DMARC are all TXT records; debug mail auth with dig +short TXT example.com, dig +short TXT _dmarc.example.com, and dig +short TXT selector._domainkey.example.com.

CNAME Chains and Flattening

A CNAME says "this name is really that name". Resolvers follow the chain and return both the alias and the final address:

dig +noall +answer www.github.com
# www.github.com.  3600  IN  CNAME  github.com.
# github.com.        60  IN  A      140.82.121.3   # apex A; low TTL, value rotates

The Apex Gotcha

A CNAME cannot coexist with any other record at the same name, and a zone apex (example.com itself) must carry SOA and NS records. Therefore CNAME at the apex is illegal — you cannot CNAME example.com → myapp.cloud.net. Symptoms: the zone fails to load, or the provider rejects the record. The fixes:

  • Use the DNS host's ALIAS/ANAME pseudo-record, which flattens to A/AAAA at query time.
  • Or use a provider feature like Cloudflare's CNAME flattening.
  • Or point the apex at static IPs with A/AAAA and CNAME only the www subdomain.

Chains Too Long or Looping

Chains should be short. Each hop is another lookup (latency), and a self-referential or mutually-referential pair (a → b → a) is a loop. Most resolvers cap the chain length and return SERVFAIL rather than spinning. Spot it by reading the answer section: if you see many CNAME hops or the chain never reaches an A/AAAA, that is the bug.

TTLs and Caching

The TTL governs how long every cache between the authoritative server and the application may keep an answer. This is why a DNS change is not instant.

"It Hasn't Propagated"

DNS does not push; it expires. After you change a record, old answers persist in caches until their TTL elapses. Check what the authoritative server now serves versus what a public resolver still has:

# Truth, fresh from the source
dig +norecurse @ns.example.com example.com A

# What a public cache is still handing out (watch the TTL count down)
dig @1.1.1.1 example.com A
dig @8.8.8.8 example.com A

If the authoritative answer is new but resolvers still return the old one, you are simply waiting out the TTL. There is no "force propagation" — only patience or a cache flush on resolvers you control.

Lower the TTL Before a Migration

The standard play: at least one full TTL before a cutover, drop the record's TTL to 60–300s. Wait for the old (high) TTL to expire everywhere, do the migration on the low TTL so rollback is fast, then raise it again afterwards. Plan changes around the old TTL, not the new one.

Negative Caching

NXDOMAIN and NODATA answers are cached too, for the duration of the zone's SOA minimum (last SOA field). Create a record that "should exist" and still get NXDOMAIN? A resolver may be holding a negative answer from before the record existed. Check the SOA minimum and wait it out, or flush.

Flushing OS and Browser Caches

# Linux — systemd-resolved
resolvectl flush-caches

# Linux — nscd (if running)
sudo nscd -i hosts

# macOS — flush the system resolver cache
sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder

Browsers keep their own in-process DNS caches (and may use DNS-over-HTTPS, bypassing the OS entirely): Chrome at chrome://net-internals/#dns, Firefox via a restart or about:networking#dns. If dig is correct but only the browser is wrong, suspect the browser cache or its DoH setting.

DNSSEC

DNSSEC signs records so a validating resolver can prove the answer was not tampered with. As an operator you mostly read its signals rather than manage keys.

The ad Flag and RRSIG

# A signed zone, validated by the resolver → ad in the flags
dig cloudflare.com A | grep flags
# ;; flags: qr rd ra ad; ...

# An unsigned zone → no ad flag (this is normal, not an error)
dig github.com A | grep flags
# ;; flags: qr rd ra; ...

# See the signatures themselves
dig +dnssec +noall +answer cloudflare.com A
# ... A   104.16.132.229
# ... RRSIG A 13 2 300 20260616212320 ... (signer, validity window, signature)

ad present means the resolver fetched the signatures and verified them up the chain of trust. ad absent on a zone that is not signed is expected — most zones are unsigned.

The Classic Symptom — SERVFAIL on Validation Failure

When a signed zone's signatures are broken (expired RRSIG, missing DS at the parent, key rollover botched), a validating resolver returns SERVFAIL — not a wrong answer. So if a name gives SERVFAIL from 8.8.8.8/1.1.1.1 (both validate) but resolves fine from a non-validating resolver, suspect DNSSEC:

# Validating resolver — SERVFAIL if DNSSEC is broken
dig dnssec-failed.org @8.8.8.8 | grep status
# ;; ->>HEADER<<- ... status: SERVFAIL

# Ask without validation (+cd = checking disabled) — if THIS succeeds, it's DNSSEC
dig +cd dnssec-failed.org @8.8.8.8 | grep status
# ;; ->>HEADER<<- ... status: NOERROR

+cd (checking disabled) tells the resolver to skip validation. SERVFAIL that clears under +cd is a DNSSEC validation failure, full stop.

Validating Locally — delv

delv (ships with BIND tools) performs full validation client-side and tells you the verdict in plain language:

delv cloudflare.com A
# ; fully validated
# cloudflare.com.  300  IN  A  104.16.132.229
# cloudflare.com.  300  IN  RRSIG A ...

delv dnssec-failed.org A
# ;; resolution failed: failure
# Current BIND (9.20+) prints the terse "failure"; older versions said
# "no valid signature found". Either way: validation failed, chain of trust broken.

drill -T / drill -S do the same from the ldns side. Use whichever is installed to confirm whether the break is signing-side (zone) or just your resolver.

Split-Horizon (Split-Brain) DNS

The same name resolves to different answers depending on who is asking — typically a private RFC 1918 address for clients inside the network or on the VPN, and a public address for everyone else. It is a deliberate design, and a frequent source of "works for me, not for them".

Detecting It

Compare an internal resolver against a public one:

# The internal / corporate resolver's answer
dig @10.0.0.53 app.corp.example.com

# A public resolver's answer for the same name
dig @8.8.8.8  app.corp.example.com

Different addresses (or NXDOMAIN from the public side) confirm split-horizon. On a host using systemd-resolved, the deciding factor is which link's resolver the query is routed to:

resolvectl status                 # per-link DNS + which domains route where
resolvectl domain                 # routing/search domains per link
resolvectl query app.corp.example.com   # resolve the way the system actually will

VPN and Search-Domain Effects

When a VPN comes up it usually pushes resolver and search-domain settings. Symptoms of getting this wrong: internal names fail because the VPN's resolver was not installed for that domain, or external names break because the VPN grabbed all DNS. resolvectl status shows whether the VPN link is DNS Domain ~corp.example.com (route only that domain) or has the catch-all ~. (route everything). The ~. catch-all on the wrong link is a classic "VPN broke my internet" cause.

Common Failure Modes

The fastest path to root cause is reading the status code and which server answered, then moving along the resolution chain.

NXDOMAIN vs SERVFAIL vs REFUSED vs Timeout

Symptom Likely cause Next step
NXDOMAIN Name genuinely absent; typo; wrong zone; search-domain suffixed a bad name Check spelling and the trailing dot; dig +norecurse @authoritative; check negative cache
SERVFAIL DNSSEC validation failure; authoritative unreachable; broken upstream dig +cd (rules in/out DNSSEC); dig @authoritative directly; dig +trace
REFUSED Server not authoritative and recursion off; ACL blocks your IP You are asking the wrong server, or from a disallowed source — use a resolver that serves you
Timeout / "no servers could be reached" Resolver unreachable; UDP/53 filtered; wrong nameserver in resolv.conf resolvectl status; dig @1.1.1.1 to bypass local config; check firewall, ss -u

Resolver Not Reachable

# Is the configured resolver even responding? Bypass it with a known-good one.
dig @1.1.1.1 example.com
dig @8.8.8.8 example.com

# What does the host think its resolver is?
cat /etc/resolv.conf       # often the 127.0.0.53 stub
resolvectl status          # the real upstreams behind the stub

If @1.1.1.1 works but the default does not, the fault is in local resolver config (resolv.conf, resolved upstream, or a dead local cache), not in DNS itself.

Wrong Search Domain and ndots

/etc/resolv.conf carries search domains and an options ndots:N. A name with fewer than ndots dots is tried with each search domain appended first, before being tried as-is. Kubernetes sets ndots:5 by default, so dig api (and many in-cluster lookups) generate a flurry of failed suffixed queries before the bare name — a real latency and NXDOMAIN-noise source.

cat /etc/resolv.conf
# search corp.example.com
# options ndots:5

# 'api' becomes api.corp.example.com. first. To force the literal name,
# add the trailing dot — it means "fully qualified, do not append search":
dig api.example.com.        # the trailing dot disables search suffixing

A surprising NXDOMAIN for a name you know exists is very often the search domain being appended (or not appended when you expected it). The trailing dot is the cure: it marks the name as fully qualified.

/etc/hosts Shadowing

getent hosts example.com    # if this differs from dig, check /etc/hosts first
grep example.com /etc/hosts

Because files precedes dns in nsswitch.conf, a stray /etc/hosts line (often a leftover 127.0.0.1 someservice) silently overrides DNS for applications while dig — which never reads /etc/hosts — shows the "correct" DNS answer. This mismatch is exactly why getent is your first reproduction command.

EDNS, Truncation, and TCP Fallback

Large answers (many records, DNSSEC signatures) can exceed the UDP size and come back truncated (tc flag set), at which point the resolver retries over TCP. Middleboxes that block DNS-over-TCP, or break EDNS, cause intermittent failures on exactly the big responses.

# See truncation, then confirm TCP works
dig +dnssec example.com         # look for 'tc' in the flags on a big answer
dig +tcp example.com            # force TCP — if this works but UDP fails, suspect EDNS/MTU
dig +bufsize=512 example.com    # shrink the advertised EDNS buffer to dodge fragmentation

A name that resolves with +tcp but fails or truncates over plain UDP points at an EDNS/MTU or middlebox problem on the path.

The Missing Trailing Dot

In zone files and many tools, a name without a trailing dot is relative and gets the origin appended; with the dot it is absolute. The classic bug: a CNAME target written www.example.com (no dot) in a zone for example.com silently becomes www.example.com.example.com.. From the operator side, the same logic explains surprise suffixing in lookups — when in doubt, append the dot and re-query.

Quick Reference

dig Flags and Options

Option Effect
+short Answer data only, one value per line
+noall +answer Just the ANSWER section, no preamble or stats
@server Query a specific resolver/authoritative server
-t TYPE Query a record type (A, AAAA, MX, TXT, NS, SOA, …)
-x ADDR Reverse (PTR) lookup
+trace Walk the delegation from the root downward
+norecurse Do not request recursion (ask an authoritative server directly)
+dnssec Request DNSSEC records (RRSIG, etc.)
+cd Checking disabled — skip DNSSEC validation (diagnose SERVFAIL)
+tcp Force the query over TCP
+nssearch Find authoritative servers and query each for the SOA
+ttlunits Render TTLs as 5m, 1h, etc.
+bufsize=N Set the advertised EDNS UDP buffer size

Record Types

Type Purpose
A / AAAA IPv4 / IPv6 address
CNAME Alias to another name (never at a zone apex)
MX Mail exchanger (priority + host; 0 . = null MX)
TXT Free text — SPF, DKIM, DMARC live here
NS Authoritative nameservers (delegation)
SOA Zone apex metadata + negative-cache TTL
PTR Reverse mapping (IP → name)
SRV Service location (host, port, priority, weight)
CAA Which CAs may issue certs for the name
ALIAS/ANAME Provider apex "CNAME" flattened to A/AAAA

Resolver / Host Commands

Task Command
Resolve as the app does (NSS) getent hosts NAME
Resolver config behind the stub resolvectl status
Resolve through resolved resolvectl query NAME
Flush resolved cache resolvectl flush-caches
Flush macOS cache sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder
Quick human-readable lookup host NAME
Validate DNSSEC locally delv NAME / drill -T NAME

Common Issues and Solutions

Issue Cause Solution
Change "hasn't propagated" Old answer still within its TTL in caches Compare dig +norecurse @authoritative vs dig @1.1.1.1; wait out the TTL; lower TTL before the next change
New record returns NXDOMAIN Negative answer cached (SOA minimum) Check the SOA minimum; flush resolvers you control; wait it out
getent and dig disagree /etc/hosts or mDNS shadowing DNS via nsswitch ordering Inspect /etc/hosts and /etc/nsswitch.conf; dig cannot see either
SERVFAIL on a signed zone DNSSEC validation failure (expired RRSIG, bad DS) dig +cd to confirm; delv/drill -T to locate the break; fix signing or DS
dig works, app doesn't App hits the stub/NSS, not your @server Reproduce with getent hosts and resolvectl query, not bare dig
Internal name fails on VPN Per-link DNS / search-domain routing wrong resolvectl status; check ~domain vs ~. catch-all on each link
Same name, different answers Split-horizon DNS dig @internal vs dig @8.8.8.8; expected by design — query the right resolver
Surprise NXDOMAIN for a real name Search domain appended (or not) per ndots Append a trailing dot to force the literal name; check resolv.conf search/ndots
Big lookups fail, small ones fine EDNS/truncation, TCP fallback blocked dig +tcp and dig +bufsize=512; fix the middlebox/MTU on the path
REFUSED from a server Not authoritative + recursion off, or ACL Query a resolver that serves you, or the zone's authoritative NS
CNAME rejected at apex CNAME illegal alongside SOA/NS at the zone root Use ALIAS/ANAME/flattening, or A/AAAA at the apex

Related Topics

The following topics complement this DNS troubleshooting cheatsheet:

  1. Linux Network Tools — getent, resolvectl, and ip route get for the host-side view of resolution and routing
  2. tcpdump / Wireshark — capture port-53 traffic to see the exact queries and answers on the wire (udp port 53)
  3. Nginx — resolver directives, upstream name resolution, and DNS-driven service discovery for proxied backends
  4. Cert-Manager (ACME) — DNS-01 challenges, CAA records, and propagation checks that gate certificate issuance
  5. Consul — service discovery that answers DNS for .consul, and the split-horizon patterns it introduces
  6. Kubernetes Networking (CoreDNS) — in-cluster resolution, the ndots:5 default, and ClusterFirst vs Default dnsPolicy