Skip to content

Day 17 of 100

The DNS Failure

CloudNetworkingSREIntermediate15–20 min

🔥 Problem Statement

A service that normally calls a downstream API without issue starts throwing connection errors — not consistently, but often enough to show up in the error rate. The calling application's own health checks are green. Its CPU, memory, and thread pool all look normal. Nothing about the application itself looks unhealthy.

The errors are specifically about not being able to reach the downstream service, not about the downstream service returning bad responses — which means, whatever's wrong, it's happening before a request even lands anywhere.

🏗️ Environment

  • Cloud: multi-service architecture, private networking
  • Service discovery: internal DNS for downstream service
  • Health checks: application-level, all green
  • Symptom: intermittent connection failures to one downstream service

🤔 Your Challenge

How would you prove — not just suspect — whether DNS resolution is actually the problem here?

  • What's the difference between a connection error and an application error, in terms of what layer it points to?
  • If DNS resolution is intermittent, what would that look like in application logs versus in a direct DNS query?
  • What role does TTL play in how quickly a DNS change propagates — or how long a bad entry lingers?
  • How would you test DNS resolution independently of the application, to isolate it as a variable?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

A connection error — as opposed to an HTTP error code or a slow response — means the request never successfully established a connection to begin with. That points upstream of the application layer entirely: something in resolving the destination address or establishing the network path is failing, intermittently, which is a different category of problem than "the downstream service is unhealthy." The calling application being otherwise healthy supports this — there's nothing wrong with how it's making the request, only with whether the request can find its target.

Step 2 — Identify the likely bottleneck

Intermittent, connection-level failures to one specific destination, with everything else about the environment normal, is a common signature of DNS resolution issues — a resolver that's occasionally slow or failing, a DNS entry that's stale or was recently changed, or a TTL that's shorter or longer than the systems around it assume. Unlike an application bug, DNS problems are often genuinely intermittent by nature — some queries hit a cache with the right answer, others hit a resolver having a bad moment, which produces exactly this kind of "sometimes it works, sometimes it doesn't" pattern.

Step 3 — Investigation

  • Query DNS directly and repeatedly, independent of the application — a simple loop resolving the downstream service's hostname every few seconds will show whether resolution itself is intermittently failing or slow, without any application code in the way.
  • Check the TTL on the relevant DNS record — a very short TTL means frequent re-resolution (more exposure to a flaky resolver); a very long TTL means a recent change could still be serving a stale, wrong answer to some callers.
  • Check whether this is private/internal DNS or public DNS — internal service discovery has its own failure modes (a service registry lagging, a recent deployment changing an internal endpoint) distinct from public DNS resolver issues.
  • Correlate failure timing against any recent infrastructure change — a service migration, a load balancer replacement, or a scaling event that changed the downstream service's actual address is a common trigger for exactly this kind of transient DNS mismatch.
  • Check resolver-level metrics or logs, if available — many environments expose resolver query success/failure rates, which would directly confirm or rule out DNS as the cause rather than inferring it from symptoms.

If direct DNS queries confirm intermittent resolution failures, the fix depends on what's actually failing: correct a stale record, adjust TTL to match how often the underlying address legitimately changes, or address resolver capacity/reliability if it's a shared resolver under load. If DNS turns out to be fine and the issue is elsewhere in the network path, this investigation has still ruled out an entire category of causes quickly and cheaply — which is the point of testing DNS in isolation rather than guessing from application-level symptoms alone.

Don't restart the application or add retries as a first response — retries can paper over a DNS problem just enough to reduce visible errors while the underlying cause (a bad or stale record) keeps affecting a fraction of requests indefinitely.

Step 5 — Engineering lesson

A connection failure and an application failure look similar from a dashboard showing error rates, but they point to completely different places to investigate. DNS is invisible infrastructure — it works so reliably, so often, that it's easy to skip past as a possible cause and go straight to application-level debugging. Testing it directly and in isolation takes a few minutes and either confirms or rules it out decisively, which is faster than working through application code that was never actually the problem.

🧠 Today’s Takeaway

When a service is unreachable, don't assume the application is broken.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.