Skip to content

Day 16 of 100

The Memory Leak Nobody Noticed

SREApplication ReliabilityIntermediate15–20 min

🔥 Problem Statement

A production service has been running fine all morning. By mid-afternoon, one of the pods restarts. Then another. Nobody paged anyone — restarts aren't alerting yet, they're just showing up as a rising number on a dashboard nobody's watching closely.

Pulling up the memory graph for the affected pods tells a clear story: 42% at 9 a.m., climbing steadily, hour over hour, to 91% by 3 p.m., right before each restart. CPU is flat the entire time. Traffic is flat the entire time. Nothing else on the dashboard looks unusual — just memory, climbing in a straight line until the container gets killed and starts fresh.

🏗️ Environment

  • Platform: Kubernetes, Deployment with 3 replicas
  • Language runtime: long-running Node.js service
  • Monitoring: CPU, memory, restart count
  • Observed over: an 8-hour window

🤔 Your Challenge

What's the difference between a traffic-driven memory increase and a memory leak, and how would you tell which one this is?

  • If traffic is flat, what would explain memory climbing steadily instead of staying roughly constant?
  • Why would CPU stay normal while memory climbs — what does that rule out?
  • What would you expect the memory graph to look like right after a restart if this is a leak?
  • What tools would let you see what's actually holding onto memory inside the process?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

Flat traffic ruling out load as the cause is the key detail. If more requests were driving more memory use, the graph would track traffic — rising and falling with it, not climbing in a straight line regardless of demand. A steady, traffic-independent climb that resets to baseline after every restart is close to the textbook shape of a memory leak: something is allocating memory during normal operation and never releasing it, so usage grows with time and workload volume (total requests processed), not with concurrent load.

Step 2 — Identify the likely bottleneck

CPU staying flat while memory climbs actually narrows things usefully — it rules out a computation-heavy cause and points at something more passive: objects being retained in memory that should have been garbage collected, a cache with no eviction policy quietly growing forever, event listeners or timers accumulating without being cleaned up, or a connection/resource pool that's not releasing what it opens. None of these show up as extra CPU work; they just sit there, taking up space.

Step 3 — Investigation

  • Confirm the reset-on-restart pattern across multiple cycles — if memory reliably returns to the same low baseline after every restart and climbs at a similar rate each time, that's strong confirmation of a leak rather than a one-off anomaly.
  • Take a heap snapshot (or the runtime's equivalent — Node's --inspect and Chrome DevTools, a JVM heap dump, etc.) at two points hours apart and diff them — this shows exactly what object types are growing in count, which is usually the fastest way to point at the actual leaking code path.
  • Check for unbounded in-memory caches — a Map or object used as a cache with no TTL or size limit is one of the most common leak sources in long-running services.
  • Check for event listener or timer accumulation — a setInterval or event subscription created per-request but never cleaned up will grow linearly with request volume.
  • Correlate the climb rate against request volume, not traffic level — if a leak grows per-request rather than per-second, its rate will roughly track cumulative requests, not current load, which is a very telling signal.

Fix the actual retention point the heap diff surfaces — add an eviction policy to an unbounded cache, ensure listeners/timers are cleaned up on completion, close pooled resources properly. Don't just raise the memory limit or increase restart frequency as a workaround; that hides the symptom (fewer visible restarts) while leaving the underlying leak to keep consuming more memory per hour of uptime, and eventually the limit increase runs out of headroom too.

In the meantime, if restarts are currently silent and unalerted, add monitoring on restart count and memory trend specifically — a slow, one-directional climb over hours is a distinctive enough pattern that it's worth its own alert, separate from a generic memory-threshold alert that might not catch a leak until it's already causing restarts.

Step 5 — Engineering lesson

CPU and memory tell different stories, and a dashboard that only draws attention to CPU spikes will miss this kind of incident entirely — the graph was there, climbing in plain sight, for six hours before anyone noticed. A memory leak is exactly the kind of failure that's invisible until it isn't: quiet, gradual, and easy to dismiss as normal until the restarts start.

🧠 Today’s Takeaway

A stable CPU graph doesn't mean the application is healthy.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.