Day 6 of 100
The Alert Storm
🔥 Problem Statement
It's 2:00 a.m. and the on-call phone won't stop buzzing. In four minutes, over 200 alerts fire across roughly 40 different services — elevated error rates, latency spikes, a few pods restarting. Slack is a wall of red.
The first instinct in the incident channel is to treat this as a mass outage: "100 things are broken, we need 100 people looking at 100 things." But nobody's actually confirmed that yet — it's just what 200 alerts in four minutes *feels* like at 2 a.m.
🏗️ Environment
- Architecture: ~30 microservices, shared dependency graph
- Alerting: per-service threshold alerts (CPU, latency, error rate)
- Observability: metrics + logs + traces
- On-call: single SRE, PagerDuty
🤔 Your Challenge
How would you determine whether these are genuinely independent incidents, or one underlying failure triggering alerts across everything downstream of it?
- If every one of these services shares a common dependency, what would you expect their alerts to look like when that dependency fails?
- What's the difference between an alert on a symptom and an alert on a root cause, and which kind fired here?
- How would a service dependency map change how you'd read this alert list?
- What's the fastest way to find the one thing that's actually broken, versus the forty things reacting to it?
Solution Hidden
Think through the problem yourself before looking at the answer.
💡Solution
Step 1 — Understand the symptoms
Two hundred alerts across forty services in four minutes is a pattern, not a coincidence. Independent failures don't usually line up that tightly in time — services failing for forty unrelated reasons tend to trickle in over minutes to hours, each on its own schedule. A near-simultaneous burst across a wide swath of the architecture is the signature of something shared breaking underneath all of them.
Step 2 — Identify the likely bottleneck
The question worth asking isn't "which forty services are broken," it's "what do these forty services have in common." If they share a dependency — a database, a caching layer, an internal auth service, a shared network path — the far more likely story is: that shared dependency degraded or failed, and every service downstream of it threw its own error-rate and latency alerts as a direct symptom, at roughly the same moment, because they were all affected by the same upstream event simultaneously.
That reframes 200 alerts from "200 problems" to "one problem, with 200 symptoms." The alerting system did its job — it's just alerting on symptoms (this service's error rate) rather than root cause (the shared dependency is down), and without a dependency map in hand, that distinction is easy to miss under 2 a.m. pressure.
Step 3 — Investigation
- Build or pull up the service dependency graph — even a rough one. Which of the alerting services share a common upstream dependency?
- Look at alert timestamps precisely, not just "they all fired around 2 a.m." A shared-cause failure tends to show a tight clustering — seconds apart, not spread across minutes — while independent failures cluster far more loosely.
- Check the shared dependency's own health directly — its own metrics, not the forty services reacting to it. If it's a database, check its connections and query latency. If it's an internal auth or config service, check its own error rate and availability.
- Sample a few traces from the affected services and see where they spend their time — if they all show elevated latency at the same specific downstream call, that call is the actual failure point.
Step 4 — Recommended action
Don't split the incident into forty separate investigations. Assign one person (or a small group) to trace the shared dependency, and treat the forty service-level alerts as confirmed symptoms rather than forty separate mysteries to solve independently — that's both faster and avoids forty people converging on the same root cause from forty different angles.
Once the shared dependency is identified and restored, most of the downstream alerts should clear on their own without individual intervention — which is itself useful confirmation that the root-cause theory was correct, versus forty coincidentally unrelated failures that would need forty separate fixes.
Step 5 — Engineering lesson
Alert volume measures how visible a failure is, not how many failures there are. A well-connected architecture means a single upstream failure can legitimately look like a mass outage in the alert list, and the fix for that isn't fewer alerts — it's correlation. SLO-based alerting on the services that matter most, paired with a dependency map that's actually kept up to date, turns "200 alerts, panic" into "one root cause, forty confirmed symptoms" in the first five minutes instead of the first hour.
🧠 Today’s Takeaway
A thousand alerts don't necessarily mean a thousand problems.
🔔 Don’t Miss Tomorrow’s Challenge
A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

