Day 15 of 100
The SLO Is Green, But Customers Are Complaining
π₯ Problem Statement
The reliability dashboard is telling a clean story:
```text Availability SLO: 99.95% Error Budget: Healthy Infrastructure: Healthy ```
Meanwhile, customer support is fielding a steady stream of complaints about slow responses. Both things are true at once, which is the actual puzzle β the number that's supposed to represent "is the service working for users" says yes, and the users are saying no.
ποΈ Environment
- SLO: 99.95% availability, error-budget tracked
- SLI: aggregate success rate across all endpoints/regions
- Support: customer complaints about slow responses
- Monitoring: dashboard-level aggregate metrics
π€ Your Challenge
How can the SLO stay green while customers are genuinely unhappy, and what additional signals would you look at?
- Does an aggregate availability number across all endpoints and regions hide problems concentrated in just one of them?
- Is this SLO measuring availability (did the request succeed) or latency (did it succeed fast enough) β and does that distinction matter to a frustrated user?
- Could a small, high-value subset of users or transactions be badly affected while the aggregate stays fine?
- What would synthetic monitoring or real user monitoring show that a server-side aggregate wouldn't?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
An aggregate availability SLO across every endpoint and every region is exactly that β an aggregate. It can stay comfortably above 99.95% while a specific region, a specific high-traffic endpoint, or a specific class of transaction is performing badly, as long as the badly-affected slice is small relative to total volume. The dashboard isn't lying; it's answering a broader question than the one customers are actually experiencing.
There's a second gap worth naming directly: this SLO measures availability β did the request return a successful response β not latency. A request that eventually succeeds after three seconds counts identically to one that succeeds in 80 milliseconds, as far as this specific SLI is concerned. "Slow" and "unavailable" are different complaints, and this SLO was only ever built to catch one of them.
Step 2 β Identify the likely bottleneck
Two plausible, non-exclusive explanations fit this pattern well: either the complaints are concentrated in a specific segment β one region, one endpoint, one customer-facing flow β that's genuinely degraded while the aggregate dilutes it into invisibility, or the SLI itself is measuring the wrong thing for what "healthy" means to users, tracking success/failure while users are actually experiencing a latency problem the current SLO was never designed to catch.
Step 3 β Investigation
- Break the availability metric down by region and by endpoint, not just the aggregate. A regional or endpoint-specific problem is one of the most common causes of exactly this "SLO green, users unhappy" gap.
- Check latency percentiles specifically β P95, P99 β not just success/failure counts. A rising tail latency can make the experience feel broken for a meaningful slice of users while never touching an availability SLI at all.
- Look at specific business transactions, not just endpoint-level technical success. A checkout flow or a login flow failing partway through β even if each individual request "succeeds" at the HTTP level β is exactly the kind of partial failure an availability-only SLO won't catch.
- Pull actual customer complaint details β timestamps, regions, specific actions being attempted β and correlate them directly against the segmented metrics above rather than the aggregate dashboard.
- Compare against synthetic monitoring or real user monitoring (RUM) if available β these measure the experience from the user's actual vantage point, which can diverge meaningfully from server-side aggregate metrics.
Step 4 β Recommended action
Fix the immediate gap two ways: address whatever the segmented investigation actually finds (the specific region, endpoint, or transaction that's degraded), and separately, treat the SLO definition itself as something that needs revisiting. If latency is part of what "healthy" means to users β and for most services it is β a latency-based SLO alongside the availability one closes exactly the gap this incident exposed. If certain transactions matter more than others to the business, an SLO scoped to those specific critical paths, not just an all-endpoints aggregate, would have surfaced this far sooner.
Step 5 β Engineering lesson
An SLO is only as good as the SLI underneath it, and an SLI is only useful if it actually represents what "working" means to the people it's meant to protect. A green dashboard measuring the wrong thing β or measuring the right thing at too coarse a level β isn't lying, but it also isn't doing its job. Customer complaints contradicting a green SLO aren't a sign to ignore the complaints; they're a sign the SLO needs to get closer to what users actually experience.
π§ Todayβs Takeaway
If your SLI doesn't represent user experience, a green SLO can still hide a real problem.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

