Day 11 of 100
The Database Is Healthy, But the Application Is Slow
π₯ Problem Statement
The dashboard is almost aggressively unremarkable:
```text Database CPU: 45% Database memory: Normal Application CPU: 50% Application memory: Normal Network: Normal ```
And yet API latency has climbed significantly, and users are noticing. Every metric anyone would normally check first says "everything is fine," which is exactly the problem β those metrics were supposed to explain a real, measurable slowdown, and they don't.
ποΈ Environment
- Architecture: API service calling internal + external dependencies
- Database: healthy per all standard metrics
- Observability: distributed tracing available
- Monitoring: CPU, memory, network
π€ Your Challenge
If CPU, memory, and database metrics all look normal, where would you investigate next?
- What kinds of latency wouldn't show up as elevated CPU or memory usage anywhere?
- Could a request spend most of its time waiting rather than computing, and what would that look like on a trace?
- Are there dependencies in this request path that aren't the primary database β a cache, a third-party API, another internal service?
- What would a connection pool or thread pool near its limit look like from the outside, before it shows up as a resource metric?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
CPU, memory, and database load all being normal rules out the causes that show up as resource saturation β the application isn't computing too much, the database isn't struggling under query load. That's useful, but it only narrows the field; it doesn't explain the latency. What none of those metrics capture is waiting β time spent blocked on a lock, queued for a connection, or waiting on a network call to something that isn't the primary database at all.
Step 2 β Identify the likely bottleneck
Latency without resource saturation is close to a textbook signature of a waiting problem rather than a computing problem. A few candidates fit this shape specifically: a connection pool (to the database, or to another service) that's near its limit, causing requests to queue for a connection even though the database itself has headroom; a downstream call to something other than the primary database β a cache, a third-party API, an internal service β that's slow or intermittently timing out; or lock contention inside the database on specific rows or tables, which can slow individual queries significantly without ever showing up as elevated overall CPU.
Step 3 β Investigation
- Pull distributed traces for the slow requests specifically. This is the tool built exactly for this situation β it shows where time is actually spent across every hop in the request, not just at the two endpoints (application and database) that the dashboard happens to show.
- Look for time spent outside the primary database entirely β a cache lookup, an external API call, an internal service-to-service call that doesn't appear on the standard dashboard at all.
- Check connection pool metrics specifically, not just database load β a pool sitting near its max size means requests queue for a connection before a query is even issued, and that queueing time is invisible to database-side metrics.
- Check for lock waits inside the database β most databases expose this directly (e.g., PostgreSQL's
pg_stat_activityshowing sessions waiting on a lock), and it's a specific, checkable thing rather than a guess.
Step 4 β Recommended action
Resist the pull to add more database resources β the database's own metrics already say it isn't the bottleneck, so more database capacity is very unlikely to fix a problem the database isn't causing. Follow the trace to find the actual hop where time accumulates, and fix that specific thing: increase a connection pool's size if it's genuinely undersized for current concurrency, fix or add a timeout to a slow downstream dependency, or resolve the specific lock contention a database investigation surfaces.
Step 5 β Engineering lesson
The standard four metrics β CPU, memory, database load, network β are a good first check precisely because they're fast to look at, but they don't cover the whole space of ways a distributed system can be slow. Time spent waiting on a connection, a lock, or a downstream call doesn't register on any of them. When the obvious metrics all say "healthy" and the user experience says otherwise, that gap is the signal to reach for a trace β it's the tool that actually answers "where did the time go," which none of those four metrics were ever built to answer on their own.
π§ Todayβs Takeaway
Healthy infrastructure metrics don't always mean a healthy application.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

