Skip to content

Day 1 of 100

The 2-Second API Mystery

SREObservabilityIntermediate15–20 min

πŸ”₯ Problem Statement

It's a normal Tuesday until the on-call channel lights up: the main API's P95 latency has gone from 300ms to 2.4 seconds. Traffic is running about 2x higher than usual β€” nothing unheard of, this happens during a marketing push a few times a quarter.

You pull up the dashboard. Application CPU is sitting at 92%. Database CPU is at 95%. Database connections are at 490 out of a 500 max pool. Error rate is still low β€” requests are succeeding, they're just slow.

The application team is already in the channel with a plan: "CPU is high. Let's scale the application." Someone's finger is hovering over the button to double the Cloud Run instance count.

πŸ—οΈ Environment

  • Cloud: GCP
  • Application: Node.js API
  • Platform: Cloud Run
  • Database: PostgreSQL
  • Observability: OpenTelemetry
  • Monitoring: Metrics + Logs + Traces

πŸ€” Your Challenge

Would you scale the application immediately, or is there something else worth checking first?

  • Which of these two maxed-out metrics β€” application CPU or database CPU β€” is the actual bottleneck, and which is downstream of the other?
  • What does a connection pool sitting at 490/500 usually mean under 2x traffic?
  • If you scaled the application right now, what would you expect to happen to the database?
  • What would you want to see in a trace before deciding anything?

Solution Hidden

Think through the problem yourself before looking at the answer.

πŸ’‘Solution

Step 1 β€” Understand the symptoms

Two things are maxed out at the same time: application CPU and database CPU. It's tempting to read that as "everything is under load, scale everything" β€” but two metrics being high simultaneously doesn't mean they're two independent problems. It's just as likely to mean one is causing the other.

The detail that actually matters here is the connection pool: 490 out of 500. That's not "busy." That's a pool one bad afternoon away from exhaustion. When a connection pool gets that close to its ceiling under elevated traffic, requests start queuing for a connection before they even reach the database β€” and a request sitting in a queue burns CPU on both sides of that queue while doing zero useful work.

Step 2 β€” Identify the likely bottleneck

Application CPU at 92% looks like "the app needs more capacity." But if every one of those Node.js event-loop cycles is spent waiting on a database connection that isn't available yet β€” checking the pool, retrying, holding open requests β€” then the CPU number is measuring contention, not throughput. The application isn't doing more work under 2x traffic. It's doing the same work, slower, with more of it stuck waiting.

That reframes the database CPU number too. At 95% database CPU with the pool nearly exhausted, the most likely story is: traffic doubled, query volume roughly doubled, and something about how queries execute under load β€” a missing index, a query that scans instead of seeks, lock contention on a hot row β€” means the database can't clear requests fast enough to keep the pool cycling. Connections pile up waiting on slow queries, the pool fills, new requests queue behind it, and application CPU climbs handling all that waiting and retrying.

Step 3 β€” Investigation

This is exactly what a distributed trace is for. Pull a handful of P95/P99 traces from the slow window and look at where the time actually goes β€” specifically, how much of a request's total latency is spent inside the database span versus everywhere else. If the database span dominates the trace, that settles it.

From there:

  • Slow query logs β€” which queries are taking longest right now, and is that different from a normal-traffic baseline?
  • Query plans β€” for the slowest queries, is PostgreSQL doing a sequential scan where an index would help? EXPLAIN ANALYZE on the worst offenders.
  • Lock contention β€” pg_stat_activity for anything sitting in active with a long query_start, or waiting on a lock.
  • Connection pool metrics over time β€” did the pool fill gradually as traffic ramped, or did it spike suddenly? Gradual points at query performance degrading under load; sudden points at something else entirely (a deploy, a lock, a runaway query).

Don't scale the application yet. If the database is the actual bottleneck, doubling application instances just means twice as many processes competing for the same constrained pool and the same overloaded database β€” CPU usage might even go up while P95 gets worse, not better, because you've added more contention without adding any actual database capacity.

The right sequence: confirm via traces and query plans where the time is going, fix the specific slow-query or indexing problem if one exists, and only then decide whether the application layer itself needs more capacity β€” which it might, but that's a decision to make with evidence, not as the first reflex when a dashboard looks red.

Step 5 β€” Engineering lesson

The system was telling the truth the whole time β€” application CPU really was at 92%. It just wasn't telling you why. Two metrics maxing out together, under elevated traffic, is close to the textbook shape of a downstream bottleneck creating an upstream symptom: the database is slow, so requests queue for connections, so the application spends its CPU waiting instead of working.

The instinct to reach for the metric that's easiest to act on β€” "CPU is high, add more compute" β€” is understandable under pressure. It's also how you can make an incident worse while believing you're fixing it. A five-minute look at a trace would have told the team exactly where the 2 seconds were going before anyone touched the instance count.

🧠 Today’s Takeaway

Don't fix the loudest metric. Find the bottleneck.

πŸ”” Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days β†’ 100 Challenges β†’ 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.