Skip to content

Day 18 of 100

The Certificate That Expired

SecuritySREIntermediate15–20 min

🔥 Problem Statement

Users start reporting they can't reach the site at all — browsers show certificate warnings, and API clients fail the TLS handshake before a single request byte gets through. The application logs show nothing wrong; from the service's own point of view, it's healthy and idle, because the failure is happening before any request reaches it.

A quick check confirms it: the TLS certificate expired at midnight. It was provisioned months ago, worked fine the entire time, and nobody was watching the calendar.

🏗️ Environment

  • Public-facing API behind a load balancer, TLS terminated at the edge
  • Certificate: previously provisioned, no automated renewal in place
  • Application: fully healthy, unaffected by the change
  • Monitoring: no certificate-expiry alerting configured

🤔 Your Challenge

Beyond renewing the certificate right now, how would you make sure this specific failure mode can't happen again?

  • Why does a certificate expiry produce a total outage rather than a partial one?
  • What's the difference between fixing this incident and fixing the process that allowed it?
  • At what point before expiry should renewal actually happen, and why not wait until the last day?
  • What would make certificate expiry visible days or weeks in advance instead of showing up as a customer-facing outage?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

Certificate expiry causes a total, hard failure rather than a degraded one because TLS handshake validation happens before any application logic runs — a browser or client that can't validate the certificate refuses to send the actual request at all. That's why the application's own logs and health checks show nothing wrong: from its perspective, no requests are arriving to fail. The break is entirely at the edge, between the client and the point where TLS is terminated.

Step 2 — Identify the likely bottleneck

This wasn't caused by a bug — it was caused by an absence: no automated renewal, and no monitoring that would have surfaced the approaching expiry date as a problem before it became an outage. A certificate provisioned once and left alone will always eventually expire; the actual root cause here is a process gap, not a technical one, which matters because it changes what "fixing" this incident actually requires.

Step 3 — Investigation

  • Confirm the exact expiry timestamp and renew immediately — this restores service and is the first priority, separate from the process fix.
  • Check whether this certificate was meant to auto-renew and didn't, or whether it was always a manual process — a broken automation is a different fix than a manual process that was simply forgotten.
  • Check for other certificates in the environment with similarly no renewal automation — if this one slipped through, others provisioned the same way are worth auditing before they cause the same outage independently.
  • Check what monitoring exists (or doesn't) for certificate expiry — most environments have no default alerting on this, since a certificate "working" gives no signal that it's approaching its end date.
  • Review who owned certificate renewal and whether that ownership was clear, documented, and actually assigned to someone's regular responsibilities — an unowned recurring task is exactly the kind of thing that gets missed.

Beyond the immediate renewal, this calls for automated renewal wherever the platform supports it — most modern certificate authorities and cloud load balancers support automatic renewal well before expiry, removing the human-memory dependency entirely. Where automation genuinely isn't possible, add explicit expiry monitoring that alerts 30 and 14 and 7 days out, escalating — not a single alert on the day of expiry, which leaves no time to react, but a graduated warning that gives whoever owns it real lead time.

This is also worth treating as a small platform-wide audit, not just a one-certificate fix — if this certificate had no renewal automation, it's worth checking whether that's true of every certificate in the environment, since the same gap likely exists in more than one place.

Step 5 — Engineering lesson

An expired certificate isn't a technical failure in the usual sense — nothing broke, nothing crashed, the code did exactly what it was supposed to do. It's a process failure: a recurring maintenance task with no automation and no monitoring, waiting for the one day someone doesn't remember. The fix that actually prevents a repeat isn't "renew the certificate" — it's making sure the next expiry date, whenever it is, can't quietly arrive unnoticed again.

🧠 Today’s Takeaway

Certificate management is an operational responsibility, not a one-time setup.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.