Skip to content
SRE

Day 14: Systems Will Fail. Speed of Recovery Wins.

Chasing a longer mean time between failures is chasing an asymptote. Chasing a shorter mean time to recover is chasing something you can actually engineer, this quarter, with tools you already have.

Thamunkpillai 4 min read

There are two numbers reliability engineering could optimize for, and most organizations instinctively reach for the wrong one. MTBF — Mean Time Between Failures — measures how long the system runs before something breaks. MTTR — Mean Time to Recover — measures how long it takes to get back to healthy once it does. The instinct is to chase MTBF: build a system so robust it simply doesn't fail. The problem is that distributed systems fail. Networks partition, disks fill, dependencies time out, and no amount of hardening eliminates that — it just makes failures rarer and, often, stranger and harder to diagnose when they do happen.

MTTR is the number you can actually engineer down, reliably, with tools available today: better alerting, better runbooks, automated rollback, practiced incident response. And unlike MTBF, improving it doesn't require predicting every way the system might break — it just requires getting good at responding, regardless of cause.

Note

This isn't an argument that MTBF doesn't matter — obviously a system that fails less often is better than one that fails more. It's an argument about where to point limited engineering time. A system with an MTBF of 30 days and an MTTR of 2 minutes causes less user-visible pain over a year than one with an MTBF of 90 days and an MTTR of 45 minutes, even though the second system "fails less."

Doing the arithmetic

Text
System A: MTBF = 10 days, MTTR = 45 minutes
System B: MTBF = 7 days,  MTTR = 5 minutes
 
Over 90 days:
  A: ~9 failures × 45 min  = 405 minutes of downtime
  B: ~13 failures × 5 min  = 65 minutes of downtime

System B fails almost 50% more often and still causes six times less downtime, because recovery speed dominates the equation once failures are frequent enough to be a statistical certainty rather than a rare event. This is the whole argument in one calculation: past a certain point, "how often" matters less than "how fast you're back."

What actually moves MTTR

MTTR isn't a single lever — it's the sum of every stage between "something's wrong" and "it's fixed," and each stage has its own set of improvements:

StageWhat slows it downWhat speeds it up
DetectNobody's watching the right metricAlerting tied to user-facing SLIs, not just infra metrics
DiagnoseTracing across services doesn't existDistributed tracing (yesterday's article) pinpoints the hop
DecideNo one's sure who owns the fixClear on-call ownership, documented escalation
FixThe fix is manual and error-proneAutomated rollback, feature flags, runbooks with copy-paste commands
VerifyNo way to confirm the fix workedSLO dashboards that show recovery in real time
Bash
# The fastest MTTR improvement most teams haven't made:
# an automated rollback that doesn't wait for a human to decide to run it.
kubectl argo rollouts undo payment-service   # seconds, not a 20-minute
                                              # "let's get on a call" delay

Runbooks are an MTTR investment, not busywork

A runbook written calmly, in advance, with exact commands to copy-paste, routinely cuts diagnosis-and-fix time by more than half compared to reconstructing the fix from memory during an active incident. The five minutes it takes to write "if X alert fires, check Y, then run Z" pays for itself the first time it's used at 2am by someone who isn't the person who wrote it.

Where this fits the bigger reliability picture

MTTR is the practical, mechanical cousin of everything else this phase of the series has covered. Day 10 argued reliability has to be designed in. Day 11's error budget is the number that tells you how much unreliability is acceptable. Day 12's toil reduction is what frees up the time to actually build the automation that lowers MTTR. This article is where those ideas turn into a specific target: not "be more reliable" in the abstract, but "get the average incident from detection to resolution under N minutes," which is a number you can actually track sprint over sprint.

The teams with the best reputation for reliability rarely have the fewest incidents. They have the incidents nobody outside the team ever notices, because they were detected, diagnosed, and fixed before the impact accumulated into something visible. That's an MTTR story, not an MTBF one — and it's exactly what tomorrow's article on blameless postmortems is built to protect, by making sure every incident actually shortens tomorrow's MTTR instead of just repeating.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.