There are two numbers reliability engineering could optimize for, and most organizations instinctively reach for the wrong one. MTBF — Mean Time Between Failures — measures how long the system runs before something breaks. MTTR — Mean Time to Recover — measures how long it takes to get back to healthy once it does. The instinct is to chase MTBF: build a system so robust it simply doesn't fail. The problem is that distributed systems fail. Networks partition, disks fill, dependencies time out, and no amount of hardening eliminates that — it just makes failures rarer and, often, stranger and harder to diagnose when they do happen.
MTTR is the number you can actually engineer down, reliably, with tools available today: better alerting, better runbooks, automated rollback, practiced incident response. And unlike MTBF, improving it doesn't require predicting every way the system might break — it just requires getting good at responding, regardless of cause.
Note
This isn't an argument that MTBF doesn't matter — obviously a system that fails less often is better than one that fails more. It's an argument about where to point limited engineering time. A system with an MTBF of 30 days and an MTTR of 2 minutes causes less user-visible pain over a year than one with an MTBF of 90 days and an MTTR of 45 minutes, even though the second system "fails less."
Doing the arithmetic
System A: MTBF = 10 days, MTTR = 45 minutes
System B: MTBF = 7 days, MTTR = 5 minutes
Over 90 days:
A: ~9 failures × 45 min = 405 minutes of downtime
B: ~13 failures × 5 min = 65 minutes of downtimeSystem B fails almost 50% more often and still causes six times less downtime, because recovery speed dominates the equation once failures are frequent enough to be a statistical certainty rather than a rare event. This is the whole argument in one calculation: past a certain point, "how often" matters less than "how fast you're back."
What actually moves MTTR
MTTR isn't a single lever — it's the sum of every stage between "something's wrong" and "it's fixed," and each stage has its own set of improvements:
| Stage | What slows it down | What speeds it up |
|---|---|---|
| Detect | Nobody's watching the right metric | Alerting tied to user-facing SLIs, not just infra metrics |
| Diagnose | Tracing across services doesn't exist | Distributed tracing (yesterday's article) pinpoints the hop |
| Decide | No one's sure who owns the fix | Clear on-call ownership, documented escalation |
| Fix | The fix is manual and error-prone | Automated rollback, feature flags, runbooks with copy-paste commands |
| Verify | No way to confirm the fix worked | SLO dashboards that show recovery in real time |
# The fastest MTTR improvement most teams haven't made:
# an automated rollback that doesn't wait for a human to decide to run it.
kubectl argo rollouts undo payment-service # seconds, not a 20-minute
# "let's get on a call" delayRunbooks are an MTTR investment, not busywork
A runbook written calmly, in advance, with exact commands to copy-paste, routinely cuts diagnosis-and-fix time by more than half compared to reconstructing the fix from memory during an active incident. The five minutes it takes to write "if X alert fires, check Y, then run Z" pays for itself the first time it's used at 2am by someone who isn't the person who wrote it.
Where this fits the bigger reliability picture
MTTR is the practical, mechanical cousin of everything else this phase of the series has covered. Day 10 argued reliability has to be designed in. Day 11's error budget is the number that tells you how much unreliability is acceptable. Day 12's toil reduction is what frees up the time to actually build the automation that lowers MTTR. This article is where those ideas turn into a specific target: not "be more reliable" in the abstract, but "get the average incident from detection to resolution under N minutes," which is a number you can actually track sprint over sprint.
The teams with the best reputation for reliability rarely have the fewest incidents. They have the incidents nobody outside the team ever notices, because they were detected, diagnosed, and fixed before the impact accumulated into something visible. That's an MTTR story, not an MTBF one — and it's exactly what tomorrow's article on blameless postmortems is built to protect, by making sure every incident actually shortens tomorrow's MTTR instead of just repeating.





