Skip to content
SRE

Day 10: Reliability Isn't a Nice-to-Have. It's the Product.

A feature nobody can reach because the service is down isn't a feature. Kicking off the SRE Foundations phase of this series: why reliability has to be engineered in, not hoped for after launch.

Thamunkpillai 5 min read

The first nine days of this series were about DevOps culture and getting infrastructure under control — silos, automation, code as the source of truth. Today starts a different phase: Site Reliability Engineering, the discipline of making sure the systems that culture and automation produce actually stay up. It's worth being direct about why this gets nine days of its own rather than a paragraph inside the platform engineering material: reliability isn't a property that shows up automatically once you've automated enough. It has to be engineered in, on purpose, the same way a feature does.

Most teams don't disagree with that in principle. They disagree with it in practice, every time a roadmap gets prioritized and "make the checkout flow faster" beats "add a retry budget to the payment integration" because one of those has a visible demo and the other doesn't — until the day it does, in production, in front of customers.

Reliability is the feature users notice by its absence

Nobody opens a support ticket that says "thanks for the 99.95% uptime this quarter." They open one the moment the number dips below whatever they'd silently assumed was a given. That asymmetry is the whole argument: reliability doesn't earn credit when it's present, but it spends credit fast when it's missing.

Note

This is the same shape as security and performance — features that don't show up on a demo but absolutely show up in churn, trust, and word of mouth. The difference with reliability is how directly it compounds: an unreliable system doesn't just lose the user who hit the outage, it slows down every future feature because the team's now firefighting instead of building.

What it costs to treat reliability as an afterthought

Without reliability engineered inWith it
Incidents are frequent and mostly reactiveIncidents are rarer, and mostly caught before users notice
Every outage is an all-hands fire drillResponse is defined, rehearsed, and boring on purpose
Trust erodes a little more each incidentTrust compounds — the product is dependable by reputation
Engineers spend increasing time firefightingEngineers spend most of their time building
Downtime cost climbs as the business scalesDowntime cost is bounded by design, not luck

The multiplier in that last row is easy to underestimate early on. A two-hour outage costs a five-person startup an afternoon of apologies. The same two-hour outage at a company processing payments for ten thousand merchants is a different order of magnitude entirely — the cost of unreliability doesn't grow linearly with scale, it grows with however much now depends on the system staying up.

Reliability is designed, not hoped for

The instinct to treat reliability as something you add later — "we'll get to hardening this once we've shipped the feature" — comes from a real, well-intentioned trade-off: ship the thing people are waiting for. What that instinct misses is that most of what makes a system reliable is cheap to design in from the start and expensive to retrofit:

  • Design for failure — assume dependencies will time out, disks will fill, and nodes will die, because on a long enough timeline they will.
  • Observability — you cannot fix what you cannot see; instrumenting after an incident means the next one looks just as opaque as the last.
  • SLOs and error budgets — an explicit, agreed number for "how much unreliability is acceptable" turns an endless debate into a shared budget. Tomorrow's article is entirely about this.
  • Automation — most reliability work is removing manual steps from recovery, because manual steps are exactly where 2am mistakes happen.
  • Fast, practiced recovery — a system that fails and heals in ninety seconds behaves, from a user's perspective, almost like a system that didn't fail.
  • Blameless learning culture — an org where postmortems hunt for "why did the system allow this" instead of "who broke it" actually fixes root causes instead of just consequences. Day 16 of this series goes deep on that specifically.

None of these are exotic. They're ordinary engineering decisions, made deliberately, early, instead of accidentally, late, under pressure.

Reliability compounds like technical debt, just in the other direction

A small investment now — a retry with backoff, a health check that actually checks something meaningful, a runbook written before the incident instead of during it — is cheap today and valuable for years. Skipping it is also cheap today, and that's exactly the trap: the cost doesn't disappear, it just gets deferred to whichever engineer is on call when it finally comes due.

What the next few days build on this

Today's argument is deliberately abstract — reliability matters, here's why, here's the shape of what building it in looks like. The next several days in this series make it concrete: SLIs, SLOs and error budgets as the actual measurement system (Day 11), toil as the tax you pay for not automating (Day 12), monitoring versus observability as genuinely different capabilities, not synonyms (Day 13), and blameless postmortems as the mechanism that turns incidents into permanent fixes instead of a monthly rerun of the same page.

The through-line across all of it: reliability isn't a phase you complete once. It's a discipline you keep practicing, and the teams that treat it as a first-class feature from the start spend far less of their lives firefighting than the ones who bolt it on after the outage that finally made it unavoidable.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.