Skip to content
SRE

Day 4: Automation Beats Heroics, Every Single Time

The engineer who can fix anything at 3 a.m. is not your most reliable system — they're your biggest single point of failure. Automation is what makes reliability survive someone taking a vacation.

Thamunkpillai 4 min read

Every engineering org has one. The person who gets paged and, within minutes, has diagnosed something nobody else on the team fully understands, run three commands from memory, and quietly fixed production before most people noticed anything was wrong. It feels like a superpower to witness. It is, structurally, one of the worst reliability postures an organization can have — because it means your system's actual uptime depends on one specific human being awake, available, and not on a plane.

Heroics are a symptom, not a strength

The instinct is to treat that engineer as an asset to protect and reward, which isn't wrong, exactly — but it misses the systemic problem their heroics are covering for. Every time a production issue gets solved by tribal knowledge instead of a documented, automated response, the organization learns nothing structurally. The fix lives in one person's head. The next time the same class of problem occurs, at 3 a.m., on a weekend, with that person unreachable, the organization is exactly as unprepared as it was the first time.

Warning

A team with a hero is a team with an undocumented single point of failure wearing a hoodie. The fact that it hasn't failed yet is not evidence it won't — it's evidence you haven't had bad enough timing yet.

What "automate it" actually means in practice

This isn't an argument against skilled, fast incident response — it's an argument for turning skilled, fast incident response into something the system can do, not just something one skilled, fast person can do. Concretely:

  • Runbooks that are actually run, not just written. A runbook nobody has followed under real pressure is a document, not a process. The gap between "documented" and "automated" is usually small — most runbooks are a sequence of commands someone already knows to type. Scripting that sequence turns tribal knowledge into a repeatable action.
  • Auto-remediation for known failure modes. If a specific alert has the same fix every time — restart a stuck pod, fail over a connection pool, clear a queue backlog — that fix belongs in automation, not in a human's muscle memory.
  • Self-healing built into the platform, not bolted onto the incident process. Kubernetes restarting a failed container is the simplest version of this. The more useful version is purpose-built automation for your specific, recurring failure modes — the ones generic infrastructure doesn't already handle.

The trade nobody states out loud

Heroics feel good to be recognized for, and they feel good to receive as a team — someone showed up, fixed it, everyone's relieved. Automation is comparatively unglamorous: writing a script that handles a problem before anyone notices it happened doesn't generate a Slack thread of gratitude. It generates silence, which is the actual goal, but it's a much harder thing to get organizational credit for.

A useful reframe

If an incident gets solved the same way twice by the same person, that's not consistency — that's a missed automation opportunity with a two-strike warning already served. The third occurrence is the one that happens while they're unreachable.

What this costs you if you don't fix it

The organizational risk compounds quietly. The hero becomes a bottleneck for anything requiring their judgment, not just incidents — they get pulled into every ambiguous problem because they're the fastest path to an answer, which leaves them no time to build the automation that would remove the need for them to be pulled in. Onboarding new engineers takes longer, because the actual operational knowledge of the system isn't written down anywhere a new hire can read it. And the hero, eventually, burns out or leaves — at which point the organization discovers, all at once, exactly how much undocumented judgment was walking around in one person's head.

No kubectl heroics. No manual changes at 2 a.m. that live only in one person's memory. Just Git, automation, and the kind of boring peace that comes from a system that can fix its own known problems without waiting for someone to wake up.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.