Skip to content
SRE

Day 12: Toil Doesn't Build Value. It Steals Time.

Every hour spent manually restarting a service is an hour not spent making it stop needing restarts. Toil is the specific, measurable enemy SRE was invented to fight.

Thamunkpillai 4 min read

Not all operational work is toil, and the distinction matters more than it sounds like it should. Writing a runbook is work. Following that runbook by hand, at 2am, every time the same alert fires, because nobody's automated the fix yet — that's toil. The difference isn't effort, it's whether the work leaves anything behind. Engineering work compounds: today's automation script keeps paying off next month, next quarter, next year. Toil doesn't compound. It resets to zero the moment it's done and waits for the next occurrence.

Google's SRE book gives toil a precise definition, and it's worth using rather than the vague "annoying ops work" most teams settle for: toil is work that's manual, repetitive, automatable, tactical, devoid of enduring value, and scales linearly with service growth. All six properties matter, but the last one is the trap — toil that scales with growth means your operational burden grows in lockstep with your success, which is exactly backwards from what a healthy system should do.

Note

Not everything unpleasant is toil, and not everything toil-shaped needs to disappear entirely. Reading a paging alert and deciding it's a false positive is real judgment, even if it happens often. The test isn't "does this feel tedious" — it's "would automating this away lose anything a human was actually contributing." If the answer's no, it's toil.

What toil actually costs

SignalHigh-toil teamLow-toil team
On-call loadConstant interruptions, same fixes repeatedRare pages, mostly novel problems
Time allocationDays consumed by manual ticketsMost time spent building
Team moraleBurnout, context-switching fatigueFocus, ownership, energy for hard problems
Growth costOps burden scales with trafficOps burden stays roughly flat
Engineer retentionPeople leave for less firefighting elsewherePeople stay because the work is interesting

That last row is the one leadership tends to underweight. Toil doesn't just cost hours — it costs the specific engineers who are good enough to be trusted with the mundane, repetitive fixes, and who eventually leave for a role where they're trusted with something harder instead.

The identify-then-automate loop

Text
1. Notice the pattern
   "I've restarted this service manually four times this week."
 
2. Measure it
   How often? How long each time? Who's doing it?
 
3. Ask: is this actually automatable?
   Most toil is. Some genuinely needs human judgment — don't automate that part.
 
4. Automate the mechanical part
   A script, a runbook trigger, a self-healing check — whatever removes
   the repetition without removing the judgment call.
 
5. Verify it actually reduced toil
   Did the manual restarts stop, or did they just move somewhere else?

The step teams skip most often is the last one. It's easy to write automation that feels like progress without checking whether the underlying toil actually went away or just changed shape — a script that still needs someone to trigger it manually at 2am hasn't eliminated the interrupt, it's just made the fix faster once someone's already awake.

Google's own guideline

Google SRE targets keeping toil under roughly 50% of an SRE's time, with the rest reserved for engineering work that actually reduces future toil. The number itself is less important than the principle: if toil is allowed to grow unchecked, it will consume 100% of available time, because it scales with the system and engineering time doesn't scale to match on its own.

Why this connects directly to yesterday's error budget

Toil and error budgets are two sides of the same resource-allocation problem. An error budget tells you how much unreliability you're allowed to spend before you have to stop and fix things. Toil tells you how much of your time is being spent on work that doesn't reduce future unreliability at all. A team that's burning its error budget while also drowning in toil has a compounding problem: they don't have the time to fix the reliability issue, because all their time is going to manually working around the symptoms of it.

That's the argument for treating toil reduction as a first-class engineering priority rather than something squeezed in "when things calm down." Things calm down because toil goes away — not the other way around. Tomorrow's article covers the specific measurement gap that makes a lot of toil invisible in the first place: the difference between monitoring, which tells you something's wrong, and observability, which tells you why.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.