Day 4 of 100
The Cloud Bill Suddenly Increased 35%
π₯ Problem Statement
The monthly billing summary lands, and it's 35% higher than last month with no corresponding launch, no traffic spike anyone remembers, and no new product line that would explain it. Finance wants an explanation this week.
Someone pulls up the top-line cost breakdown and immediately proposes the fastest fix available: "Let's shut down the expensive resources." A few non-production instances are the obvious first target β nobody's actively watching them day to day, and stopping them would show an immediate, visible drop on next month's bill.
ποΈ Environment
- Cloud: GCP
- Workloads: Compute + Storage + Logging + Monitoring + Data Transfer
- Environments: Production + Non-Production, shared billing account
- Cost visibility: monthly billing export, no per-team budgets yet
π€ Your Challenge
What would you investigate before deleting or stopping anything?
- Does a top-line cost increase tell you which specific resource or team caused it, or just that something changed?
- What's the difference between a resource that's expensive and a resource that's newly expensive?
- Could a 35% increase come from something that isn't compute at all β logging volume, data transfer, storage class?
- What would you want to know about a resource's owner and purpose before shutting it down?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
"35% higher" is a single aggregate number, and it tells you almost nothing about what changed β only that something did. Jumping straight to "shut down the expensive resources" skips the step that actually matters: the biggest line item on the bill isn't necessarily the thing that changed. A resource that's always been expensive and stayed flat this month isn't the cause of a 35% increase. Something that moved is.
This distinction matters because acting on "biggest" instead of "newly biggest" risks shutting down a stable, load-bearing resource while the actual cause β something new, misconfigured, or newly scaled β keeps running untouched.
Step 2 β Identify the likely bottleneck
A month-over-month cost jump this size, with no known launch or traffic change, tends to come from one of a fairly short list: a newly idle or forgotten resource left running (a test environment spun up and never torn down), a logging or monitoring pipeline that started ingesting significantly more volume (a verbose log level left on, a new service instrumented without sampling), a change in storage class or data transfer pattern (data moved to a more expensive tier, or a new cross-region transfer path introduced), or a genuine, legitimate scaling event that just hasn't been communicated to whoever owns the budget conversation yet.
Without a cost breakdown by service and by time, "shut down the expensive resources" is guessing at which of these it is β and non-production instances being an easy, low-drama target to stop doesn't mean they're actually the cause.
Step 3 β Investigation
- Cost breakdown by service, compared month over month. Which specific line items grew, and by how much? This turns "35% higher" into "logging costs are up 4x" or "data transfer in us-east1 tripled" β an actual lead to chase.
- Cost breakdown by project/label/team, if tagging discipline allows it. Even partial tagging usually narrows the field significantly.
- Usage graphs for the top movers, not just cost. A cost increase with flat usage points at a pricing or tier change; a cost increase with matching usage growth points at genuine new load.
- Logging ingestion volume specifically β this is one of the most common, least obvious sources of a cost spike, especially if a recent deploy changed a log level or added new instrumentation without a sampling or retention policy.
- Ownership of the specific resources actually driving the increase β not the non-production instances that happen to be visible and easy to stop, but whatever the breakdown in the first bullet actually points to.
Step 4 β Recommended action
Don't stop anything until the breakdown identifies what actually grew. If it turns out to be a forgotten non-production environment, stopping it is the right call β but that's a conclusion to reach from evidence, not a first move made because it's the fastest visible action available under pressure from Finance.
If the cause turns out to be something load-bearing β genuine production scaling, a new feature that's legitimately using more infrastructure β the right response isn't to cut it, it's to right-size it, communicate the cost to whoever owns that budget, and decide deliberately whether the increase is worth the value it's producing. Cost optimization done under time pressure, aimed at whatever's easiest to touch rather than what's actually responsible, has a real chance of cutting something that matters while missing the real cause entirely β which means next month's bill looks the same, minus whatever got broken by shutting down the wrong thing.
Step 5 β Engineering lesson
A big bill and a big change in the bill are different problems, and only one of them is what actually needs investigating here. The fastest available action β stop the visibly expensive stuff β feels productive under Finance pressure, but "visibly expensive" and "newly expensive" answer different questions, and only the second one explains this month's number.
Good cost optimization looks a lot like good incident response: find out what actually changed before you change something else in response.
π§ Todayβs Takeaway
Optimize based on evidence, not just the largest bill.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

