Cost and Architectural Decisions
Not the Expensive Things — the Idle Ones
In one line
Cloud costs attach not to expensive things but to things that are switched on while doing nothing.
Why read the bill
Cloud costs are pay-as-you-go, so there is no approval step in the middle. Starting an instance needs no sign-off, and neither does stopping one. So starting things keeps happening, and stopping them almost never does.
After a few months the bill grows, and nobody knows what made it grow. The person who built it moved to another team, and finance sees only the total. If you say "let's cut costs" in this state, you end up touching whatever stands out first, and what stands out is usually small.
So the order is fixed — first split the bill into line items, then distinguish what that money is buying. Savings started without splitting are usually small, or they cut something that should not be cut.
You cannot do anything until you split by line item
If you start from "cloud costs are too high," it usually ends like this — you touch whatever stands out first, and what stands out is usually small.
First look at each item's share. Splitting the lab's bill gives this.
batch compute 1,401.60$ 43.9%
db managed_db 747.52$ 23.4%
web compute 525.60$ 16.5%
snapshots snapshot 210.00$ 6.6%
...
합계 3,191.17$
The top item alone is 44%. All the rest combined do not add up to it.
And the top item is usually idle
Read the memo for batch and it says this.
A batch job runs for 2 hours a day, but the instance stays on all the time.
It works 2 hours × 30 days = 60 hours, yet 730 hours are billed. 92% of the cost attaches to hours when nothing is being done.
It is hard to see from CPU utilization alone. You see an average of 11% and dismiss it as "it's a batch job, so that's normal." You need to look at when it works.
Distinguish what the money is buying
Not every cost is waste.
The 3000GB of cross-az traffic arises because the web tier and the DB are in different AZs. Move them into the same AZ and it disappears. But if that AZ goes down, both go down together.
That is not waste but the price paid for availability. To reduce it, you must first answer "is it acceptable for one AZ to go down?" Reducing it without that answer does not save money; it buys risk.
The standby DB (1 of the 2 db instances) is the same. It does nothing in normal times, but that is exactly what the money is for.
Costs that disappear by changing the path
This is the most pleasant kind. The traffic stays the same and only the cost disappears.
A NAT gateway has an hourly charge and a separate per-throughput charge. If traffic to object storage is passing through the NAT, using a gateway endpoint means that path no longer goes through the NAT.
The hourly charge remains — the NAT itself is still needed. But the throughput charge drops to 0.
The things nobody deletes
The 4200GB of snapshots is seven months of daily snapshots. Why has nobody deleted them?
Because nobody has set a retention period. To delete, you have to answer "how far back in time must we be able to go?", and without that answer, nobody can take responsibility for deleting.
The retention period is set not by cost but by recovery requirements. Once you write down that it is 30 days, everything beyond that is deleted automatically.
Where to start
Going in order of share seems like the right answer, but in practice it is better to start with what is not being used.
| Order | What | Why |
|---|---|---|
| 1 | Unused resources | Easy to find and no risk |
| 2 | Idle hours | The amount is large and it is easy to undo |
| 3 | Retention policy | Once you set it, it shrinks automatically |
| 4 | Changing paths | Requires a design change |
| 5 | Changing structure | Requires debate |
No one objects to item 1. The team learns the method while producing results, and that builds the trust to do the next one. If you start from item 5, it usually ends in debate.
What you see in the field
- The top item on the bill was a batch instance, and that batch ran only two hours a day. It ended with a one-line change setting the uptime to 60 hours.
- Someone saw a CPU average of 11%, said "it's a batch job, so that's normal," and left it for three years. What needed to be looked at was not utilization but when it works.
- Cross-AZ traffic was treated as waste, and the web tier and DB were moved into one AZ; when that AZ failed, the whole service stopped. The loss was larger than the amount saved.
- Seven months of snapshots sit there and nobody can delete them, because there is no document that answers whether deleting is okay.
- A team tried to start with the biggest item and two months passed in debates between teams. The team that started by cleaning up unused resources produced results in the same period.
Summary
- Split by line item and look at the shares
- Look at how much the top item actually works
- Distinguish waste from the cost of availability
- Find what disappears by changing the path
- Start with what is not used