TT Lab
Get started
Learn Learning paths Courses

Cost and Architectural Decisions

Not the Expensive Things — the Idle Ones

Continue in TT Lab

In one line

Cloud costs attach not to expensive things but to things that are switched on while doing nothing.

Why read the bill

Cloud costs are pay-as-you-go, so there is no approval step in the middle. Starting an instance needs no sign-off, and neither does stopping one. So starting things keeps happening, and stopping them almost never does.

After a few months the bill grows, and nobody knows what made it grow. The person who built it moved to another team, and finance sees only the total. If you say "let's cut costs" in this state, you end up touching whatever stands out first, and what stands out is usually small.

So the order is fixed — first split the bill into line items, then distinguish what that money is buying. Savings started without splitting are usually small, or they cut something that should not be cut.

You cannot do anything until you split by line item

If you start from "cloud costs are too high," it usually ends like this — you touch whatever stands out first, and what stands out is usually small.

First look at each item's share. Splitting the lab's bill gives this.

batch        compute          1,401.60$   43.9%
db           managed_db         747.52$   23.4%
web          compute            525.60$   16.5%
snapshots    snapshot           210.00$    6.6%
...
합계                          3,191.17$

The top item alone is 44%. All the rest combined do not add up to it.

And the top item is usually idle

Read the memo for batch and it says this.

A batch job runs for 2 hours a day, but the instance stays on all the time.

It works 2 hours × 30 days = 60 hours, yet 730 hours are billed. 92% of the cost attaches to hours when nothing is being done.

It is hard to see from CPU utilization alone. You see an average of 11% and dismiss it as "it's a batch job, so that's normal." You need to look at when it works.

Distinguish what the money is buying

Not every cost is waste.

The 3000GB of cross-az traffic arises because the web tier and the DB are in different AZs. Move them into the same AZ and it disappears. But if that AZ goes down, both go down together.

That is not waste but the price paid for availability. To reduce it, you must first answer "is it acceptable for one AZ to go down?" Reducing it without that answer does not save money; it buys risk.

The standby DB (1 of the 2 db instances) is the same. It does nothing in normal times, but that is exactly what the money is for.

Costs that disappear by changing the path

This is the most pleasant kind. The traffic stays the same and only the cost disappears.

A NAT gateway has an hourly charge and a separate per-throughput charge. If traffic to object storage is passing through the NAT, using a gateway endpoint means that path no longer goes through the NAT.

The hourly charge remains — the NAT itself is still needed. But the throughput charge drops to 0.

The things nobody deletes

The 4200GB of snapshots is seven months of daily snapshots. Why has nobody deleted them?

Because nobody has set a retention period. To delete, you have to answer "how far back in time must we be able to go?", and without that answer, nobody can take responsibility for deleting.

The retention period is set not by cost but by recovery requirements. Once you write down that it is 30 days, everything beyond that is deleted automatically.

Where to start

Going in order of share seems like the right answer, but in practice it is better to start with what is not being used.

Order What Why
1 Unused resources Easy to find and no risk
2 Idle hours The amount is large and it is easy to undo
3 Retention policy Once you set it, it shrinks automatically
4 Changing paths Requires a design change
5 Changing structure Requires debate

No one objects to item 1. The team learns the method while producing results, and that builds the trust to do the next one. If you start from item 5, it usually ends in debate.

What you see in the field

Summary