TT Lab
Get started
Learn Learning paths Courses

Cost and Architectural Decisions

Read the Bill and Cut It

Continue in TT Lab

Goal

Cloud costs cannot be reduced if you do not know where they attach. And they usually attach not to "expensive things" but to "things that are switched on while doing nothing."

One team's monthly bill has been broken out by line item. You read it, fix it, and check how much it went down.

Getting started

cp -r /opt/lab/cost/* . && python3 cost.py

The total comes to roughly $3,191.

Files

File What it does
architecture.json Rates and resources. You will keep editing this
cost.py Calculates per line item. Do not edit it

Be sure to read the 메모 (memo) of each resource. The waste is written in the memos, not in the numbers.

Grading

From step 3 on, the grader calculates directly from your architecture.json. Submitting only numbers will not pass.

Steps

  1. Share per line item → 01-breakdown.txt
  2. Idle hours → 02-idle.md
  3. How much the fix saves → 03-fix-idle.txt
  4. Snapshot retention → 04-snapshot.txt
  5. Changing the path → 05-path.txt
  6. Cross-AZ traffic → 06-az.md
  7. Unused items → 07-unused.txt
  8. Summary → 08-notes.md

Notes

Rates differ by provider and region. What matters is the order of magnitude and "what is proportional to what."

Split the bill by line item

Copy /opt/lab/cost/, run python3 cost.py, and record the three largest items and the share of each in 01-breakdown.txt.

cp -r /opt/lab/cost/* . && python3 cost.py

This is the very first thing you do when reducing costs. Until you split by line item you cannot tell where to start, and people usually touch whatever stands out (which means the small things).

The total should come to roughly $3,191.

How much does the largest item actually work

Read the 메모 (memo) of the top item, compare the hours it actually works with the hours billed, and write it in 02-idle.md.

Look at architecture.json, specifically its 메모 (memo) field. It is a batch job that runs 2 hours a day, but the instance is on for 730 hours.

More than 90% of the cost attaches to hours when nothing is being done. This is the most common waste in cloud costs, and it is hard to see from CPU utilization alone — even when the average is low, people dismiss it as "that's just how the workload is."

How much does fixing it save

Leave architecture.json exactly as architecture.json, make a copy, and fix the batch so it runs only the hours it actually needs. For the file name, you may overwrite architecture.json. After fixing, record the result of python3 cost.py in 03-fix-idle.txt.

2 hours a day × 30 days = 60 hours. Change 가동시간 (uptime) from 730 to 60.

The grader checks by calculating directly from your architecture.json. Submitting only numbers will not pass.

See how much a one-line fix reduces the bill. That is the answer to "where do I start?"

The things nobody deleted

Look at the memo of the snapshot item, set a retention policy, reduce it accordingly, and write it in 04-snapshot.txt. Also write why you chose that period.

Seven months of daily snapshots have accumulated. With 30-day retention it shrinks to roughly 1/7.

The retention period is set not by cost but by recovery requirements. There is an answer to "how far back in time must we be able to go?", and without that answer nobody can delete — which is why seven months pile up.

Costs eliminated by changing the path

Look at the NAT item's memo, change the structure so that the throughput charge disappears, and record it in 05-path.txt.

A NAT has an hourly charge and a separate per-throughput charge. If most of the throughput is traffic to object storage, using a gateway endpoint means that path no longer goes through the NAT.

Try reducing 처리GB (throughput in GB). The traffic stays the same and only the path changes, and the cost disappears. The hourly charge remains — because the NAT itself is still needed.

Traffic that crosses AZs

Find in the memo why the cross-az item is 3000GB, and write how to reduce it in 06-az.md.

The web tier and the DB are in different AZs, so every request crosses an AZ. The per-GB price is small, but it adds up because the volume is large.

There is a trap here, though — if you move them into the same AZ, both go down together when that AZ dies. This item may not be "waste to cut" but the price paid for availability. Distinguish what the money is buying, and write it down.

What nobody uses

Skim the memos, find the resources nobody is using right now, delete them, and record the final total in 07-unused.txt.

A load balancer used three months ago is still running.

The amount is small. But things like this are the easiest to find and the lowest in risk. If you start here when beginning cost reduction, the team learns the method while producing results. If you start with the big things, it usually ends in debate.

Summarize

Write at least three lines in 08-notes.md: where to start, how to distinguish waste from the cost of availability, and an example of a cost that disappears by changing the path.

The text must include 비중, 가용성, and 경로 (share, availability, path).