Cost and Architectural Decisions
Commit First and the Waste Becomes a 3-Year Contract
Goal
In the reading you saw the order. Delete what is unused first, and buy commitment discounts last. Here you check that order with real numbers.
You experience two things in particular by hand.
- Downsizing based only on CPU causes incidents. Memory and burst credits are often the real constraint.
- If you buy a commitment before cleaning up, you lock the waste into a 3-year contract. You calculate in dollars how much more you pay.
Getting started
mkdir -p /root/rightsize
cp -r /opt/lab/rightsize/* /root/rightsize/
cd /root/rightsize
python3 plan.py
The total comes to $4,607.76.
Files
| File | What it does |
|---|---|
fleet.json |
The instance list and the last 14 days of metrics. Leave the original as it is |
plan.py |
Calculates one month of cost. Do not edit it |
after.json |
A copy of fleet.json that you will edit |
Be sure to read the metrics. Looking only at CPU최대_14일 (14-day peak CPU), it seems all five machines
could be downsized, but for three of them another metric is blocking it.
Grading
For steps 3 and 5, the grader calculates directly from your after.json.
Submitting only numbers will not pass, and downsizing something that must not be downsized is also caught.
What to produce
01-baseline.txt 지금 얼마이고 어디에 몰려 있나
02-signals.md 줄일 수 있는 것과 못 줄이는 것, 그리고 그 근거 지표
after.json 실제로 고친 구성 (3단계와 5단계가 함께 쓴다)
04-blocked.md 무엇이 막고 있고, 무엇을 재면 풀리나
06-commit.txt 정리된 기준선 위에서의 약정 금액
07-order.txt 정리 전에 샀다면 얼마를 더 내나
08-keep.md 줄인 것을 유지하는 절차
How much is it now
Run plan.py and write the current total and the three largest items, with their shares, in 01-baseline.txt.
python3 plan.py calculates per line item and prints the shares as well.
Writing only the amounts is not enough. You need the shares to decide where to start. Cutting a 1% item in half gives only 0.5%.
What is blocking the size reduction
Sort the five compute instances into those that can be downsized and those that cannot, and write the supporting metric for each in 02-signals.md.
In fleet.json, look at CPU최대_14일, 메모리최대_14일, and 버스트크레딧최소_14일 together (14-day peak CPU, 14-day peak memory, and 14-day minimum burst credits).
A low CPU alone is not sufficient grounds. If memory is full, reducing vCPU also reduces memory, and the workload no longer fits. For burstable instances, even with a low average CPU, if the credits are at zero they are already performance-throttled.
Reduce only one step at a time
Copy fleet.json to after.json, then downsize one step at a time only what the metrics allow. Do not touch those that are blocked.
The size ladder is at fleet.json, under 크기_사다리 (size ladder) — s → m → l → xl.
You must not reduce two steps at once. When a problem occurs, you have no basis to tell which step was the cause. Reduce one step, observe, and then decide again.
The grader reads after.json directly to see which ones you reduced and by how much.
Do not just leave the blocked ones alone
For each of the three you could not downsize, write in 04-blocked.md what is blocking it and what you would measure or change to unblock it. Include the metric values.
If you stop at "we cannot reduce it because memory is high," it will be in the same place next quarter.
Ask about each one. Why is this value so high, what would have to change in the code or structure to lower it, and what would you measure to be able to decide?
For a burstable instance, the answer may not be to reduce it but to measure whether to move it to a fixed-performance family.
Delete what is unused first
Find the resources nobody has used in 14 days and delete them from after.json. Keep the ones in use.
Look at the metrics of the items other than compute — 대상수, 요청수_14일, 연결됨, 연결수_14일 (target count, 14-day request count, attached, 14-day connection count).
This is item 1 in the order from the reading. Zero risk, immediate effect. The reason to do it before right-sizing is that it is easy to undo and needs no measurement.
After deleting them all, the total from python3 plan.py after.json becomes $3,153.60.
Calculate the commitment on the cleaned-up baseline
Apply the commitment discount to the compute amount after cleanup, and write the monthly amount and the 3-year total in 06-commit.txt. Also write which baseline you calculated on.
The commitment terms are in fleet.json, under 약정 (commitment) — a 40% discount, 36 months.
A commitment means paying the committed hourly amount for all three years. Even if you actually use less, there is no refund. So what you take as the baseline is exactly your three-year spend.
If you had bought before cleaning up
If you had bought the same commitment on the compute amount from before cleanup, how much would it be over 3 years? Write that, and the difference from after cleanup, in 07-order.txt. Also write why that money simply disappears.
Compute before cleanup was $4,204.80. Calculate with the same discount rate and term, and subtract the result of step 6.
This difference is the answer of this lab. It shows in numbers why the reading said "commitment last." If you buy the commitment first, that amount keeps going out even if you delete resources later.
To keep it that way half a year from now
In 08-keep.md, write the order in which to make changes and the process for keeping what you reduced. Include who should be able to see the costs.
Even if you cut once, in six months it goes back to how it was. That is the difference between a one-off task and a process.
The reading lists four — budgets and alerts, enforced tags, regular reviews, and visibility. Of these, visibility is the key. If the person who built it cannot see their own costs, nobody reduces them.
Write it so that next quarter someone else can run it exactly as written using only this document.