Cost and Architectural Decisions
What Do You Look At First When You Open the Bill
In one line
Cost reduction looks at the largest items first, and the lowest-risk actions first. Doing it the other way round yields little effect for the effort, and breaks the service.
Why this was needed
If you start with discount contracts or architecture changes, you may lock in even unused resources for a long period, or raise service risk for the sake of a small saving. A sequence that moves from removing waste, which is easy to undo, to right-sizing, which requires measurement, protects both cost and operational stability.
How it works
The order
1. 안 쓰는 것 지우기 ← 위험 0, 효과 즉시
2. 안 쓸 때 끄기 ← 위험 낮음, 효과 큼
3. 크기 줄이기(right-sizing) ← 측정 필요
4. 저장 계층 옮기기 ← 수명주기 정책
5. 아키텍처 바꾸기 ← 효과 크지만 오래 걸림
6. 약정 할인 구매 ← 위 5개를 끝낸 뒤에 한다
It is important that item 6 comes last. Buying a commitment while leaving the waste in place amounts to locking the waste into a 3-year contract. Clean up first, and attach the commitment to the cleaned-up baseline.
1. What is not used
Make it a checklist and go through it every quarter.
- Detached volumes (unattached)
- Unused static IPs
- Load balancers with no targets
- Old snapshots and AMIs
- Empty database instances
- Log groups with no logs for 30 days
- Tables no one queries
This alone often cuts 10–20%.
2. Schedules
Turn non-production environments on only during business hours. It is simple to do and the effect is large.
The thing to watch is to send a notification before turning them off — someone may be
working late. And set up an exception tag (AlwaysOn=true) to keep what needs to stay.
3. Right-sizing
Doing it without metrics causes incidents. At a minimum, look at it like this.
- CPU — if the peak stays under 40% for two weeks, it is a candidate to drop one size
- Memory — this is often the real constraint. Always look at it together
- Burst credits — for burstable instances, you must check whether the credits are exhausted. Even with a low average CPU, if the credits are at zero it is already performance-constrained.
Reduce one step at a time and observe. If you reduce two steps at once, you have no basis to decide on when rolling back.
4. Storage lifecycle
Apply policies to logs, backups, and images.
0~30일 자주 접근 계층
30~90일 저빈도 계층
90~365일 아카이브
365일 이후 삭제
Before adding a deletion rule, you must check the retention requirements. Data that regulation requires you to keep for years may be mixed in.
5. Architecture
These are the things covered earlier — introducing endpoints, CDNs, AZ placement, moving to or away from managed services. The effect is large but the lead time is long, so run it in parallel with the first four steps.
How to keep the savings
Even if you cut once, in six months it goes back to how it was. Keeping it requires a process.
- Budgets and alerts — set budgets per team, and alert when an overrun is projected
- Enforce tags — block the creation of untagged resources
- Regular reviews — look together at the top 10 items every quarter
- Visibility — let teams see their own costs. If it is not visible, it does not shrink
The last one is the key. If only the finance team looks at costs, nobody reduces them. The person who built it has to be able to see the price.
What you see in the field
- Bought a commitment first and then could not change the architecture → the order was wrong.
- Reduced the size by two steps and caused an outage → the decision was made without metrics.
- Reverted within six months of the savings → it was a one-off task, not a process.
What you will do in the following lab
You check this order with real numbers. If you try to downsize based on metrics alone, three of the five machines are blocked, and what blocks them is not the CPU but memory and burst credits. And you calculate for yourself how much more you pay over 3 years if you bought the commitment before cleaning up.
What to check before reducing
The five steps above tell you what to reduce, but whether it is okay to reduce it must be looked at separately. Incidents caused by skipping this check are often more expensive than what the savings saved.
Is it really unused? A metric of 0 and nobody using it are different things. A settlement batch that runs once a month, a reporting instance used every quarter, and resources turned off for disaster recovery all look like 0 in normal times. The observation period must be longer than the resource's longest cycle for you to decide; otherwise, ask the owner before deleting.
Do you know who uses it? Resources with no tags, whose owner you do not know, are the hardest to delete. In that case, instead of deleting, turn it off first. Leave it a few days, and if nobody comes looking for it, then delete it. Just inserting one reversible step makes this decision much easier.
Is anything depending on it? Before lowering a storage tier, you must check whether the readers can tolerate the latency, and before changing an instance type, you must check whether a license or an IP is pinned to it.
Have you decided when to roll back? After reducing the size, write down in advance what you will watch to decide on rolling back, and for how many days you will watch. The earlier case in which reducing two steps at once caused an outage comes up, and the essence of that incident is not the big step but that no rollback condition was set.
Finally, record the amount saved. If a record remains of what was reduced, when, and how much it saved, you do not have to start the same discussion from scratch next quarter, and above all, when you have to roll back you can state the cost of doing so in numbers.
What to look at next
Availability has a price too. How to set that price with RTO/RPO.