Cost and Architectural Decisions
The Four Lines of the Bill
In one line
Cloud pricing boils down to four axes — hours left on, amount stored, amount moved, and number of calls. Once you know which axis the money flows out of, the direction of optimization is settled.
Why this was needed
Cloud pricing is not one server price but the sum of run time, storage capacity and requests, data transfer, and managed features. Reducing only one axis can increase the cost on another, so read the bill by the cause of each cost rather than by resource type.
How it works
The four axes
| Axis | Billing basis | How to reduce |
|---|---|---|
| Compute | Instances × hours | Downsize, turn off when not in use, commitment discounts |
| Storage | GB × month | Lifecycle policies, cheaper tiers, deleting |
| Data transfer | GB (mostly in the outbound direction) | Redesign paths, cache, compress |
| Requests/operations | Number of calls | Batch processing, cache |
The key is that "turn off when not in use" is on the first line. If you run development and staging environments only during weekday business hours, that alone cuts about 70% (only 9×5 = 45 hours out of 24×7 = 168 hours are used).
The structure of commitment discounts
Even for the same instance, the price differs greatly depending on how you buy it.
| Method | Discount | Conditions |
|---|---|---|
| On-demand | None | Start and stop whenever you like |
| Commitment (1–3 years) | 30–60% | You promise a term. You pay even if you do not use it |
| Spot/preemptible | 60–90% | Can be reclaimed at any time |
Spot offers an overwhelming discount but is reclaimed suddenly. So use it only for work that can be interrupted — batch jobs, CI runners, retryable queue consumers. Using it for a service that holds state causes incidents.
Commitments are the opposite. The standard approach is to commit only to the minimum baseline and fill the variable part with on-demand. If you commit at the peak, you lose money on whatever sits idle.
Storage tiers
Depending on access frequency, the price differs by 10× or more.
자주 접근 → 가끔 → 드묾 → 보관용(아카이브)
비쌈 가장 쌈, 꺼내는 데 시간·요금
Move data automatically with a lifecycle policy — a rule like "after 30 days, infrequent access; after 90 days, archive; after 365 days, delete." If you do not set this, logs pile up for years and stay in the most expensive tier.
Note: the archive tier charges a fee and takes time to retrieve from. Putting data you will read often there actually makes it more expensive.
Hidden charges
When you first look at a bill, a few lines are hard to understand.
- Allocated but unused things — volumes detached from an instance, static IPs that are not attached. Both are billed just for existing.
- Snapshot accumulation — if automatic snapshots have no retention policy, they keep piling up.
- Log storage — when log collection is on and retention is left unlimited.
- NAT gateway — the one covered in the earlier course.
- Cross-region replication — easy to turn on and forget.
Without tags you can do nothing
To reduce costs you first need to know who created what and why. Tags serve that purpose.
Owner = platform-team
Env = prod | stage | dev
Service = payments
CostCenter = 1042
A tagging rule has to be set up while resources are few for it to be followed. After hundreds have piled up without tags, applying it retroactively is practically impossible. And if you block the creation of untagged resources by policy, the rule enforces itself.
What you see in the field
- The development environment is left on at night, so compute cost ends up similar to production → look at the schedule first.
- A detached volume cannot be deleted → there is no owner tag, so you cannot tell whether it is in use.
- Costs rose after moving to archive → the total cost, including restore frequency, was not considered.
What to look at next
The one of the four axes that is most often underestimated — data transfer.