Make the Cloud Decision in Numbers
Goal
If you learn the cloud only as concepts, you cannot use it when you have to decide. This lab works out, in numbers, what you need for decisions.
- Where to put the region → the physical limit that comes from distance
- Whether to use a managed service → a comparison that includes people's time
- What you are responsible for protecting → the boundary of shared responsibility
Getting started
cp -r /opt/lab/cloudbasics/* .
python3 latency.py 11000 --real 180
python3 tco.py --self 400 --managed 900 --hours 20 --rate 60
Files
| File | What it does |
|---|---|
latency.py |
Distance → minimum round-trip time |
tco.py |
Cost comparison that includes people's time |
incidents.txt |
10 incidents. You fill this in |
Steps
- The physical limit →
01-physics.txt - Choosing a region →
02-region.md - AZs and regions →
03-az.md - Shared responsibility →
incidents.txt - TCO →
05-tco.txt - Break-even →
06-breakeven.md - When the cloud is not the answer →
07-notcloud.md - Summary →
08-notes.md
Notes
Distances and rates are approximations. What matters is the order of magnitude — 1 ms and 100 ms are not the same kind of number.
What money cannot reduce
Work out the minimum round-trip time for Seoul ↔ Virginia (about 11,000 km) and save it in 01-physics.txt. The measured value is roughly 180 ms — also write down the difference.
python3 latency.py 11000 --real 180
Light travels at 300,000 km/s in a vacuum, but in optical fiber it is about 200,000 km/s (refractive index 1.5).
110 ms is the physical limit. You cannot reduce it by making the instance bigger or spending more money. It is the only thing in the cloud you "cannot buy," so you have to solve it with architecture.
So where do you put the region
In 02-region.md, write where you would put the region for a service whose users are mostly in Korea, and justify it with numbers.
Use the numbers from step 1. For the Seoul region, how many ms to the users? For Virginia, how many?
Then think about how many round trips happen. If one page calls the API 5 times, that is 5 round trips. 110 ms × 5 = 550 ms, and that is time in which nothing was done.
Why AZs are separate
Compare the orders of magnitude of inter-AZ latency and inter-region latency in 03-az.md, and explain why you split across AZs.
AZs are different buildings in the same metropolitan area — they are tens of kilometers apart, so the round trip is about 1 ms. Between regions it is tens to hundreds of ms.
You buy something different for the same price. If you split across AZs, you survive the loss of one building while latency barely grows. If you split across regions, you survive the loss of a city, but latency becomes a hundred times greater.
Who protects what
Fill in incidents.txt by writing cloud or me for each of the 10 incidents. The format is 번호|사고|답 (number, incident, answer).
One criterion — what I can configure is my responsibility.
Bucket public-access settings, IAM key management and OS updates are all done by me. Hypervisor patching and physical access control are something I cannot touch, so they are the provider's responsibility.
Two confusing cases — an AZ power failure is the provider's responsibility, but having deployed only there is my responsibility. Minor patches of a managed DB are done by the provider.
Are managed services expensive
Compare the monthly cost of running it yourself and of the managed service including people's time, and save it in 05-tco.txt.
python3 tco.py --self 400 --managed 900 --hours 20 --rate 60
If you look only at infrastructure charges, the managed service looks more than twice as expensive. Once you include 20 hours of people's time, the managed service is $700 a month cheaper.
People's time does not show up in the ledger, so it keeps being treated as zero. And that time is usually spent at night, when an incident has happened.
Find the break-even
Work out how many hours a month or fewer it must take for running it yourself to be cheaper, and write in 06-breakeven.md whether that is realistic.
tco.py tells you the break-even. Under the conditions above it is 8.3 hours.
8 hours a month — for backup checks, patching, monitoring, capacity reviews and even one incident response. Is that realistic?
There is also the cost that the person who spends that time cannot do other work (opportunity cost). It does not appear in the ledger.
When the cloud is not the answer
Find at least two cases where the cloud is more expensive or unsuitable and write them in 07-notcloud.md, each with the reason.
Things to think about — large-scale compute with a steady, predictable load that you will use for 3 years or more, workloads that must send out large volumes of data (egress charges), data residency regulation, and special hardware.
It is not "the cloud is cheap" but "the cloud is flexible." If you don't need flexibility, there is no reason to pay for it.
Summarize
Write at least three lines in 08-notes.md: what money cannot buy, the criterion for shared responsibility, and what is easy to leave out when looking at managed-service costs.
The text must include 지연, 설정 and 사람 (these are the Korean words for latency, configuration and people; the grader checks for them).