Cost and Architectural Decisions
What One More Nine Costs
In one line
Availability can be bought. The problem is that every additional 9 raises the cost by roughly an order of magnitude. So you must first decide "how much do we need," not "as safe as possible."
Why this was needed
If you make every service equally available, you over-invest in redundancy and standby resources even for unimportant features. Only by writing down separately, as RTO and RPO, the damage that downtime and data loss do to the business can you compare whether the extra cost of a recovery strategy is smaller than the loss it actually avoids.
How it works
What the 9s mean
| Availability | Allowed downtime per year | Per month |
|---|---|---|
| 99% | About 3.65 days | About 7 hours |
| 99.9% | About 8.8 hours | About 43 minutes |
| 99.95% | About 4.4 hours | About 22 minutes |
| 99.99% | About 52 minutes | About 4.3 minutes |
| 99.999% | About 5 minutes | About 26 seconds |
99.99% is 52 minutes a year. If each deployment stops things for 5 minutes, ten deployments a year use up the whole budget. So a high availability target automatically requires zero-downtime deployment. The moment you set the target, follow-on work appears.
RTO and RPO
| Meaning | Question | |
|---|---|---|
| RTO (recovery time objective) | How quickly the service is back | Recover within how many minutes or hours? |
| RPO (recovery point objective) | How recent the data that is saved is | How many minutes or hours of data may we lose? |
The two are independent. A setup that "recovers within 10 minutes but loses a day of data" is possible, and so is one that "loses not even a second of data but takes 6 hours to recover." Set them differently for each service — the RPO of a payment ledger must be close to 0, but a recommendation cache can lose a day.
Strategies and values
| Strategy | RTO | RPO | Relative cost |
|---|---|---|---|
| Backup and restore | Several hours to 1 day | The backup interval | Cheapest |
| Pilot light | Tens of minutes | Minutes | Low |
| Warm standby | A few minutes | Minutes | Medium |
| Active-active | Nearly 0 | Nearly 0 | Most expensive |
Pilot light is a method in which only the DB replication is kept on and everything else is kept off. In an emergency you turn on the rest. It is the strategy in which the value of pay-as-you-go shows best, and because it gives good value for the cost, it is often chosen in practice.
The order of deciding
- Divide services into tiers — not every one can be the top tier. Three tiers are usually enough (critical / important / general).
- Set the RTO/RPO of each tier as numbers — you must agree with the business side. "How much do we lose per hour of downtime?" is the basis.
- Choose the minimum configuration that satisfies those numbers — not a better one, but a sufficient one.
- Verify — an RTO that has not been rehearsed is wishful thinking.
Step 4 is, in practice, the one most often skipped. And in a real outage the RTO comes out at three times the plan.
What a rehearsal reveals
When you run a recovery drill, problems that were not in the documents come out.
- The recovery runbook is not up to date
- The person in charge does not have the permissions needed
- There is a backup but the decryption key cannot be found
- The recovery order of dependent services has not been decided
- DNS and certificates do not point to the new environment
If you discover these during an outage, the RTO will not be met.
Balancing against cost
If the availability investment exceeds the expected loss, it is over-investment.
기대 손실 = 연간 예상 중단 시간 × 시간당 손실
For a service that loses 1 million won per hour, raising 99.9% (8.8 hours a year) to 99.99% (52 minutes a year) saves about 8 hours, and that is worth 8 million won a year. If that configuration costs 30 million won a year, it is over-investment. If it is the other way round, you should of course do it.
Choosing to speed up recovery instead of buying availability
The phrase "we aim for 99.99%" usually comes up without any talk of budget. The cost of adding one more 9 is exponential, not linear, so to decide how far to buy, you have to look at the other side as well.
Calculate the value of downtime first. Add up revenue per hour, users who churn, penalty fees, and the people's time spent on recovery. Only with this number can you say whether "2 million won a month for redundancy" is expensive or cheap. For most internal systems, the calculation concludes that 99.9% (43 minutes a month) is enough.
For the same money, speeding up recovery is usually better. Redundancy prevents only certain kinds of failures (hardware failure, AZ outage). A large share of the outages you actually experience are bad deployments and configuration changes, and a redundant system deploys those changes identically to both sides. The ability to roll back within 5 minutes covers a wider range than redundancy.
If you have never measured recovery time, that number is a hope. How many minutes it takes to restore from a backup, and how long it takes to come up in another region, you only know by doing it. Usually it comes out at two or three times the value written in the documents. DNS TTL, image pulls, and the first traffic hitting an empty cache all eat up time.
Single points of failure are more often in people and process than in infrastructure. It is common for only one person to be able to bring that system back, for the recovery procedure to exist only in that person's head, or for no one to know the password of the emergency account. Before spending the redundancy budget, it is more valuable to run a recovery drill without that person, during a weekday daytime, just once.
Availability that is not measured is not managed. Availability measured from the inside differs from the availability users experience. There are situations where the server returns 200 but the user cannot see the screen. Even a single synthetic monitor that measures a real user flow from the outside narrows the gap between "we were at 99.95%" and "I failed three times."
What you see in the field
- Set a "99.99% availability target" and kept manual deployment → deployments use up the whole budget.
- Backups run every day but recovery has never been tried → the RTO is unknown.
- Made every service active-active → the cost was unaffordable, and it was eventually rolled back.
What to look at next
How to record these decisions so that later people can read them.