TT Lab
Get started
Learn Learning paths Courses

Cost and Architectural Decisions

What One More Nine Costs

Continue in TT Lab

In one line

Availability can be bought. The problem is that every additional 9 raises the cost by roughly an order of magnitude. So you must first decide "how much do we need," not "as safe as possible."

Why this was needed

If you make every service equally available, you over-invest in redundancy and standby resources even for unimportant features. Only by writing down separately, as RTO and RPO, the damage that downtime and data loss do to the business can you compare whether the extra cost of a recovery strategy is smaller than the loss it actually avoids.

How it works

What the 9s mean

Availability Allowed downtime per year Per month
99% About 3.65 days About 7 hours
99.9% About 8.8 hours About 43 minutes
99.95% About 4.4 hours About 22 minutes
99.99% About 52 minutes About 4.3 minutes
99.999% About 5 minutes About 26 seconds

99.99% is 52 minutes a year. If each deployment stops things for 5 minutes, ten deployments a year use up the whole budget. So a high availability target automatically requires zero-downtime deployment. The moment you set the target, follow-on work appears.

RTO and RPO

Meaning Question
RTO (recovery time objective) How quickly the service is back Recover within how many minutes or hours?
RPO (recovery point objective) How recent the data that is saved is How many minutes or hours of data may we lose?

The two are independent. A setup that "recovers within 10 minutes but loses a day of data" is possible, and so is one that "loses not even a second of data but takes 6 hours to recover." Set them differently for each service — the RPO of a payment ledger must be close to 0, but a recommendation cache can lose a day.

Strategies and values

Strategy RTO RPO Relative cost
Backup and restore Several hours to 1 day The backup interval Cheapest
Pilot light Tens of minutes Minutes Low
Warm standby A few minutes Minutes Medium
Active-active Nearly 0 Nearly 0 Most expensive

Pilot light is a method in which only the DB replication is kept on and everything else is kept off. In an emergency you turn on the rest. It is the strategy in which the value of pay-as-you-go shows best, and because it gives good value for the cost, it is often chosen in practice.

The order of deciding

  1. Divide services into tiers — not every one can be the top tier. Three tiers are usually enough (critical / important / general).
  2. Set the RTO/RPO of each tier as numbers — you must agree with the business side. "How much do we lose per hour of downtime?" is the basis.
  3. Choose the minimum configuration that satisfies those numbers — not a better one, but a sufficient one.
  4. Verify — an RTO that has not been rehearsed is wishful thinking.

Step 4 is, in practice, the one most often skipped. And in a real outage the RTO comes out at three times the plan.

What a rehearsal reveals

When you run a recovery drill, problems that were not in the documents come out.

If you discover these during an outage, the RTO will not be met.

Balancing against cost

If the availability investment exceeds the expected loss, it is over-investment.

기대 손실 = 연간 예상 중단 시간 × 시간당 손실

For a service that loses 1 million won per hour, raising 99.9% (8.8 hours a year) to 99.99% (52 minutes a year) saves about 8 hours, and that is worth 8 million won a year. If that configuration costs 30 million won a year, it is over-investment. If it is the other way round, you should of course do it.

Choosing to speed up recovery instead of buying availability

The phrase "we aim for 99.99%" usually comes up without any talk of budget. The cost of adding one more 9 is exponential, not linear, so to decide how far to buy, you have to look at the other side as well.

Calculate the value of downtime first. Add up revenue per hour, users who churn, penalty fees, and the people's time spent on recovery. Only with this number can you say whether "2 million won a month for redundancy" is expensive or cheap. For most internal systems, the calculation concludes that 99.9% (43 minutes a month) is enough.

For the same money, speeding up recovery is usually better. Redundancy prevents only certain kinds of failures (hardware failure, AZ outage). A large share of the outages you actually experience are bad deployments and configuration changes, and a redundant system deploys those changes identically to both sides. The ability to roll back within 5 minutes covers a wider range than redundancy.

If you have never measured recovery time, that number is a hope. How many minutes it takes to restore from a backup, and how long it takes to come up in another region, you only know by doing it. Usually it comes out at two or three times the value written in the documents. DNS TTL, image pulls, and the first traffic hitting an empty cache all eat up time.

Single points of failure are more often in people and process than in infrastructure. It is common for only one person to be able to bring that system back, for the recovery procedure to exist only in that person's head, or for no one to know the password of the emergency account. Before spending the redundancy budget, it is more valuable to run a recovery drill without that person, during a weekday daytime, just once.

Availability that is not measured is not managed. Availability measured from the inside differs from the availability users experience. There are situations where the server returns 200 but the user cannot see the screen. Even a single synthetic monitor that measures a real user flow from the outside narrows the gap between "we were at 99.95%" and "I failed three times."

What you see in the field

What to look at next

How to record these decisions so that later people can read them.