TT Lab
Get started
Learn Learning paths Courses

Capacity Planning and Change Management — Calculate When It Fills, Write Down When to Stop

Calculate When It Fills — Organic Growth, Steps, and the Order-By Date

Continue in TT Lab

In one line

The answer in capacity planning is not "what percent are we at now" but "by when do we have to order what." That date is fixed by three numbers — the speed at which usage normally grows (organic growth), the line you must not cross (the threshold), and the time from ordering to installation (the lead time). If you mix a one-time large increase (a step) into the trend, the forecast goes wrong altogether.

Why you need this

If you request expansion after the disk reaches 95%, it is already too late. Servers and storage take weeks for purchase orders, delivery, rack work, and securing a work window, and even in the cloud, reserved capacity and budget approval take time. So the capacity owner must answer not "are we fine now" but "what is the latest day we have to press the order button."

The introduction of the Google SRE book lists what capacity planning needs: an accurate organic demand forecast, separately accounting for inorganic demand such as a new feature launch or a customer migration, and looking ahead over a period longer than the procurement lead time. This module builds hands-on skill in those three things with a single disk usage table.

How it works

The trend line. Take the usage measured once a day, put x = the number of days since the first day and y = the usage (GB), and draw a least-squares line. The slope is given by the following formula.

기울기 = Σ (x − x̄)(y − ȳ) / Σ (x − x̄)²        단위: GB/일

Excel's SLOPE or a few lines of Python will do. The line is rough, but the wobble of less accumulating on weekends and more on weekdays has almost no effect on the slope once you gather a few months.

Remove the step. If you draw the table, there is a place where hundreds of GB were added in a single day. It is the day you moved data over from an old server, took on a new customer, or extended the log retention period. This is something that will not happen again. If you draw a line over the whole period, that step gets mixed into the slope and is translated as "it grows this much every day," and the forecast becomes far more pessimistic than reality. So you find the biggest one-day increase and recompute the slope from only the days after that. If there is a step planned for the future (next month's migration plan), do not mix it into the trend; add it separately.

Days remaining. If the threshold is 80%, the days remaining is 올림((용량 × 0.8 − 지금 사용량) ÷ 기울기) (that is, the ceiling of the capacity times 0.8 minus the current usage, divided by the slope). If you work out the time to 100% with the same formula, you can say "how long it holds even after crossing the threshold." The reason for rounding up is to be conservative about "how many days from now."

Order deadline. 임계 도달일 − 리드 타임 (the threshold date minus the lead time) is the latest day you must order. If that date has already passed or there are only a few days left, it should be the first line of the report.

Why the threshold is not 100%. A filesystem slows down well before it is full (fragmentation, reserved blocks), a database needs free space for housekeeping, and a forecast can always be wrong. The gap between the threshold and 100% is a buffer that absorbs forecast error and sudden growth. So the plan is fitted to the threshold, and the days remaining to 100% are used to say "how long it holds if the forecast is wrong."

Number Where it comes from If wrong
Organic slope The trend line after the step Mixing in the step leads to over-ordering, ignoring growth leads to a stockout
Threshold Operating criterion (headroom, the point where performance degrades) If too high, there is no time to respond
Lead time Purchase, delivery, work window If set too short, the order is late

What it looks like in the field

The most common mistake is to trust the "forecast" on the monitoring screen as is. Many tools draw a line from the slope of the last few days or hours. If backup files piled up all at once yesterday, "full in 3 days" appears, and when they are cleaned up the next day, "in 400 days" appears. A forecast whose period you have not checked is not a number but noise.

The second is not knowing about steps. If a report right after a migration says "the growth rate tripled," someone sets a year's budget from that number. Conversely, if you leave out the migration scheduled for next month, the trend is perfectly fine and one day suddenly there is not enough. That is why a capacity report writes the trend (organic) and the scheduled events (inorganic) separately.

Know where a straight line is wrong, too. For a service where growth itself speeds up as users increase, a straight line always warns late. In that case, write the slope of the recent segment next to the overall slope to reveal the acceleration, and add more margin to the lead time. Conversely, for data that nears equilibrium as old data is deleted, such as a log volume with a retention policy, a straight line is overly pessimistic. Before making the forecast model more sophisticated, the first thing is to state in one line in the report what assumptions the line was drawn under.

The third is doing the same calculation by hand every month. If you make it a small script that gives the same answer even when the table changes, next month it is a single command and it works as it is on other volumes.

What you will do in the next lab

You receive a 120-day disk usage table and the operating criteria (capacity, threshold, lead time). You write down the current state, compute the slope over the whole period, find the day the step occurred, and recompute the organic slope after it. You calculate the days remaining to the threshold and to 100%, and the order deadline, and finally you turn this calculation into a script that can run on any table. The grader also runs that script on a table it creates anew each time it grades.