TT Lab
Get started
Learn Learning paths Courses

Capacity Planning and Change Management — Calculate When It Fills, Write Down When to Stop

Write Down When to Stop First — Change Types, Blast Radius, Rollback Criteria

Continue in TT Lab

In one line

A good change request writes "when to stop and how to roll back" before "what to do." Subtracting the time it takes to roll back backward from the work window gives the stop decision time, and that time becomes the trigger for rolling back. A rehearsal is the only way to learn before the work whether that plan was optimistic.

Why you need this

The Google SRE book writes that about 70% of outages in a production system are due to changes. You cannot eliminate changes, so all you can do is narrow the paths by which a change becomes an incident. The book gives three methods — rolling out progressively, detecting problems quickly and accurately, and safely rolling back when something goes wrong.

A change request and a work plan are tools for confirming these three before the work, on paper. If you leave the judgment of "it seems like a little more will do it" to an operator who is 20 minutes behind plan at three in the morning, they usually continue. If you write the stop criterion as a number in advance, you do not have to leave that judgment to a tired person.

How it works

Change types. The change enablement practice of ITIL 4 divides changes into three kinds.

Type Meaning Example
Standard A recurring change that is low risk and whose procedure is pre-approved Adjusting a configuration value using an approved runbook
Normal A change that is scheduled after assessment and approval Expanding a DB volume with a new runbook
Emergency A change that must be handled quickly because, if not done now, it will soon become an outage Replacing a certificate that expires in a few hours

The reason for dividing the types is to match the approval cost to the risk. If you take every change to a committee, people bypass the procedure, and if you just do every change, accidents happen. An emergency change is not without approval either; it is approved through a short path and the records are filled in afterward.

Impact scope. The impact of "restarting one DB" is not only that DB. You have to follow the dependencies all the way — the services that use that DB, and the services that use those services — for the notification targets and the verification targets to come out. If you notify by looking only at direct dependencies, the settlement report team two hops away sees a blank screen in the morning.

Work window and stop decision time. If the work window is 02:00–04:00 and rolling back needs 50 minutes in the worst case plus 5 minutes of buffer, 03:05 is the stop decision time. If the work is not finished, including verification, by this time, you have to start rolling back in order to return to the original state within the window. If the work is 60 minutes in the plan, it ends at 03:00, so it "fits" in the window. But what if the number 60 minutes is optimistic?

The rollback criterion must be measurable. "Roll back if a problem occurs" is not a criterion. Write it in time, rate, or count, such as "if the S5 check is not finished by 03:05" or "if the error rate exceeds 1% for more than 5 minutes." For each work step, also write the verification method by which you judge it done (verify). A step with no verification method is a step you do not know is finished.

Pre-check and notification. Before starting work, the list that checks whether the basis for rolling back is alive (did last night's backup succeed, is there space to take a snapshot) and whether the work target is the same as the plan's premises (version, capacity, path) is the pre-check. The rule is that if even one is off, you do not start the work. You also write in the request who notifies the owners of the affected services of the start, the end, and the rollback decision, and through what channel.

Rehearsal. You actually run the same procedure in a development or replica environment and record the start and end times of each step. Comparing the plan with the actual shows which step was optimistic. If the rehearsal passed the stop decision time, the real work is likely to pass it too — you first find a way to extend the window, split the work, or shorten the slow steps.

What it looks like in the field

In Korean SI and operations settings, these documents circulate under the names "work plan," "work result report," and "change request (CR)." The forms differ from company to company, but the reasons reviewers reject them are nearly the same — the rollback procedure is the single word "restore," the impacted services list has only direct dependencies, the sum of work time fills the window completely, and the pre-check has no backup verification.

After the work is finished, a result report remains. For each planned step, it writes the actual start and end times, the verification result, and what differed from the plan, and if it was rolled back, the time of that decision and its rationale. The plan for the next similar work starts from the measured times in this result report. If you estimate anew each time without a result report, the same step is late by the same amount every time.

The most expensive mistake is not calculating the rollback time. With work of 60 minutes in a window of 120 minutes, it looks ample, but if the rollback takes 50 minutes, the margin is only 5 minutes. With a rehearsal record, you know before the work that this margin is actually negative.

What you will do in the next lab

You classify the types of three change requests, and follow the impact scope of a DB restart through the service dependency list to the end. From the work plan table and the window, you calculate the work and rollback times and the stop decision time, and use them to write a JSON change request — the head (type, window, impact, pre-check) and the body (verification method for each step, and a measurable rollback criterion). Finally, you find the slowest step in the rehearsal record and judge whether it passed the stop time.