TT Lab
Get started
Learn Learning paths Courses

In Front of an Unfamiliar System

Draw the Blast Radius First

Continue in TT Lab

In one line

Before making any change, you must be able to draw on paper who and what is affected. If you cannot draw it, you are not yet ready to make that change. And that picture must not be a guess but something dug out of logs, sockets, configuration, and schedules.

Why this was needed

Say the order server's responses were slow, so you raised the connection pool from 10 to 30. It is one line of configuration, and it looks easy to revert. But four of these servers were running, and all four use the same database. The pool total goes from 40 to 120. PostgreSQL's max_connections default is usually 100, and this value changes only when the server is restarted. New connections start being rejected with sorry, too many clients already, and the first thing to fall over is not the order server but the settlement batch that used the same DB. The settlement batch was not drawn in the system diagram.

The reason outages that start with "we only need to change one line of configuration" keep recurring is that nobody checks what that one line touches. It is especially so on an unfamiliar system — to us it is a new system, but to the customer it is a 10-year-old system, and in that time dependencies that nobody recorded have piled up.

The four questions for drawing the impact scope

1. Who calls this component (upstream) Even when you edit a single configuration file, you need to know who sends requests to that process. The distribution of source IPs in the access log shows the callers over the past period, and the sockets attached right now show the current callers.

awk '{print $1}' access.log | sort | uniq -c | sort -rn | head
ss -tn state established '( sport = :8080 )'

The first command shows the number of requests per IP in descending order. Do not ignore an IP with only a few requests. It may be a batch that runs once a day or an outside institution that comes only at the end of the month. In the second command, the peer address of a connection whose local port is the service port is the upstream.

2. What does this call (downstream) As a result of your change, load may concentrate on the downstream. A change that lengthens a timeout is the typical case — we think we gave some slack, but on the downstream the connections stay held longer. Pool size, retry count, and concurrency also push on the downstream in the same direction.

grep -nE '[a-z0-9.-]+:[0-9]{2,5}' /etc/app/*.yml
ss -tnp state established '( dport = :5432 )'

If you extract everything shaped like host:port from the configuration, you get the downstream candidates, and if you count how many connections are going out to that port right now, you can estimate the number of connections after the change. If there are several instances, do not forget that you must multiply one machine's value by the number of machines.

3. Does it share state Whether there is another system using the same DB, the same filesystem, or the same cache. Shared state is rarely drawn in the system diagram, yet accidents come from here. A quick way is to compare whether the same values come out of the configuration files of different services.

grep -hE 'db|dsn|path|dir|cache' service-a.yml service-b.yml | sort | uniq -d

uniq -d shows only the lines that appear two or more times. The DB address or directory that appears here is a resource the two services use together.

4. When is it safe The time the batch runs, the days with deadlines, the settlement day. Even if it is technically safe, the timing being wrong makes it an accident. Schedules are not in only one place.

crontab -l; ls /etc/cron.d /etc/cron.daily
systemctl list-timers --all

The first five fields of a crontab are, in order, minute, hour, day, month, and weekday. 30 23 * * 5 is "every Friday at 23:30". The systemd timer shows the next run time (NEXT) and the last run time (LAST) together. However, even though what time the batch starts shows up here, what time it ends does not. You can learn that only by asking the logs or the customer.

Write the rollback first

The first line of a change plan must not be the change itself but how to roll it back.

변경:   /etc/app/config.yml 의 pool_size 10 → 30
백업:   cp -p config.yml config.yml.2026-08-20
되돌림: cp -p config.yml.2026-08-20 config.yml && systemctl reload app
확인:   curl -s localhost:8080/healthz 가 200, 에러율 5분간 관찰
소요:   되돌림 2분

If you can write "Rollback: 2 minutes", that change may be made. If you cannot, it is not yet time. When writing it, check three things.

Decide the observation window

Decide in advance how long you will watch after the change. If you watch for 5 minutes and leave, a problem that blows up 10 minutes later is dropped from the list of candidate causes. Conversely, you cannot stay attached indefinitely either, so state it explicitly, such as "observe for 30 minutes, three metrics (error rate, latency, queue length)".

And write down the value before the change first. An error rate of 0.4% can be judged only if you know whether the normal level is 0.1% or 0.5%. If the time the batch runs falls inside the observation window, you cannot tell whether a change at that time is due to your change or to the batch, so move the window.

What it looks like in the field

What you will do in the lab that follows

You are given a customer system with no system diagram and no wiki. Two configuration files, an access log, the rotation configuration, and a crontab — that is all.

There you dig out the upstream, the downstream, the shared state, and the batch window with the commands in this section, and build a change plan in which you wrote the rollback first. The grader deliberately changes the configuration and then runs your rollback script to check that it returns to the original, so before submitting, run it once on a copy as described above.