In Front of an Unfamiliar System
Draw the Blast Radius First
In one line
Before making any change, you must be able to draw on paper who and what is affected. If you cannot draw it, you are not yet ready to make that change. And that picture must not be a guess but something dug out of logs, sockets, configuration, and schedules.
Why this was needed
Say the order server's responses were slow, so you raised the connection pool from 10 to 30. It is one line of
configuration, and it looks easy to revert. But four of these servers were running,
and all four use the same database. The pool total goes from 40 to 120.
PostgreSQL's max_connections default is usually 100, and this value changes only when the server is
restarted. New connections start being rejected with sorry, too many clients already,
and the first thing to fall over is not the order server but the settlement batch that used the same DB.
The settlement batch was not drawn in the system diagram.
The reason outages that start with "we only need to change one line of configuration" keep recurring is that nobody checks what that one line touches. It is especially so on an unfamiliar system — to us it is a new system, but to the customer it is a 10-year-old system, and in that time dependencies that nobody recorded have piled up.
The four questions for drawing the impact scope
1. Who calls this component (upstream) Even when you edit a single configuration file, you need to know who sends requests to that process. The distribution of source IPs in the access log shows the callers over the past period, and the sockets attached right now show the current callers.
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head
ss -tn state established '( sport = :8080 )'
The first command shows the number of requests per IP in descending order. Do not ignore an IP with only a few requests. It may be a batch that runs once a day or an outside institution that comes only at the end of the month. In the second command, the peer address of a connection whose local port is the service port is the upstream.
2. What does this call (downstream) As a result of your change, load may concentrate on the downstream. A change that lengthens a timeout is the typical case — we think we gave some slack, but on the downstream the connections stay held longer. Pool size, retry count, and concurrency also push on the downstream in the same direction.
grep -nE '[a-z0-9.-]+:[0-9]{2,5}' /etc/app/*.yml
ss -tnp state established '( dport = :5432 )'
If you extract everything shaped like host:port from the configuration, you get the downstream candidates, and if you count how many connections are going out to that
port right now, you can estimate the number of connections after the change.
If there are several instances, do not forget that you must multiply one machine's value by the number of machines.
3. Does it share state Whether there is another system using the same DB, the same filesystem, or the same cache. Shared state is rarely drawn in the system diagram, yet accidents come from here. A quick way is to compare whether the same values come out of the configuration files of different services.
grep -hE 'db|dsn|path|dir|cache' service-a.yml service-b.yml | sort | uniq -d
uniq -d shows only the lines that appear two or more times. The DB address or
directory that appears here is a resource the two services use together.
4. When is it safe The time the batch runs, the days with deadlines, the settlement day. Even if it is technically safe, the timing being wrong makes it an accident. Schedules are not in only one place.
crontab -l; ls /etc/cron.d /etc/cron.daily
systemctl list-timers --all
The first five fields of a crontab are, in order, minute, hour, day, month, and weekday. 30 23 * * 5 is
"every Friday at 23:30". The systemd timer shows the next run time (NEXT) and the last run
time (LAST) together. However, even though what time the batch starts shows up here, what time it ends does not.
You can learn that only by asking the logs or the customer.
Write the rollback first
The first line of a change plan must not be the change itself but how to roll it back.
변경: /etc/app/config.yml 의 pool_size 10 → 30
백업: cp -p config.yml config.yml.2026-08-20
되돌림: cp -p config.yml.2026-08-20 config.yml && systemctl reload app
확인: curl -s localhost:8080/healthz 가 200, 에러율 5분간 관찰
소요: 되돌림 2분
If you can write "Rollback: 2 minutes", that change may be made. If you cannot, it is not yet time. When writing it, check three things.
- Back up with
cp -p. If you copy without-p, the backup becomes owned by the account that copied it, and its permissions are newly set according to the umask. If you put that backup back in place withmv, the owner and permissions of the configuration file change, and a second accident occurs: "I rolled back but the service can't read its configuration".-ppreserves the permissions, the owner (when you have the privilege), and the modification time together. - Check that
reloadreally rereads that value.systemctl reloadasks the service itself to reread its configuration, and if the service does not support reload, it fails. Even if it does, some values are read only at startup. In that case a restart is needed, and a restart cuts open connections, so the time required and the impact change. - Run the rollback once in advance. Change the configuration on a copy, then run the rollback
script, and compare with
difforsha256sumwhether the result is the same as the original. A rollback that has never been run even once is not a plan but a hope.
Decide the observation window
Decide in advance how long you will watch after the change. If you watch for 5 minutes and leave, a problem that blows up 10 minutes later is dropped from the list of candidate causes. Conversely, you cannot stay attached indefinitely either, so state it explicitly, such as "observe for 30 minutes, three metrics (error rate, latency, queue length)".
And write down the value before the change first. An error rate of 0.4% can be judged only if you know whether the normal level is 0.1% or 0.5%. If the time the batch runs falls inside the observation window, you cannot tell whether a change at that time is due to your change or to the batch, so move the window.
What it looks like in the field
- You lengthened the timeout and the downstream DB connections were exhausted → you did not look at the downstream. The symptom appears first not in our service but as connection refusals in another service that uses the same DB.
- You deployed at night and the settlement batch was running at that time → you did not ask about the timing. The crontab had only the start time, and how many hours the batch takes was not written anywhere.
- When you tried to roll back, nobody knew the original configuration → you edited without taking a backup.
- You rolled back with the backup but the service couldn't read its configuration → a backup taken without
-pwas put back withmv, and the owner and permissions changed. - You missed a caller that comes only once a day → you looked at only an hour of the access log. When looking for the upstream, look as widely as the log retention period allows.
What you will do in the lab that follows
You are given a customer system with no system diagram and no wiki. Two configuration files, an access log, the rotation configuration, and a crontab — that is all.
There you dig out the upstream, the downstream, the shared state, and the batch window with the commands in this section, and build a change plan in which you wrote the rollback first. The grader deliberately changes the configuration and then runs your rollback script to check that it returns to the original, so before submitting, run it once on a copy as described above.