In Front of an Unfamiliar System
Draw the Map Before You Touch Anything
Goal
On a customer system you are seeing for the first time, you draw a map before touching anything. What to request, how far the investigation scope extends, what touches what, and when it is safe. And you write the way to roll back first.
Environment
A set of customer systems is in /opt/data/site. To you it is a new
system, but to the customer it is a 10-year-old system, and in that time dependencies
that nobody recorded have piled up. You make a copy and then work on that copy.
cp -r /opt/data/site /root/site
cd /root/site
ls -R | head -40
That is all you received. There is no system diagram and no wiki.
What you will build
inventory.txt 무엇이 있는지 + 첫날 요청할 것
retention.txt 로그 보존 기간과 그것이 조사에 뜻하는 것
upstream.txt 이 서비스를 부르는 쪽
downstream.txt 이 서비스가 부르는 쪽
shared.txt 구성도에 없는 공유 상태
timing.txt 언제가 안전하지 않은가
plan.md 변경 계획 — 첫 줄이 되돌리는 방법
rollback.sh 실제로 되돌리는 스크립트
snapshot.sh 재시작 전에 상태를 남기는 스크립트
How it is graded
In steps 7 and 8, the grader runs your scripts directly.
rollback.sh 채점기가 설정을 일부러 바꾼 뒤 돌려서, 원본으로 돌아오는지 봅니다
snapshot.sh 프로세스·메모리·열린 파일·최근 로그 네 가지가 시각과 함께 나오는지 봅니다
A script that only talks and does nothing fails. And in every case the grading puts your configuration back as it was.
Steps
- Write down what is there, and make a list of what to request at the same time on the first day. Approval systems are mostly serial, so you have to throw everything on the first day to have it run in parallel.
- Find the log retention period. That value decides the investigation scope.
- Find all of the upstream from the source IPs in the access log. Also how many requests from each.
- Find the downstream in the configuration, and write down what changes put the downstream at risk.
- Compare the two configuration files and find the shared state.
- Find the batch window in the crontab and decide a safe time.
plan.mdandrollback.sh— the first line of a change plan is not the change but how to roll it back. Also the time the rollback takes and the observation window.snapshot.sh— leave the state before a restart.
Notes
The reason for step 8 is what is most often forgotten in this lab. A restart erases the evidence along with the symptom. Even if a restart is necessary, if you leave a snapshot before it, that one file will later become half of the report.
What is there and what to request
Write down what is there, and make a list of what to request at the same time on the first day. Approval systems are mostly serial, so you have to throw everything on the first day to have it run in parallel.
Start by looking at how many configuration files there are. And the list to request on the first day — approval systems are mostly serial, so the account goes up only after the VPN is done, and the permission comes only after the account comes out. You have to throw everything at once on the first day to have it run in parallel.
The retention period decides the investigation scope
Find the log retention period. That value decides the investigation scope.
It is written in etc/logrotate.conf. Check whether it also matches the number of rotated files. If a report comes in that goes beyond that value, it cannot be confirmed from the logs.
Who calls this
Find all of the upstream from the source IPs in the access log. Also how many requests from each.
Count the source IPs in the access log. Do not stop after finding just one — the places that stop along with this service when it is stopped are the impact scope.
What this calls
Find the downstream in the configuration, and write down what changes put the downstream at risk.
Look at downstream in the configuration. And also write down what changes put the downstream at risk — if you lengthen the timeout, we think we gave some slack, but on the downstream the connections stay held longer.
Shared state that is not in the system diagram
Compare the two configuration files and find the shared state.
There are two configuration files. Look for lines that point to the same thing. Shared state is rarely drawn in the system diagram, yet accidents come from here.
When is it not safe
Find the batch window in the crontab and decide a safe time.
Look at crontab. When the daily close, the nightly report, and the monthly settlement each run. Even if it is technically safe, the timing being wrong makes it an accident.
Write the rollback first
plan.md and rollback.sh — the first line of a change plan is not the change
but how to roll it back. Also the time the rollback takes and the observation window.
The first line of the change plan is the way to roll back. Also the time the rollback takes and the observation window (what to watch and for how long). The grader deliberately changes the configuration and then runs your rollback.sh to check that it returns to the original.
Leave it before the restart
snapshot.sh — leave the state before a restart.
A restart erases the evidence along with the symptom. Leave the four things — processes, memory, open files, recent logs — together with the time. This one file will later become half of the report.