Investigating a Window With No Logs
Goal
You received a report that "a few failed yesterday afternoon" but there are no logs for that interval. You narrow the time with traces other than logs, and go as far as making sure logs are left next time.
Environment
You work under /root/nolog. You build the site yourself — it is preparation, not
the assignment.
mkdir -p /root/nolog/logs /root/nolog/etc && cd /root/nolog
python3 - <<'PY'
import sqlite3, random, datetime
random.seed(11)
base = datetime.datetime(2026, 9, 7, 12, 0, 0)
con = sqlite3.connect('app.db'); cur = con.cursor()
cur.execute("create table orders(id integer primary key, status text, created_at text)")
rows, oid = [], 1
for m in range(240):
t = base + datetime.timedelta(minutes=m)
fail = random.randint(18, 26) if 123 <= m <= 126 else (1 if random.random() < 0.15 else 0)
for _ in range(random.randint(8, 14)):
rows.append((oid, 'PAID', (t + datetime.timedelta(seconds=random.randint(0, 59))).isoformat())); oid += 1
for _ in range(fail):
rows.append((oid, 'FAILED', (t + datetime.timedelta(seconds=random.randint(0, 59))).isoformat())); oid += 1
cur.executemany("insert into orders values (?,?,?)", rows); con.commit(); con.close()
with open('logs/access.log', 'w', encoding='utf-8') as f:
for m in range(240):
if 123 <= m <= 127: continue
t = base + datetime.timedelta(minutes=m)
for i in range(random.randint(5, 9)):
f.write('%s GET /api/pay 200\n' % (t + datetime.timedelta(seconds=i * 6)).isoformat())
with open('logs/error.log', 'w', encoding='utf-8') as f:
for m in range(130, 240):
if random.random() < 0.2:
f.write('%s WARN slow query 1200ms\n' % (base + datetime.timedelta(minutes=m)).isoformat())
PY
printf 'pool_size=2\ntimeout=1\n' > etc/app.conf
printf 'log_level=WARN\n' > etc/other.conf
touch -t 202609071402 etc/app.conf
touch -t 202608311000 etc/other.conf
ls -l --time-style=+%Y-%m-%dT%H:%M etc/
What gets created is this.
app.db 주문 2,000여 건 (status, created_at)
logs/access.log 접근 로그 — 사고 구간이 비어 있다
logs/error.log 오류 로그 — 보존 기간 때문에 뒷부분만 남았다
etc/app.conf 설정 파일
etc/other.conf 설정 파일
What to create
minute.txt 사고가 시작된 분과 그 근거
gap.txt 비어 있는 구간과, 로그가 사라진 범위
traces.txt 로그가 아닌 흔적 세 가지 이상
cause.txt 의심 대상과 배제한 후보
watch.sh 지금 진행 중인 문제에 붙일 사후 계측기
next.md 다음 조사를 줄이는 네 가지
report.md 정리
Steps
- Build the site.
- The data is the log. Group
created_atby minute and find the minute where failures are concentrated. You must also write the usual value for 'concentrated' to be proven. - Find the empty interval in the access log, and also write from where the error log remains. You must know the range that disappeared to set the scope of the investigation.
- Gather at least three traces that are not logs. Not only file modification times but also at least one among boot, process, package, and certificate.
- Cross-check the times and narrow the candidates. Also write the candidates you ruled out and the grounds.
watch.sh— leaves times and observed values periodically. The grader runs it directly.next.md— the four things, log statements · retention period · correlation ID · metric, made concrete.- Wrap up.
Notes
The minute to point to in step 2 is not just one. The interval where failures are concentrated spans several minutes, so you may point to any of them — but you must write that minute's count and the usual value together.
The times matching in step 5 is correlation, not causation. Writing that distinction in the document prevents you from later reverting the wrong thing.
Build the site where the logs have disappeared
Build the site.
Run the preparation block in the instructions as it is. That the error log is missing from the incident interval onward is the premise of this lab.
The data is the log
The data is the log. Group created_at by minute and
find the minute where failures are concentrated. You must also write the usual value for 'concentrated' to be proven.
Group created_at by minute (substr(created_at,1,16)) and count FAILED. The minute where the value spikes is the time of the incident. You must also write the usual value for the comparison to work.
Absence is evidence too
Find the empty interval in the access log, and also write from where the error log remains. You must know the range that disappeared to set the scope of the investigation.
If you count the access log by minute, you see an interval with 0 lines. And also check what time the first line of the error log is — you must know how far the loss extends to set the scope of the investigation.
Gather traces that are not logs
Gather at least three traces that are not logs. Not only file modification times but also at least one among boot, process, package, and certificate.
At least three among file modification time (ls -l --time-style), boot time (uptime -s), process start (ps -eo lstart), and package install (/var/log/dpkg.log). A trace can be cross-checked only if it comes with a time.
Cross-check the times to narrow down
Cross-check the times and narrow the candidates. Also write the candidates you ruled out and the grounds.
Look for what changed right before the incident start time. Also write the candidates you ruled out and the grounds. And also that times matching is correlation, not causation.
If it is in progress now, post-hoc instrumentation
watch.sh — leaves times and observed values periodically. The grader runs it directly.
While it reproduces, leave times and observed values periodically. Even if rough, it is better than nothing, and it becomes the only basis at the next meeting. The grader runs this script directly.
Faster next time than this time
next.md — the four things, log statements · retention period · correlation ID · metric, made concrete.
They are the four: log statements · retention period · correlation ID · metric. You must write down 'what, and at what value' for it to be carried out — including the fact that retention is currently 24 hours.
Wrap up
Wrap up.
Time of the incident · the fact that there were no logs · the suspect · preventing recurrence. And do not leave out the heart of this investigation, 'absence was evidence too.'