TT Lab
Get started
Learn Learning paths Courses

In Front of an Unfamiliar System

What to Ask For the Moment You Arrive at the Customer

Continue in TT Lab

In one line

What blocks you on the first day is not skill but permissions. If you go knowing what to request, you gain a day, and if you do not know, you wait three. And once you have received permissions, you check on the spot whether what you received is really what you requested.

Why this was needed

A common first-visit day at a customer site goes like this. Introductions in the morning, approval to bring in the laptop in the afternoon, a VPN account the next day, server access the day after. You actually get your hands on things on the fourth day. But when you open the logs on the fourth day, the retention period is two weeks, so the logs of the outage the customer mentioned from three weeks ago have already been deleted, and you found the problem but your permission to fix it is read only, so you have to file a new change request. Most of this delay arises because you did not say in advance what you needed.

Approval systems are mostly serial. The account request goes up only after the VPN approval is finished, and the permission request can be made only after the account comes out. Each step has a different approver, and each approver takes a day. So you have to throw everything at once on the first day to have it cooked in parallel. Even for things that have to follow an order, if you tell them in advance "what will be requested next", the next approval goes up the moment the previous step finishes.

What to request

Request Why it is needed What is often missed
Network access (VPN/leased line) You can do nothing Registering the 2FA device is a separate procedure
Account + permissions You cannot even query Read permission and execute permission are separate
Server list and system diagram You do not know where to look The latest copy is not on the wiki but on someone's PC
Log location and retention period It decides the investigation scope You find out later that nothing is left after 30 days
People in charge and the contact chain Where to ask when stuck Night and weekend contact rules
Change procedure Absolutely necessary to fix anything Even an emergency change may need approval

One more on top of this, the one most often missing — whether a test environment exists, and if so, what differs from production. The answer "it is the same" is usually not true. The data volume, the external integrations, and the certificates differ. Before spending a day trying to reproduce in the test environment a problem that occurs only in production, get the list of differences first.

How to check the permissions you received

When you get word that the account has come out, look at a few things as soon as you connect. All of them are read-only commands, so they are safe even on an unfamiliar server.

id; groups                 # which groups this account belongs to
sudo -l                    # what may be run as root, without running it
chage -l "$USER"           # account and password expiry dates
timedatectl                # is the clock synced, which time zone

sudo -l shows only the list without actually using the execute permission. If a command you need is missing here, file the additional request right on the first day. The expiry date in chage -l gets hit surprisingly often. Partner accounts are often issued for a short time, and the account gets locked when the investigation is in full swing. The clock and the time zone are needed later when you line up the logs of several machines.

Do not take the log retention period from what you are told; check it in the configuration. File logs are usually managed by logrotate. rotate is the number of files to keep before deleting, and you have to multiply it by the rotation interval (daily, weekly) to get a period. With weekly and rotate 4, the file in use and about four weeks of past files remain. If there is a maxage, files older than that number of days are deleted regardless of the count.

grep -nE 'daily|weekly|monthly|rotate|maxage' /etc/logrotate.conf /etc/logrotate.d/*
journalctl --disk-usage
ls -d /var/log/journal     # absent: the journal may live in memory only
journalctl --list-boots | head -3

The systemd journal is by default deleted by size, not by period. The default cap is 10% of the file system size and at most 4G, and the period limit (MaxRetentionSec) is off by default. So the more logs a server produces, the shorter the period the journal covers. There is a more important trap. With the default setting (Storage=auto), it stores to disk only when the /var/log/journal directory exists, and otherwise keeps it in memory only. On that server, the moment it reboots, the entire previous journal disappears. If journalctl --list-boots shows only a single boot, suspect this case.

What you do not touch

The default posture of the first week is read only. This is not timidity but arithmetic. On an unfamiliar system, you cannot predict the impact scope of a single change.

Before you touch anything, check three things.

  1. Can it be reverted? — You must be able to explain in words how to revert it.
  2. Who is affected? — If you do not know who uses this server, it is still too early.
  3. Should it be done now? — If you start with fixing during the investigation stage, you lose the cause.

Especially the third. A restart erases the evidence along with the symptom. Even if a restart is necessary, leave the state before it — the process list, memory, open files, recent logs. It takes only a few seconds.

date -u +%FT%TZ                          # when this snapshot was taken
ps auxf                                  # process tree
free -m                                  # memory
ss -tanp                                 # sockets and their owners
lsof -p "$PID"                           # files the process holds open
journalctl -u "$UNIT" --since "30 min ago"

The reason to stamp the time first is that when several snapshots pile up, you can no longer tell which one is from before the restart. This one snapshot will later become half of the report.

Trust is decided in the first week

Even if you say something technically right, if there is no relationship in which people will listen, you can change nothing. The way to build trust in the first week is simple.

What it looks like in the field

What to check in the learning that follows

The reading right after this covers how to draw what touches what before making a change. Then in the lab, you dig out for yourself the first-day request list, the retention period, upstream and downstream, shared state, and the batch window, and build a rollback script and a pre-restart snapshot script. The retention-period calculation and the snapshot commands you saw in this section are used as they are.