In Front of an Unfamiliar System
What to Ask For the Moment You Arrive at the Customer
In one line
What blocks you on the first day is not skill but permissions. If you go knowing what to request, you gain a day, and if you do not know, you wait three. And once you have received permissions, you check on the spot whether what you received is really what you requested.
Why this was needed
A common first-visit day at a customer site goes like this. Introductions in the morning, approval to bring in the laptop in the afternoon, a VPN account the next day, server access the day after. You actually get your hands on things on the fourth day. But when you open the logs on the fourth day, the retention period is two weeks, so the logs of the outage the customer mentioned from three weeks ago have already been deleted, and you found the problem but your permission to fix it is read only, so you have to file a new change request. Most of this delay arises because you did not say in advance what you needed.
Approval systems are mostly serial. The account request goes up only after the VPN approval is finished, and the permission request can be made only after the account comes out. Each step has a different approver, and each approver takes a day. So you have to throw everything at once on the first day to have it cooked in parallel. Even for things that have to follow an order, if you tell them in advance "what will be requested next", the next approval goes up the moment the previous step finishes.
What to request
| Request | Why it is needed | What is often missed |
|---|---|---|
| Network access (VPN/leased line) | You can do nothing | Registering the 2FA device is a separate procedure |
| Account + permissions | You cannot even query | Read permission and execute permission are separate |
| Server list and system diagram | You do not know where to look | The latest copy is not on the wiki but on someone's PC |
| Log location and retention period | It decides the investigation scope | You find out later that nothing is left after 30 days |
| People in charge and the contact chain | Where to ask when stuck | Night and weekend contact rules |
| Change procedure | Absolutely necessary to fix anything | Even an emergency change may need approval |
One more on top of this, the one most often missing — whether a test environment exists, and if so, what differs from production. The answer "it is the same" is usually not true. The data volume, the external integrations, and the certificates differ. Before spending a day trying to reproduce in the test environment a problem that occurs only in production, get the list of differences first.
How to check the permissions you received
When you get word that the account has come out, look at a few things as soon as you connect. All of them are read-only commands, so they are safe even on an unfamiliar server.
id; groups # which groups this account belongs to
sudo -l # what may be run as root, without running it
chage -l "$USER" # account and password expiry dates
timedatectl # is the clock synced, which time zone
sudo -l shows only the list without actually using the execute permission. If a command you need is missing here,
file the additional request right on the first day. The expiry date in chage -l gets hit surprisingly often. Partner accounts are often
issued for a short time, and the account gets locked when the investigation is in full swing. The clock and the time zone are
needed later when you line up the logs of several machines.
Do not take the log retention period from what you are told; check it in the configuration. File logs are usually managed by
logrotate. rotate is the number of files to keep before deleting, and you have to multiply it by the rotation
interval (daily, weekly) to get a period. With weekly and rotate 4, the file in use
and about four weeks of past files remain. If there is a maxage, files older than that
number of days are deleted regardless of the count.
grep -nE 'daily|weekly|monthly|rotate|maxage' /etc/logrotate.conf /etc/logrotate.d/*
journalctl --disk-usage
ls -d /var/log/journal # absent: the journal may live in memory only
journalctl --list-boots | head -3
The systemd journal is by default deleted by size, not by period. The default
cap is 10% of the file system size and at most 4G, and the period limit (MaxRetentionSec) is
off by default. So the more logs a server produces, the shorter the period the journal covers. There is a more important
trap. With the default setting (Storage=auto), it stores to disk only when the
/var/log/journal directory exists, and otherwise keeps it in memory only. On that server, the moment it reboots,
the entire previous journal disappears. If journalctl --list-boots shows only a single boot, suspect this case.
What you do not touch
The default posture of the first week is read only. This is not timidity but arithmetic. On an unfamiliar system, you cannot predict the impact scope of a single change.
Before you touch anything, check three things.
- Can it be reverted? — You must be able to explain in words how to revert it.
- Who is affected? — If you do not know who uses this server, it is still too early.
- Should it be done now? — If you start with fixing during the investigation stage, you lose the cause.
Especially the third. A restart erases the evidence along with the symptom. Even if a restart is necessary, leave the state before it — the process list, memory, open files, recent logs. It takes only a few seconds.
date -u +%FT%TZ # when this snapshot was taken
ps auxf # process tree
free -m # memory
ss -tanp # sockets and their owners
lsof -p "$PID" # files the process holds open
journalctl -u "$UNIT" --since "30 min ago"
The reason to stamp the time first is that when several snapshots pile up, you can no longer tell which one is from before the restart. This one snapshot will later become half of the report.
Trust is decided in the first week
Even if you say something technically right, if there is no relationship in which people will listen, you can change nothing. The way to build trust in the first week is simple.
- Return small things quickly — A one-hour check result must go out before a three-day analysis.
- Say you do not know what you do not know — If you state a guess as a fact, you lose everything at once.
- Report at the promised time — Even with no progress, give "nothing yet" on time.
What it looks like in the field
- While investigating, you learn that the logs cover only 20 days → had you asked about the retention period on the first day, you would have set the investigation scope differently. Had you told the customer on the first day that "the outage three weeks ago cannot be confirmed from the logs", you would have looked for other evidence (monitoring metrics, the customer's own records) in advance.
- The journal is empty after the server rebooted → the storage method was memory. If the cause of the outage was right before the reboot, that record could never have been kept in the first place.
- You found the problem but have no permission to fix it → the result of not checking the change procedure in advance.
Had you seen on the first day with
sudo -lthat read permission and execute permission are separate, you would not have been blocked. - The account got locked in the middle of the investigation → you did not check the expiry date of the partner account.
- The symptom disappeared with a restart and it was closed as cause unknown → the price of restarting without a snapshot.
What to check in the learning that follows
The reading right after this covers how to draw what touches what before making a change. Then in the lab, you dig out for yourself the first-day request list, the retention period, upstream and downstream, shared state, and the batch window, and build a rollback script and a pre-restart snapshot script. The retention-period calculation and the snapshot commands you saw in this section are used as they are.