In Front of an Unfamiliar System
Starting When You Know Nothing
In one line
In front of an unfamiliar system, what you do in the first 30 minutes is not fix the problem but establish what you do not need to look at.
Why this was needed
If it is your own team's service, your hands move first the moment you get an outage report. The dashboard's location, the directory where the logs pile up, and who deployed what yesterday are already in your head. At a customer site you have none of those three. You do not know what the monitoring is, you do not know where the logs are, you do not know what changed yesterday, and meanwhile the person next to you asks when it will be fixed.
Under these conditions, what separates experienced people from the rest is not the amount of knowledge but the order. Without an order, you end up opening whichever log catches your eye, and the logs of an unfamiliar environment are an ocean, so two hours later you are still in the same place.
So the first step is always reconnaissance. You are not looking for something to fix; you are drawing the terrain.
How it works
There are four things to establish in reconnaissance.
First, access and permissions. Instinct says to open the logs first, but if you discover mid-diagnosis that you lack permission, the time spent on that request and its approval is entirely wasted. Worse is acting beyond your permissions. On a customer's production, a single command that exceeds your permissions causes a greater loss of trust than the outage. Start by writing at the very top of the ticket whether what you hold now is read or write, and whether it is production or staging.
Second, what the system stands on. The distribution, the kernel, the running processes, the open ports. A single line of /etc/os-release decides the syntax of every later command. Even the default log path differs: the RHEL family uses /var/log/messages and the Debian family uses /var/log/syslog. When you carry over someone else's documentation as is, this is where it breaks most often.
Third, where the data is and how much of it there is. Which file is the largest, how many lines the log has, what format it is in. If you do not know the size, you cannot predict whether a single grep will finish in seconds or take minutes, and on a system that is already under load, the diagnostic command itself makes the outage worse.
Fourth, how much headroom the resources have. Here you must look twice. Check capacity with df -h and inodes separately with df -i. This is because capacity and inodes are completely different resources.
What it looks like in the field
Let me tell you about the most common trap in advance. The application died with "No space left on device", yet df -h says there is 30% free. The culprit is then one of three things. The inodes are exhausted, a file that was deleted but is still open is holding on to blocks, or that path is actually a different filesystem.
On Linux, a file name is not the thing itself but only a reference pointing to it, so even if you delete the name, the blocks are not released as long as some process has that file open. That is why, if a process does not reopen the file after log rotation, the deleted log keeps eating the disk. If the values of df and du differ greatly, it is almost always this story.
Knowing just this one thing saves you 20 minutes over other people on a "the disk is full" report. And those 20 minutes are what an FDE sells.
What recon leaves behind
The output of reconnaissance must be a document, not something in your head. There are two reasons. First, the next time you look at the same system (or when someone else does), you do not have to spend 30 minutes again. Second, showing the customer "what we have figured out" builds trust in itself. Even if you fixed nothing on the first day, if one map comes out, that time is not wasted.
What goes into the environment map is mostly fixed.
- Access path and permissions — which account you use to get in where, whether it is read or write, and whether that server is production or not.
- What is running — the processes, the open ports, and how they call each other. A port list alone reveals half of the dependencies.
- Where the data is — the logs, the data directories, the configuration files. The size and format of each.
- Resource headroom — capacity and inodes together.
- What you do not know — this item is the most important. Only if you write down what you could not confirm, exactly as it is, will you not forget later that you made an assumption at that spot.
Not mixing guesses with confirmed facts is the only rule of this document. "nginx is probably in front" and "nginx is listening on port 80" are different sentences, and a few days later you can no longer tell the two apart. If you write the command you used to confirm next to each confirmed fact, you can later check again with the same command whether the state has changed.
Finally, put a time limit on recon. Decide on 30 minutes or an hour, and if you have not finished drawing by then, move on to the next step with the map left unfinished. Failing to see the actual problem while trying to draw a perfect map is the second most common failure, after diving in without a map.
What you will do in the next lab
You go into a Pod that is assumed to be the customer's server, check the distribution, build a data inventory, record capacity and inodes together, and leave behind one environment map. You fix nothing. You only draw.