In Front of an Unfamiliar System
Scouting a Customer Environment
Goal
You will be able to go into a customer server for the first time, with no documentation and no dashboard, and produce one environment map within 30 minutes.
Why it matters
The most expensive mistake in an unfamiliar environment is opening an arbitrary log. Logs are an ocean, and if you go in without knowing the size or the format, the time just vanishes. The purpose of recon is not to find the answer but to establish the search scope. If you know how many lines a file has, you know how many seconds a single grep takes, and if you know the collection interval, you can tell whether the log for the reported time exists in the first place. And in the capacity check you always look at two things together — blocks and inodes are different resources, so even when capacity remains, if inodes run out, files cannot be created. It is actually common to answer the customer "the disk has room" after looking only at df, and then have it come back to you.
Steps
- Create the
/root/recondirectory. - Create
/root/recon/os.txtso that it contains the distribution name of this system. - Save the list of files under
/opt/data, with their sizes, to/root/recon/files.txt. - Write only the file name of the largest file at the top level of
/opt/datato/root/recon/biggest.txt. - Write the total line count of
/opt/data/web.log, as a number, to/root/recon/weblog_lines.txt. - Save both the filesystem capacity and the inode usage to
/root/recon/disk.txt. - Write the earliest time and the latest time in
/opt/data/web.logto/root/recon/timerange.txt. - Summarize what you found above on one page,
/root/recon/map.md. It must contain the distribution, the largest file name, the line count of web.log, and the log collection interval.
Notes
cat /etc/os-release,ls -lS,du -h,wc -l,df -h,df -i- To gather the output of two commands into one file, append with
>>. - Common mistake 1: checking only
dfand leaving out the inodes. Capacity and inodes are separate resources. - Common mistake 2: writing the path too in step 4. Only the file name is needed.
Create the recon working directory
Create the /root/recon directory.
Gather the outputs in one place so that later they get bundled into a report. If you add -p to mkdir, it does not raise an error even if the directory already exists.
Check the distribution
Create /root/recon/os.txt so that it contains the distribution name of this system.
The distribution information is in /etc/os-release as key=value. This value decides the log paths and the package commands.
Build the data inventory
Save the list of files under /opt/data, with their sizes, to /root/recon/files.txt.
List /opt/data together with the sizes. You can send the output of ls -l or du straight into the file.
Find the largest file
Write only the file name of the largest file at the top level of /opt/data to /root/recon/biggest.txt.
Do not count by eye; sort. ls has an option to sort by size, and du is combined with sort. Write only the file name.
Count the lines of the web log
Write the total line count of /opt/data/web.log to /root/recon/weblog_lines.txt as a number.
You need to know the line count to predict how many seconds a single grep takes. Use the line-counting option of wc.
Record capacity and inodes together
Save both the filesystem capacity and the inode usage to /root/recon/disk.txt.
By default df shows only blocks. There is a separate option for viewing inodes, and both outputs must be in the one file.
Find out the log collection interval
Write the earliest time and the latest time in /opt/data/web.log to /root/recon/timerange.txt.
The second field of web.log is the time. If you extract only that field and sort it, the earliest and the latest values come at the two ends.
Summarize on one environment map
Summarize what you found above on one page, /root/recon/map.md. It must contain the distribution, the largest file name, the line count of web.log, and the log collection interval.
Gather the values you found in the earlier steps on one Markdown page. The distribution, the largest file, the log line count, and the collection interval must all be in it.