FDE Capstone: The Warehouse Got the Same Order Three Times
The install succeeded, but the port already belonged to someone else
In one line
A preflight check judges, by reading only, whether the customer's host meets the requirements before installation, and leaves the result as a report and an exit code that the next stage's automation, not a person, can read.
Why this was needed
The installation window at a customer site is usually a one-shot. The infrastructure team opens the door at 22:00 on Thursday night and closes it at midnight. In a setting like this, the fact that an install script ended with "done" guarantees almost nothing. The files may have been copied into place, but the port the agent wanted to use may already be held by an internal proxy, so it dies as soon as it starts, and the certificate the customer handed over may have expired three weeks before the installation. Both can only be found by digging through logs after the installation, but they are facts you could learn in five minutes before it.
What an FDE hands over to a customer is not code but a working result. That is why you put a separate step before the installation procedure that asks "are the conditions for the installation to succeed in place on this host?" If you leave this step as a person's checklist, someone skips it every time. If you make it a script and leave the result in a file, you can tell the customer with evidence that "these two things block the installation and the rest are warnings".
How it works
The checker takes a requirements specification as input and, for each check, produces one of pass, warn, or fail and an observed value. The observed value matters. A report that only says "port failed" cannot be refuted but cannot be verified either. Only when observed: in_use, the measured path, the free MiB, and the notAfter time are there together can the customer's owner recheck with their own eyes.
The exit codes are easy to explain if you borrow the monitoring plugin convention. The Monitoring Plugins development guidelines use 0 OK, 1 Warning, 2 Critical, and 3 Unknown, and limit 3 to cases where the arguments are wrong or the checker itself could not run. The point worth noting is that they write not to escalate higher-level errors such as a name resolution failure or a socket timeout to Unknown. A hostname that does not resolve is not "we don't know" but a fact that blocks the installation.
Each individual check has its common traps.
- Ports. To see whether it is free, you actually bind. But if you bind with no options, even a port where a connection that just ended remains in TIME_WAIT shows up as in use. In its explanation of
create_server(), the Python socket documentation says that on POSIX it turns onSO_REUSEADDRto reuse an address left in TIME_WAIT right away. When we measured on the lab image, a port with a listening socket was EADDRINUSE even with this option on, and a port with only TIME_WAIT left could be bound only when the option was on. A checker that does not turn it on blocks a perfectly good installation. - Disk. shutil.disk_usage gives total, used, and free in bytes. In our measurement, free was equal to the statvfs
f_bavail(the blocks available to an unprivileged user) multiplied by the fragment size. Do not create an installation path that does not exist yet; measure the nearest existing parent, and write the measured path in the report. - Certificate. The Python ssl module has no public function that reads a PEM file and extracts notAfter. Instead,
ssl.cert_time_to_seconds()converts a notAfter string in the form"%b %d %H:%M:%S %Y %Z"into epoch seconds. You get the string with-enddateof openssl x509. The-checkendof the same documentation ends with a nonzero value if it expires within the given seconds, but to write the number of remaining days in the report, it is better to calculate the dates yourself. - Python version. If you compare "3.12.3" and "3.9" as strings, 3.12 is smaller. Compare them as integer tuples.
- Write permission. The os documentation says that checking with
os.access()and then opening creates a gap between the check and the use, and recommends EAFP, that is, actually trying and catching the exception. There are also reasons that the permission bits do not tell you, such as a read-only mount, so you actually create a temporary file and delete it right away.
명세(spec.json) ──▶ preflight.py ──▶ report.json (점검마다 status + observed)
│
└──▶ 종료 코드 0 · 1 · 2 (명세를 못 읽으면 3, 보고서 없음)
What it looks like in the field
The most common mistake is the checker fixing things. If the data directory does not exist, it creates it, it kills the process holding the port, and it swaps an expired certificate for a self-signed one. At that moment the check result loses "what this host originally looked like". Besides, directory ownership, the internal proxy that was holding the port, and the certificate issuance procedure are all the customer organization's decisions. A preflight check reveals the facts, and fixing is done by the owner through an agreed procedure. The customer memo this time also says that the infrastructure team will create the working directory.
The second is mixing warnings and failures. If the free space exceeds the vendor's minimum but is less than the operations team's standard, the installation is possible. If you escalate even this to a failure, you waste the installation window, and if you do not report it at all, the disk fills up two months later. That is why there is a separate warning grade.
What really matters in practice
- A report has to be re-measurable. Do not write only the verdict; leave the observed value and the measured target together.
- The checker does not change the host. It even deletes the temporary file of the write test.
- An exit code is a promise. You make the install automation read 1 as proceed and 2 as stop.
- A stale report is not evidence. Run it again right before the installation, and leave in the decision record a hash that says which report it was.
What you will do in the next lab
You turn the customer memo into a specification and implement, in turn, the port, TIME_WAIT, disk, certificate, file, host, and write checks. The grader runs your checker every time with a different port and threshold values and with expiring and expired certificates it makes itself, and compares the verdicts and observed values. At the end, you make a real verdict with the customer specification and leave a go / no-go record.