FDE Capstone: The Warehouse Got the Same Order Three Times
We rebooted and nothing came back up
This lab runs on a real VM
It is one Ubuntu 24.04 VM. systemd is PID 1 and there is no Kubernetes. The preparation step installs a small order lookup app (/opt/orders-svc/app.py) and reproduces the state in which the previous owner started it from a shell with nohup. The grading is done by an agent inside the VM, so do not reboot — the grading will be cut off. Whether it comes up at boot is judged by the enable state and the target dependency.
Goal
You install the app on the customer VM as a systemd service with a dedicated user, an environment file, a restart policy, and a boot target, check for yourself what a missing enable, kill -9, a missing daemon-reload, and an environment file replacement each change, and then finish with a handoff check script.
Why it matters
"I installed it and it works well" only means it is up now. A process started from a shell is a process the service manager does not know about, so if it dies nobody revives it, and if you reboot nobody starts it. It is the same if you only run systemctl start and forget enable. Startup at boot hangs on the single symlink that the [Install] section creates at enable time, and if you forget daemon-reload after fixing a unit file, systemd keeps running on the old configuration. Only if you make each of these differences by hand once can you answer "why did nothing come up" at a customer site within 5 minutes.
The expected time is 60 minutes. The preparation takes a few minutes. When the session ends, the VM disappears, so keep any files you want to keep separately before it ends.
Steps
- Find the previous owner's process that is holding port 8181 and write it into
/root/svc/legacy.txtas five lines:pid=,user=,cgroup=(the path in/proc/<pid>/cgroup),unit=(the last piece of that path), andsurvives_reboot=(yes or no). - Create a system account
ordersthat cannot log in, and create/var/lib/orders(owner orders, 750) and/etc/orders/orders.env(root:orders, 640). PutORDERS_PORT=8181, a newORDERS_TOKENof 16 characters or more, andORDERS_DATA_DIR=/var/lib/ordersin the environment file. - Write
/etc/systemd/system/orders.service(User=orders, EnvironmentFile, ExecStart=/usr/bin/python3 /opt/orders-svc/app.py, Restart=on-failure, Wants and After=network-online.target, WantedBy=multi-user.target). If you start it after daemon-reload, it fails. Find the cause line the app left in the journal (bind failed pid=…) and copy it as it is into/root/svc/why-failed.txt. - Take down the previous owner's process and start orders.service.
curl http://127.0.0.1:8181/healthzmust respond with the orders user and the PID of the service's main process. - Enable the service and save the output that appears then (including the path of the symlink that was created) into
/root/svc/enable.txt. - Kill the main process with
kill -9and see whether systemd revives it. Writeold_pid=andnew_pid=into/root/svc/kill.txt. - Add sandboxing (NoNewPrivileges=yes, ProtectSystem=strict, ProtectHome=yes, PrivateTmp=yes, ReadWritePaths=/var/lib/orders) to the unit. Before changing it, measure the
systemd-analyze securityscore, save the daemon-reload warning that systemctl gives right after you fix the file into/root/svc/reload.txt, and then daemon-reload, restart, and measure again. Writebefore=andafter=into/root/svc/security.txt. - Change the
ORDERS_TOKENof the environment file to a new value and apply it to the service. Writeold_token_sha=andnew_token_sha=(the sha256 hex of the token) anddaemon_reload_needed=(yes or no) into/root/svc/rotate.txt. - Write the handoff check script
/root/svc/verify.sh.bash verify.sh <유닛>(the placeholder is the unit) must end with 0 if that unit meets all the conditions to come up even after a reboot (loaded, enabled, multi-user.target in WantedBy, Restart being on-failure or always, network-online.target in both Wants and After, User not empty and not root, NeedDaemonReload=no), and if it violates even one, print the violated items and end with a nonzero value. The grader also runs it on the other prepared units (orders-audit,orders-report,orders-worker,orders-sync, andorders-rootjob).
Notes
- The port owner:
ss -ltnp 'sport = :8181'. The unit that holds a process:systemctl status <pid>or/proc/<pid>/cgroup. - You look at properties not in the file but in the values systemd has loaded:
systemctl show orders.service -p User,Restart,Wants,After,NeedDaemonReload. - If you repeat a failure many times, it no longer starts because of
start-limit-hit. Aftersystemctl reset-failed orders.service, start again. - Common mistake 1: writing only
After=network-online.target. It only sets the order and does not pull in that target. - Common mistake 2: ending after seeing it up with
systemctl start. Startup at boot is decided by the symlink that enable created. - Common mistake 3: turning on only ProtectSystem=strict. The data directory also becomes read-only and the app dies with
data dir error.
The stray process holding 8181
Write the pid, user, cgroup, and unit of the process listening on 8181, and whether it survives a reboot, into /root/svc/legacy.txt.
The -p of ss shows the process that owns the socket. The path in /proc//cgroup tells you which unit this process was born in. Think about whether that unit is the one responsible for starting this app.
A dedicated account and an environment file
Create the system account orders, /var/lib/orders (orders, 750), and /etc/orders/orders.env (root:orders, 640).
The --system of useradd uses the system UID range, and --shell gives a shell that cannot log in. The environment file holds a token, so set the group and permissions so that only the service account can read it. The old token on the wiki is already a leaked value.
Find the reason the first start failed in the journal
Write /etc/systemd/system/orders.service and try to start it, then copy the failure cause line from the journal into /root/svc/why-failed.txt.
The app's standard output goes to the journal. journalctl -u 유닛 -o cat (the placeholder is the unit) shows only the message without the prefix. Even if systemctl start ends in success, Type=simple is a success the moment it starts the process, so you have to look at the actual result in status and the journal.
Take down the stray process and stand up the service
Take down the previous owner's process and make orders.service active. healthz must respond with the orders user and the main process PID.
Take down only the app that is up as root (you must not kill the orders account's process too). If the failure from the earlier step repeated and you hit the start count limit, you need reset-failed.
The single symlink that enable makes
Enable orders.service and save the output (including the symlink path) into /root/svc/enable.txt.
enable does not start the service. It only reads the WantedBy of [Install] and creates a symlink in the .wants directory of that target, and at boot that target pulls in this service. The message comes out on standard error.
Does it come back even after kill -9
Kill the main process with kill -9 and write old_pid= and new_pid=, together with the revived PID, into /root/svc/kill.txt.
The main process PID is the MainPID of systemctl show. Look at the manual's table for which terminations Restart=on-failure counts as failures. After waiting as long as RestartSec, a new PID appears.
Sandboxing and a missing daemon-reload
Measure the score, add sandboxing to the unit, leave the daemon-reload warning in /root/svc/reload.txt, apply it, and write before= and after= into /root/svc/security.txt.
The last line of systemd-analyze security is the overall exposure score (the lower, the more narrowly confined). ProtectSystem=strict makes the whole filesystem read-only, so you have to open the paths the app writes to separately. Read the systemctl status right after you fix the file.
How is a token replacement applied
Change ORDERS_TOKEN to a new value, apply it to the service, and write old_token_sha=, new_token_sha=, and daemon_reload_needed= into /root/svc/rotate.txt.
EnvironmentFile is not a unit setting but a file read right before the process is started. Think about how that differs from fixing a unit file, and when the environment of an already running process changes. Compute the hash with only the token, without a newline.
A check script that proves booting without a reboot
Write /root/svc/verify.sh, which judges whether a unit meets all the conditions to come up even after a reboot.
Do not grep the file; look at the values systemd loaded with systemctl show and is-enabled. For a nonexistent unit too, show succeeds but the LoadState is not-found. If you print all the violated items instead of ending at the first one, the person taking over can fix them in one go.