TT Lab
Get started
Learn Learning paths Courses

FDE Capstone: The Warehouse Got the Same Order Three Times

We rebooted and nothing came back up

Continue in TT Lab

This lab runs on a real VM

It is one Ubuntu 24.04 VM. systemd is PID 1 and there is no Kubernetes. The preparation step installs a small order lookup app (/opt/orders-svc/app.py) and reproduces the state in which the previous owner started it from a shell with nohup. The grading is done by an agent inside the VM, so do not reboot — the grading will be cut off. Whether it comes up at boot is judged by the enable state and the target dependency.

Goal

You install the app on the customer VM as a systemd service with a dedicated user, an environment file, a restart policy, and a boot target, check for yourself what a missing enable, kill -9, a missing daemon-reload, and an environment file replacement each change, and then finish with a handoff check script.

Why it matters

"I installed it and it works well" only means it is up now. A process started from a shell is a process the service manager does not know about, so if it dies nobody revives it, and if you reboot nobody starts it. It is the same if you only run systemctl start and forget enable. Startup at boot hangs on the single symlink that the [Install] section creates at enable time, and if you forget daemon-reload after fixing a unit file, systemd keeps running on the old configuration. Only if you make each of these differences by hand once can you answer "why did nothing come up" at a customer site within 5 minutes.

The expected time is 60 minutes. The preparation takes a few minutes. When the session ends, the VM disappears, so keep any files you want to keep separately before it ends.

Steps

  1. Find the previous owner's process that is holding port 8181 and write it into /root/svc/legacy.txt as five lines: pid=, user=, cgroup= (the path in /proc/<pid>/cgroup), unit= (the last piece of that path), and survives_reboot= (yes or no).
  2. Create a system account orders that cannot log in, and create /var/lib/orders (owner orders, 750) and /etc/orders/orders.env (root:orders, 640). Put ORDERS_PORT=8181, a new ORDERS_TOKEN of 16 characters or more, and ORDERS_DATA_DIR=/var/lib/orders in the environment file.
  3. Write /etc/systemd/system/orders.service (User=orders, EnvironmentFile, ExecStart=/usr/bin/python3 /opt/orders-svc/app.py, Restart=on-failure, Wants and After=network-online.target, WantedBy=multi-user.target). If you start it after daemon-reload, it fails. Find the cause line the app left in the journal (bind failed pid=…) and copy it as it is into /root/svc/why-failed.txt.
  4. Take down the previous owner's process and start orders.service. curl http://127.0.0.1:8181/healthz must respond with the orders user and the PID of the service's main process.
  5. Enable the service and save the output that appears then (including the path of the symlink that was created) into /root/svc/enable.txt.
  6. Kill the main process with kill -9 and see whether systemd revives it. Write old_pid= and new_pid= into /root/svc/kill.txt.
  7. Add sandboxing (NoNewPrivileges=yes, ProtectSystem=strict, ProtectHome=yes, PrivateTmp=yes, ReadWritePaths=/var/lib/orders) to the unit. Before changing it, measure the systemd-analyze security score, save the daemon-reload warning that systemctl gives right after you fix the file into /root/svc/reload.txt, and then daemon-reload, restart, and measure again. Write before= and after= into /root/svc/security.txt.
  8. Change the ORDERS_TOKEN of the environment file to a new value and apply it to the service. Write old_token_sha= and new_token_sha= (the sha256 hex of the token) and daemon_reload_needed= (yes or no) into /root/svc/rotate.txt.
  9. Write the handoff check script /root/svc/verify.sh. bash verify.sh <유닛> (the placeholder is the unit) must end with 0 if that unit meets all the conditions to come up even after a reboot (loaded, enabled, multi-user.target in WantedBy, Restart being on-failure or always, network-online.target in both Wants and After, User not empty and not root, NeedDaemonReload=no), and if it violates even one, print the violated items and end with a nonzero value. The grader also runs it on the other prepared units (orders-audit, orders-report, orders-worker, orders-sync, and orders-rootjob).

Notes

The stray process holding 8181

Write the pid, user, cgroup, and unit of the process listening on 8181, and whether it survives a reboot, into /root/svc/legacy.txt.

The -p of ss shows the process that owns the socket. The path in /proc//cgroup tells you which unit this process was born in. Think about whether that unit is the one responsible for starting this app.

A dedicated account and an environment file

Create the system account orders, /var/lib/orders (orders, 750), and /etc/orders/orders.env (root:orders, 640).

The --system of useradd uses the system UID range, and --shell gives a shell that cannot log in. The environment file holds a token, so set the group and permissions so that only the service account can read it. The old token on the wiki is already a leaked value.

Find the reason the first start failed in the journal

Write /etc/systemd/system/orders.service and try to start it, then copy the failure cause line from the journal into /root/svc/why-failed.txt.

The app's standard output goes to the journal. journalctl -u 유닛 -o cat (the placeholder is the unit) shows only the message without the prefix. Even if systemctl start ends in success, Type=simple is a success the moment it starts the process, so you have to look at the actual result in status and the journal.

Take down the stray process and stand up the service

Take down the previous owner's process and make orders.service active. healthz must respond with the orders user and the main process PID.

Take down only the app that is up as root (you must not kill the orders account's process too). If the failure from the earlier step repeated and you hit the start count limit, you need reset-failed.

The single symlink that enable makes

Enable orders.service and save the output (including the symlink path) into /root/svc/enable.txt.

enable does not start the service. It only reads the WantedBy of [Install] and creates a symlink in the .wants directory of that target, and at boot that target pulls in this service. The message comes out on standard error.

Does it come back even after kill -9

Kill the main process with kill -9 and write old_pid= and new_pid=, together with the revived PID, into /root/svc/kill.txt.

The main process PID is the MainPID of systemctl show. Look at the manual's table for which terminations Restart=on-failure counts as failures. After waiting as long as RestartSec, a new PID appears.

Sandboxing and a missing daemon-reload

Measure the score, add sandboxing to the unit, leave the daemon-reload warning in /root/svc/reload.txt, apply it, and write before= and after= into /root/svc/security.txt.

The last line of systemd-analyze security is the overall exposure score (the lower, the more narrowly confined). ProtectSystem=strict makes the whole filesystem read-only, so you have to open the paths the app writes to separately. Read the systemctl status right after you fix the file.

How is a token replacement applied

Change ORDERS_TOKEN to a new value, apply it to the service, and write old_token_sha=, new_token_sha=, and daemon_reload_needed= into /root/svc/rotate.txt.

EnvironmentFile is not a unit setting but a file read right before the process is started. Think about how that differs from fixing a unit file, and when the environment of an already running process changes. Compute the hash with only the token, without a newline.

A check script that proves booting without a reboot

Write /root/svc/verify.sh, which judges whether a unit meets all the conditions to come up even after a reboot.

Do not grep the file; look at the values systemd loaded with systemctl show and is-enabled. For a nonexistent unit too, show succeeds but the LoadState is not-found. If you print all the violated items instead of ending at the first one, the person taking over can fix them in one go.