TT Lab
Get started
Learn Learning paths Courses

Shell Scripting

set -euo pipefail, and Why That Alone Is Not Enough

Continue in TT Lab

In a nutshell

set -euo pipefail is the minimum mechanism that prevents a script from silently producing wrong results. But on its own it neither cleans up temporary files nor prevents duplicate runs.

Why this was needed

A backup script ran every night before dawn and reported success every day. On the day recovery was needed, they discovered the backup file was 0 bytes.

tar czf "$DEST" "$SRC_DIR"

SRC_DIR was empty because of a typo, the shell simply passed the empty string along, tar made an empty archive, and the exit code was 0. It is the result of three places each passing "silently".

set -u alone would have stopped it at the very first step. The value of these options lies in stopping the problem at the point where it is discovered.

How it works

The standard four-line header looks like this.

#!/usr/bin/env bash
set -euo pipefail
IFS=$'\n\t'

But set -e does not solve everything. In fact, it has traps.

So a trap is needed as its partner.

cleanup() {
  local rc=$?
  rm -f "$TMPFILE"
  exit "$rc"
}
trap cleanup EXIT

Two things are key here. First, the EXIT trap runs on normal exit, error exit, and signal exit alike. Second, if you do not grab $? at the very start and restore it at the end, the trap swallows the exit code.

Always create temporary files with mktemp. A fixed name collides when runs overlap, and in /tmp it becomes a target for symbolic link attacks.

What it looks like in the field

Preventing duplicate runs. If a job that cron runs every 5 minutes starts taking 6 minutes, two overlap. Without a lock, they write the same file at the same time and the data is corrupted. Use flock or a mkdir-based lock to make the second run exit immediately, but it is important to distinguish "why it didn't run" by exit code. That way monitoring can tell "failure" apart from "skipped because it was already running".

Exit code conventions. Everyone knows 0 for success and 1 for a general failure. Beyond that, the project decides. If you borrow the sysexits convention, 64 is a usage error and 66 is a missing input file. More important than the values themselves is using them consistently and writing them down in documentation.

Retry conditionally. A network call may succeed if you try again even after a failure. However, three conditions attach to retrying. There must be a maximum count, you must rest between attempts (if you don't, you push the other side harder), and you must stop immediately on success. And you must first confirm whether the operation is safe to retry (idempotent) - retrying a payment request pays twice.

shellcheck is a free reviewer. It catches missing quotes, unused variables, and dangerous patterns. If you put it in CI, review time drops noticeably.

Where set -e does not look after you

It is easy to think you are safe once you write set -euo pipefail at the top. But set -e has several exceptions defined by the specification, and accidents always happen in those places.

A command used as a condition does not stop when it fails. Commands in if, while, &&, ||, and to the left of ! are not termination conditions, because their failure is material for a decision. Up to here this is intended behavior. The problem is when you put a function in that position.

setup() {
  mkdir -p /srv/data      # 여기서 실패해도
  cp config.yml /srv/data # 이 줄이 실행된다
}
if setup; then echo "준비 완료"; fi

The moment a function is called as the condition of an if, set -e is turned off throughout the inside of that function. The function returns the exit code of its last command, so whatever failed earlier, as long as cp succeeds, the "ready" message is printed.

A failure in command substitution is swallowed by the assignment.

VERSION=$(cat /etc/app/version)   # 파일이 없어도 스크립트는 계속 간다

This is because the assignment statement VERSION=... succeeded. The exit code is that of the assignment, not of cat. If you put local or export in front, it is hidden even more thoroughly. If you must obtain the value, write it in two parts.

VERSION=$(cat /etc/app/version) || exit 1
: "${VERSION:?버전 파일이 비어 있습니다}"

Without pipefail, the front of a pipe might as well not exist. With set -e alone, in curl ... | jq ., even if curl dies, it passes if jq succeeds. However, once you turn on pipefail, a pipe cut short by head fails with SIGPIPE, so in such places you write the intent with || true.

Leave cleanup to trap. The more a script stops midway, the more temporary files it leaves.

tmp=$(mktemp -d) || exit 1
trap 'rm -rf "$tmp"' EXIT

EXIT runs on normal exit, on exit caused by set -e, and on any exit call. If you also want to catch INT/TERM, list them together, but do not use commands inside it that can fail again. If the cleanup function fails, what is left is not a temporary file but an infinite loop.

What you will do in the next lab

You will prove through behavior that strict mode really works, propagate pipe failures with pipefail, clean up temporary files with trap, prevent duplicate runs with a lock, safely handle paths that contain spaces, and build a backup script and a retry helper that follow the exit code convention. Finally, you will write an audit tool that inspects other scripts.