Naming an Artifact by Hash Does Not Make the Build Reproducible
One-line summary
Naming an artifact by a hash and making the same source produce the same bytes are two different stories. The former is about attaching a label accurately, and the latter is about making the build "a function where the same input gives the same output". A great many pipelines attach only the label and claim reproducibility.
Why this is needed
"We tagged it with the source hash, so it is reproducible" is something you often hear but is rarely proven. If you build the same commit twice and directly compare the SHA-256 of the two artifacts, they usually differ. Yet the name is the same. An artifact with the same name and different content breaks three things at once. The cache lies, the promotion model collapses, and in an incident retrospective nobody can answer the question "is what went out then exactly this?".
What reproducibility actually buys is three things. First, you can trust the cache. Only when the premise "the same key gives the same result" holds can you use a cache hit as grounds for "it is fine to skip the build". Second, you can check an artifact's provenance by recreating it. If you rebuild from the same source with the same procedure and the bytes match, you can verify that the artifact came from that source without even a signature. Third, you can revive that time's artifact during an incident. Even if it has been deleted from the repository, as long as the source and procedure remain, you get the same bytes again.
Where nondeterminism comes in
The sources are almost fixed. Even when you meet a new language or tool, the list does not change much.
- Build time. If you embed the build time in the artifact, it differs every time. This includes compressed-file headers, the modification times of files in an archive, and even comments in generated source.
- File order. The order in which a directory is walked is decided by the file system, and that order is not guaranteed. If the order in which things go into the archive changes, the bytes change.
- uid/gid and permissions. If the build account differs, the owner information inside the archive differs. If the umask differs, the permission bits differ.
- Locale and sorting. The sort result changes with
LC_ALL, and the sort result is reflected in list files and archive order. - Absolute paths. If the build directory path goes into the artifact, the result differs even if only the workspace name differs.
- gzip header. gzip by default writes the original file name and time into the header.
- Parallel execution order. If several jobs append to a single result file without ordering, the order is different every time.
How to pin it down
The key convention is SOURCE_DATE_EPOCH. The official documentation defines this variable as "the last modification time of something, usually of the source code, expressed as the number of seconds since the Unix epoch". It was agreed that when a build tool has to write a time, it uses this value instead of the current time. The value is usually taken from the time of that commit.
You have to be especially careful with archives. The form the official documentation recommends is as follows.
export SOURCE_DATE_EPOCH="$(git log -1 --pretty=%ct)"
tar --sort=name \
--mtime="@${SOURCE_DATE_EPOCH}" \
--owner=0 --group=0 --numeric-owner \
--pax-option=exthdr.name=%d/PaxHeaders/%f,delete=atime,delete=ctime \
-cf product.tar build
gzip -6 -n < product.tar > product.tar.gz # -n 은 원본 이름과 시각을 빼고 압축한다
--sort=name fixes the file order, --mtime the time, --owner/--group/--numeric-owner the owner, and --pax-option removes the atime/ctime that get mixed into the PAX headers. This combination requires GNU tar 1.28 or later. And the verdict is made not by eye but by hash. Making it twice and checking whether sha256sum is the same — there is no other way to confirm reproducibility.
What a digest guarantees and what it does not
The OCI image specification defines a digest as "a collision-resistant hash over bytes" and writes that this enables content addressing. If you received the digest by a safe route, then even for content received from an untrusted place, you can confirm it has not been tampered with by recomputing the hash and comparing. The format is 알고리즘:인코딩된값 (algorithm:encoded-value), and the specification recommends verifying content from untrusted sources against the digest before using it.
This property is built into the storage format as well. The OCI image layout specification requires that the content placed at blobs/<알고리즘>/<인코딩된값> (algorithm, encoded value) must match the digest 알고리즘:인코딩된값. In other words, the file name points not to a location but to the content itself. The layout must also contain oci-layout and index.json, which serves as the entry point.
Here you must draw the line exactly. What a digest guarantees goes only as far as "the bytes I received are indeed those bytes." It says nothing about which source and what procedure produced those bytes. To claim that connection, you must separately record what the build took as input and what procedure it went through, and that is what provenance does. A reproducible build is its counterpart in that it makes that claim something anyone can verify by recreating it.
What it looks like in the field
- After clearing the cache, the artifact hash changed. It means the cache had been mixing into the result, and the tests all along were of the thing the cache made.
- For the same commit, what CI built and what was built by hand differ in size by a few bytes. It is usually the file order of the archive or the owner information.
- When you rebuilt an old commit in order to roll back, something different from before came out. An input that was not pinned moved in the meantime.
References
- SOURCE_DATE_EPOCH specification: https://reproducible-builds.org/docs/source-date-epoch/
- Nondeterminism of archives: https://reproducible-builds.org/docs/archives/
- tar(1): https://man7.org/linux/man-pages/man1/tar.1.html
- OCI descriptor (digest): https://github.com/opencontainers/image-spec/blob/main/descriptor.md
- OCI image layout: https://github.com/opencontainers/image-spec/blob/main/image-layout.md
- SLSA provenance: https://slsa.dev/spec/v1.0/provenance
What you will do in the next lab
With only shell, tar and sha256sum, you directly compare whether the same source produces the same bytes. First you make an archive twice with no countermeasures and confirm that the hashes diverge, then you pin the time, order, owner and gzip header one at a time while tracing with your eyes which item moved how many bytes. At the end you organize it into a build script that pulls SOURCE_DATE_EPOCH from the commit time, and judge by whether the hashes of two runs are the same. Since containers cannot be started in this Pod, for the image side you deal with it by reading and comparing the digest of an existing oci-archive with skopeo.