TT Lab
Get started
Learn Learning paths Courses

CI/CD Pipelines

Naming an Artifact by Hash Does Not Make the Build Reproducible

Continue in TT Lab

One-line summary

Naming an artifact by a hash and making the same source produce the same bytes are two different stories. The former is about attaching a label accurately, and the latter is about making the build "a function where the same input gives the same output". A great many pipelines attach only the label and claim reproducibility.

Why this is needed

"We tagged it with the source hash, so it is reproducible" is something you often hear but is rarely proven. If you build the same commit twice and directly compare the SHA-256 of the two artifacts, they usually differ. Yet the name is the same. An artifact with the same name and different content breaks three things at once. The cache lies, the promotion model collapses, and in an incident retrospective nobody can answer the question "is what went out then exactly this?".

What reproducibility actually buys is three things. First, you can trust the cache. Only when the premise "the same key gives the same result" holds can you use a cache hit as grounds for "it is fine to skip the build". Second, you can check an artifact's provenance by recreating it. If you rebuild from the same source with the same procedure and the bytes match, you can verify that the artifact came from that source without even a signature. Third, you can revive that time's artifact during an incident. Even if it has been deleted from the repository, as long as the source and procedure remain, you get the same bytes again.

Where nondeterminism comes in

The sources are almost fixed. Even when you meet a new language or tool, the list does not change much.

How to pin it down

The key convention is SOURCE_DATE_EPOCH. The official documentation defines this variable as "the last modification time of something, usually of the source code, expressed as the number of seconds since the Unix epoch". It was agreed that when a build tool has to write a time, it uses this value instead of the current time. The value is usually taken from the time of that commit.

You have to be especially careful with archives. The form the official documentation recommends is as follows.

export SOURCE_DATE_EPOCH="$(git log -1 --pretty=%ct)"
tar --sort=name \
    --mtime="@${SOURCE_DATE_EPOCH}" \
    --owner=0 --group=0 --numeric-owner \
    --pax-option=exthdr.name=%d/PaxHeaders/%f,delete=atime,delete=ctime \
    -cf product.tar build
gzip -6 -n < product.tar > product.tar.gz   # -n 은 원본 이름과 시각을 빼고 압축한다

--sort=name fixes the file order, --mtime the time, --owner/--group/--numeric-owner the owner, and --pax-option removes the atime/ctime that get mixed into the PAX headers. This combination requires GNU tar 1.28 or later. And the verdict is made not by eye but by hash. Making it twice and checking whether sha256sum is the same — there is no other way to confirm reproducibility.

What a digest guarantees and what it does not

The OCI image specification defines a digest as "a collision-resistant hash over bytes" and writes that this enables content addressing. If you received the digest by a safe route, then even for content received from an untrusted place, you can confirm it has not been tampered with by recomputing the hash and comparing. The format is 알고리즘:인코딩된값 (algorithm:encoded-value), and the specification recommends verifying content from untrusted sources against the digest before using it.

This property is built into the storage format as well. The OCI image layout specification requires that the content placed at blobs/<알고리즘>/<인코딩된값> (algorithm, encoded value) must match the digest 알고리즘:인코딩된값. In other words, the file name points not to a location but to the content itself. The layout must also contain oci-layout and index.json, which serves as the entry point.

Here you must draw the line exactly. What a digest guarantees goes only as far as "the bytes I received are indeed those bytes." It says nothing about which source and what procedure produced those bytes. To claim that connection, you must separately record what the build took as input and what procedure it went through, and that is what provenance does. A reproducible build is its counterpart in that it makes that claim something anyone can verify by recreating it.

What it looks like in the field

References

What you will do in the next lab

With only shell, tar and sha256sum, you directly compare whether the same source produces the same bytes. First you make an archive twice with no countermeasures and confirm that the hashes diverge, then you pin the time, order, owner and gzip header one at a time while tracing with your eyes which item moved how many bytes. At the end you organize it into a build script that pulls SOURCE_DATE_EPOCH from the commit time, and judge by whether the hashes of two runs are the same. Since containers cannot be started in this Pod, for the image side you deal with it by reading and comparing the digest of an existing oci-archive with skopeo.