The Build Was Green — So Who Put That Library In?
We picked seven; twenty-five got installed
In one line
The list of dependencies lives not in the file we wrote by hand (package.json) but in the
file that records what was actually installed (package-lock.json), and the difference in
count between the two explains almost everything about why supply chain security is hard.
Why this matters
The first question that arrives the morning after an incident is always the same: "Do we use that?" The question is hard not because the answer is complicated, but because there is nowhere to look up the answer. The repository has a dependency list we wrote by hand. But the names in it each pull in other names, and those names pull in yet more names. Five lines become twenty-five without any event at all. That is simply what happens when you install.
So "what is in there?" becomes a question you have to ask again after the build is finished. No amount of reading the source will answer it. The answer is decided at build time, and once that moment passes, it is gone. The lockfile and the SBOM were created to capture that moment.
There is one more distinction to make. Not everything that is installed ships in the deployment. Test tools, bundlers, and linters run only on the build machine and never reach users. If you count the two in the same bucket, the number of "things to fix" doubles, and nobody looks at an inflated number.
How it works
A lockfile pins down "what was fetched". In lockfileVersion 3, npm's package-lock.json uses
a map called packages, where the install location is the key and the root
project goes in under the empty-string key. Each entry has a version, a resolved field that records where it was fetched from,
and an integrity field that guarantees the bytes fetched are the right ones.
An entry that belongs only to the development-only tree gets dev set to true
(npm docs).
The most common mistake here is not excluding the root entry when counting, which produces one extra.
An SBOM carries that list in a format that can travel across tools and organizations. There are two
widely used formats. CycloneDX is maintained jointly by OWASP and Ecma and standardized as ECMA-424, and
it represents components, services, and a dependency graph that holds both direct
and transitive dependencies
(CycloneDX overview). The
first two fields of the JSON document are bomFormat and specVersion, and the reason bomFormat is
pinned to the single value "CycloneDX" is interesting: BOM files have no naming convention and JSON schemas have
no namespaces, so you must be able to tell what format a file is just by looking at it
(CycloneDX 1.6 JSON).
SPDX does the same job with a different vocabulary. For each package it records a name, version, checksum, and
external references such as purl, and it expresses relationships between elements with
relationship names such as DEPENDS_ON and CONTAINS
(SPDX 2.3 Package Information).
Whichever you use, three fields really matter: what the subject is, what it contains,
and what pulled in what. An SBOM without the last field is just a flat list,
and it cannot answer "why is this here?"
What it looks like in the field
The scene you see most often is that SBOMs are generated but nobody compares two of them. Files pile up with every release, but no one looks at the difference between the last release and this one. Yet most supply chain incident signals are in that difference — a new name in the list even though nobody added a dependency. Someone left the version range of an indirect dependency wide, and the new name came along when a new version was published.
The second scene is scan results being ignored because they are bloated. If you count development-only dependencies too, the number doubles, and a list that has doubled does not get read by the development team. It gets read only if you send it split up.
The third scene is when the list does not say what its subject is. components is packed with twenty-five lines,
but nowhere does it say which artifact the list belongs to. Then someone who opens that file
months later has no way to confirm "is this really the release from back then?" The subject must be recorded not by name but by
digest — a name can point to different content under the same name, but a hash cannot.
The signing and provenance attestation in the following modules all sit on top of this "subject" field.
Finally, when you generate the list is also a design decision. A list generated by reading the source and a list generated by scanning the build output are different. The former is "what did we decide to use" and the latter is "what actually shipped". The two are usually similar, but the day they diverge is the day the incident happens.
What you will do in the next lab
From a single lockfile, you count the list yourself and write it out in CycloneDX format. You trace the path by which a package that nobody chose came in, and compare against the SBOM of the previous release to count what was added, removed, and changed. If you do by hand, once, what a tool does in one second, you can later see what is missing from the list the tool produces.