Last Week Had a Better Model. Nobody Can Find It
The Minimum Record That Rebuilds a Run
In one line
The smallest unit of record that lets you bring one training session back to life is the run, and experiment tracking is a device that forces that record into a file instead of into a person's memory.
Why this was needed
Building a model is usually quiet work. In a Jupyter notebook you raise the learning rate a little, try more epochs, change the seed and run again. When the number improves you are happy, and when it gets worse you revert. What actually remains from this process is only the output of the last cell you ran. Then two weeks later someone asks, "Didn't you say last week that something better came out?"
When you then try to reproduce it, three things become problems at once: which parameters, which data, and whether the training code back then is the same as the code now. If even one of the three is missing, the same number will not come back. Worse, a similar number does come back. If it were completely different, you would know something was wrong, but when 0.76 comes out instead of 0.78, people move on thinking, "Was it always about this much?" A record skipped over that way is never recovered.
The MLflow documentation defines a run as "one execution of data science code" and says each run records metadata (metrics, parameters, start and end times) together with artifacts (output files such as model weights and images) (MLflow Tracking). The reason the definition looks like this is clear. You need all four of these to set that run up again, and if even one is missing, that run becomes "a rumor that it existed".
How it works
What a tracking tool actually does is simple. When training starts, it opens a run, each time the code
calls log_param and log_metric it writes that value to the store, and when training ends it closes the run.
If you do not set up any server, records accumulate in a local directory. If you want to change the storage location,
you configure the tracking environment separately.
import mlflow
with mlflow.start_run():
mlflow.log_param("lr", 0.001)
# 학습 코드
mlflow.log_metric("val_loss", val_loss)
Once several runs have accumulated, what comes next is querying. MLflow supports lookups such as
"the run with the lowest validation loss in this experiment" through MlflowClient.search_runs, and from MLflow 3 on,
search_logged_models can find models by filtering on metric and parameter conditions with a SQL-like string.
What matters here is that the sort criterion is written as code. If a person
scans the table by eye and picks, choosing again next week gives a different answer.
The list of things to record differs a little from tool to tool, but working backward from the goal of reproducibility,
you usually arrive at the same place: parameters, metrics, and the identity of the input. MLflow's
dataset tracking attaches to a run, for this third item, an object that holds each dataset's name, its digest (fingerprint), and its source
location (MLflow Dataset Tracking).
A fingerprint is needed because a file name is not a version. train.csv is train.csv both yesterday and
today.
Autolog is a feature where the library writes all of this for you, and for a framework on the supported
list you turn it on with the single line mlflow.autolog()
(Automatic Logging). It is convenient, but
if you do not know what gets written, you do not know what is missing either. So it is better to write things by hand the first time.
What it looks like in the field
The most common accident is not failing to use a tool but using it halfway. The parameters are written down but there is no data version. There are metrics but no code commit. Then the table looks impressive but no row can be reproduced. Because there is a table, nobody feels a problem, until a demand such as regulatory response or incident investigation arrives, "prove how this model was built," and everything comes out at once.
The second is the storage location. If you run training in a container and write the results inside the container,
the record disappears together with that Pod. The lab Pod has no volume either, so when the session ends,
/root is gone entirely, and this is not an inconvenient restriction but a miniature of reality. The record must stay
outside the place where the computation happened.
The third is names. The habit of naming run identifiers test, test2, test_final is
convenient that day but tells you nothing a month later. An identifier exists not for people to read
but to link with other records, so short and non-overlapping is enough. Instead,
write "what was tried" in the parameters and notes.
What you will do in the next lab
You build an experiment ledger yourself from a single JSON Lines file. For each run you record the parameters, the metrics, and the SHA-256 fingerprints of the code and data, pick the best run with code by a target metric, and try to recreate the same number from that record alone. Finally, you open the log of a run that vanished without its parameters being recorded, and count as items what is missing that makes it impossible to reproduce.