Last Week Had a Better Model. Nobody Can Find It
Versions Are Facts, Aliases Are Roles
In one line
A registry is not a warehouse where model files are collected, but a device that answers "what is being served right now" in one place, using versions and aliases.
Why this was needed
Once experiments are well organized, the next question comes. "So which one is running in production right now?" Surprisingly, this question is the one that gets stuck most often. The model file sits somewhere in S3 as
model_final_v3_really.pkl, the serving configuration has that path written as a string, and which run that file came from is in the creator's head. When that person goes on vacation, the organization cannot explain its own model.
Rollback becomes hard for the same reason. If a deployment is "changing a path string," then rollback is also "changing the path string back," so when you are in a hurry, there is no record of who rolled back to which file. When you try to draw a timeline after an incident, you end up digging through chat logs.
How it works
The MLflow model registry documentation organizes this problem into four concepts (MLflow Model Registry).
| Concept | Description |
|---|---|
| Registered Model | One model with a name. It holds versions, aliases, and tags |
| Model Version | Each time one is registered under the same name, the number goes up by 1. The first registration is 1 |
| Model Alias | A changeable name that points to a specific version. Something like champion |
| Tag | A key-value label. It can also be attached to a version (for example, validation_status: approved) |
The key is the split between versions and aliases. A version is a fact that does not change once created,
and an alias is a knob that points to "what currently holds this role". If you have production look at models:/MyModel@champion,
a deployment becomes changing the version the alias points to. The documentation explains this
approach as follows: if you give an alias to the version that will receive production traffic and target that alias,
you can change the serving target just by reassigning the alias to a different version.
Lineage comes with this. Each registered version is linked to the run or logged model that produced it, so you can trace back which data and parameters it was trained with. The reason you recorded runs carefully in the earlier module pays off here — the registry stands on top of tracking, and if tracking is empty, the registry cannot be more than a file list.
registry.json 버전 1, 2, 3 … (변하지 않는 사실)
aliases.json champion → 2 (지금 누가 그 역할인가)
audit.jsonl champion 1→2, 2→1 … (어떻게 여기까지 왔는가)
Moving an alias comes with one discipline: the fact that you moved it must be written down somewhere. If you look only at the alias file, you can learn only the current state, "champion is now version 2," and the fact that it was version 1 until yesterday does not remain. So a change history goes alongside the alias. If you stack up, one line at a time, the value before the change, the value after, and who moved it and why, then replaying that record from the beginning must produce the current state exactly. If it does not, someone touched it without leaving a record.
A tag is a lighter marker than an alias. You express state by attaching validation_status: pending to a version awaiting validation and
approved to one that passed. An alias attaches only one per role, but tags can be attached in
multiples, which makes them good as conditions that automation reads.
What it looks like in the field
A team that hard-coded the version number in the serving configuration instead of using an alias has to rerun the deployment pipeline when it rolls back. The more urgent it is, the longer that pipeline takes. A team that uses aliases, in the same situation, changes only where it points, and that change stays in the history as one line.
There is a mistake in the opposite direction too: creating too many aliases. If champion, stable,
prod, prod-real, and prod-new all exist at once, nobody knows
which of them is the real one. Keep aliases only as many as there are roles. Usually one for production and one for the challenger is enough.
The third is the habit of deleting versions. If you delete a failed version from the registry, the list becomes clean, but the reason that version failed disappears with it. The person who tries a similar approach next quarter hits the same wall.
The fourth is the misunderstanding that a registry and an artifact store are the same thing. Where the bytes of the model file are placed
is the store's job, and what the registry answers is "which version number are those bytes, what role do they hold now, and where did they come from." If you mix the two,
deployments break every time you move a file, and conversely, when the file stays the same but its role changes, that fact is
recorded nowhere. This is why the lab keeps registry.json and the model file separate.
What to check in the next quiz
You check what versions and aliases are each responsible for, what happens to the number when you register a new model under the same name, and why rolling back should be a matter of moving an alias.