A Catalog Rots Over Time, Not at Registration
In one line
A catalog collapses not at first registration but half a year later. Rot almost always shows up in three shapes: entities with missing required fields, references that point to targets that do not exist, and dependency relations that point at each other. All three can be found statically, without bringing up the portal.
Why this was needed
When you first open a portal, it is usually clean. A person put things in by hand, and the person who put them in still remembers the service. The problem is what comes after.
- Teams merge and
team-paymentsdisappears, but twelve entities that list that team as owner remain. - A service moved to another system, but
spec.systemstill has the old name. - A temporary database was deleted, but
dependsOnstill has that Resource.
What matters here is that Backstage does not report these as errors. The catalog flags an entity that violates the schema as an error, but a reference to a nonexistent target simply becomes "a relation that was not computed." The screen shows only a blank cell, with no red text.
And the path by which a catalog dies is always the same. References break, the graph fragments, the screen looks empty, people stop looking, and nobody fixes it. Because an empty screen does not look like an outage, nobody reports it.
How it works
There are four things to diagnose
First, missing required fields. Required fields differ by kind. Component and API must all have spec.type, spec.lifecycle, and spec.owner; Resource must have spec.type and spec.owner; and System and Domain must have spec.owner. On top of that, every kind needs apiVersion, kind, and metadata.name. If you memorize "three required fields" wholesale, you will be wrong when the kind changes.
Second, a broken owner. This is the most expensive defect. When the owner is broken, outage calls, vulnerability tickets, cost attribution, and retirement decisions all lose their destination. And the cause is often not the entity file. Organization data sync stopped, so the Group itself never made it into the catalog. If you fix only the file, the same thing happens again next week.
Third, broken references. spec.system, spec.domain, spec.dependsOn, and spec.providesApis are the targets. The reason to look at these separately from the owner is that the response differs. The owner is an organizational problem, and the rest are usually cases where a name changed or the entity has not been created yet.
Fourth, dependency cycles. The state where A dependsOn B and B dependsOn A. This makes impact analysis lose its meaning. When you ask "what breaks if I take this down?", the answer returns to itself. The cause is usually a misunderstanding. People think that because A calls B, A must also be written on B's side, but for relations the catalog computes both directions from a declaration on one side.
Normalize references before comparing
If you skip this, the checker reports perfectly good references as defects.
team-orders → group:default/team-orders
group:team-orders → group:default/team-orders
group:default/team-orders → group:default/team-orders
The three point to the same thing. Each field has a default kind, and the default for namespace is default. If you compare strings as they are, the three notations all become different things. Normalizing and then comparing is the first step of this work.
There is an order to fixing
- Organization data first. Fill in the missing Groups and Users.
- Groupings next. Create the missing Systems and Domains, or fix references to the name of the place they moved to.
- Individual entities last. Fill in the missing required fields and clean up the remaining references.
If you reverse the order, you fix the same error twice. If you fix each individual entity's owner one by one and then create the Group, the values you just fixed become awkward again.
And one thing you must keep to. Deleting entities with defects to make the list clean is not fixing. The purpose of a catalog is to show everything we have, without omission. A service you made invisible also breaks at 3 a.m.
Stop it with a gate
Because catalog-info.yaml lives in the code repository, you can run the check in a PR. If you block files before they are merged, the speed at which the catalog rots drops sharply. A gate must return 0 for clean input and a non-zero value for input with defects.
What it looks like in the field
This repository has a painful record about gates. A check item saying "there is no plaintext password in the documentation" was green for a long time, but it was actually looking at only one shape. In the meantime, the administrator passwords of several tools were in plaintext in two documents, and since all three had different shapes they matched no pattern at all.
The lesson was not about regular expressions. Once you build a gate, you must actually feed it what it is meant to block and check that the red light turns on. If you only check that things pass, that gate goes on for months as nothing at all. A catalog lint is exactly the same. If you are satisfied just seeing 0 on a clean directory, you ship a gate that gives 0 even when defects come in.
The same sense connects to this cluster's repeated lesson that "the state Ready and actually working are different claims." The catalog screen looking normal and the relations inside it actually being connected are different facts. The latter cannot be seen with the eyes, so you must compute it.
What you will do in the next lab
In /root/cba-audit/incoming/, you reproduce as they are eight catalog entries you have taken over. Four kinds of defects are deliberately included. Then you diagnose each of the missing required fields, broken owners, broken references, and dependency cycles and extract them as lists, create a fixed copy, and make all four diagnoses come to 0. Finally you write a lint script that performs the same checks, test it on both clean input and input with defects, and record the audit result on the cluster.