TT Lab
Get started
Learn Learning paths Courses

CBA — Backstage Associate

Entities and Relations — Why Ownership Is the Catalog's Heart

Continue in TT Lab

In one line

A catalog is not a list of services but a relationship graph. Understanding what question each kind exists to answer is more useful, for both the exam and real work, than memorizing the entity kinds.

Why this was needed

With only a "list of services," you cannot answer questions like these.

That is why Backstage divides entities into several kinds and stores the relations between them.

How it works

Eight entity kinds

kind Question it answers Representative fields
Component A piece of software we build and deploy spec.type (service/website/library), spec.lifecycle, spec.owner, spec.system
API An interface a component exposes or consumes spec.type (openapi/asyncapi/graphql/grpc), spec.definition
Resource Infrastructure a component needs spec.type (database/s3-bucket/queue)
System A group of entities that work together spec.owner, spec.domain
Domain A business area that spans several systems spec.owner
Group A team or organizational unit spec.type (team), spec.children, spec.profile
User A person spec.memberOf
Location A signpost that points to where other entities are spec.type, spec.targets

Template (for the scaffolder) is added on top of these. And spec.lifecycle is a free string, but by convention experimental / production / deprecated are used — they serve to filter services slated for retirement out of the list.

Relations are computed

This is an important point. What you write in catalog-info.yaml are declarations such as spec.owner, spec.system, and spec.providesApis, and the catalog reads them and computes and stores the bidirectional relations.

spec.owner: group:team-checkout      →  ownedBy / ownerOf
spec.system: commerce                →  partOf / hasPart
spec.providesApis: [checkout-api]    →  providesApi / apiProvidedBy
spec.consumesApis: [payments-api]    →  consumesApi / apiConsumedBy
spec.dependsOn: [resource:checkout-db] →  dependsOn / dependencyOf

So even if you write providesApis on only one side, the API entity page shows "the components that provide this API." You do not need to write the opposite direction by hand. If the exam asks "Must the relation be written in both files?", the answer is no.

Entity reference format

When you write a relation, there is a fixed string format for pointing to another entity.

[<kind>:][<namespace>/]<name>

Of the three parts, kind and namespace can be omitted, and when omitted a context-dependent default applies. The default for namespace is default.

What you wrote Interpretation
team-checkout The context's default kind + the default namespace
group:team-checkout group:default/team-checkout
group:payments/team-checkout The namespace is also explicit

A field such as spec.owner has a fixed default kind, so writing just team-checkout works, but explicitly adding group: is much safer in review. That is because a human reader can see right away whether it is a team or an individual.

Why catalog-info.yaml lives next to the code

There are three ways to fill the catalog.

  1. Static registration — list URLs in the Backstage configuration file. It works at small scale but becomes unmanageable as things grow.
  2. Location entity — a signpost entity points to other files. Useful for grouping hierarchically.
  3. Discovery — scan the organization's repositories, find catalog-info.yaml, and register automatically. This is the right answer in practice.

For option 3 to work, the file must be in the same repository as the code. This arrangement has a large effect.

Conversely, if you pile the catalog into a single central repository, nobody has any motivation to edit that file. It is the most common way a catalog rots.

Why ownership is the heart

If only one thing in the catalog had to be accurate, it would be the owner.

That is why the owner must be a team (Group), not a person (User). People leave, and teams are taken over. It is a point that comes up often on the exam, and it is also the first reason catalogs collapse in practice.

What it looks like in the field

The habit, in the author's homelab, of attaching standard labels such as app.kubernetes.io/name and app.kubernetes.io/part-of to Kubernetes workloads is exactly the same thinking as this catalog. Helm chart best practices also recommend attaching six standard ones: app.kubernetes.io/name, instance, version, component, part-of, and managed-by.

That is, expressing the same ownership and membership information consistently in two places, cluster labels and catalog entities, is what it looks like in practice. And what connects the two is Backstage's Kubernetes plugin — if you attach the annotation backstage.io/kubernetes-id to an entity, it finds workloads in the cluster that have a label with the same value and shows them on the entity page.

Semantic labels such as gpu.homelab/tier=xlarge, attached to the homelab GPU nodes, are in the same spirit. Translating a physical fact into a vocabulary humans can use for decisions — it is no different from why the catalog groups services into systems and domains.

What you will do in the next lab

In /root/cba-catalog/, you write entity files for Component, API, Resource, System, Domain, Group, User, and Location yourself. Then you express the same ownership information as labels on a real cluster and check it with kubectl, and finally you extract the reference strings of all entities in their normalized form.