Plugin Architecture and Operations — Keeping the Portal Alive
In one line
Backstage's extensions come in two kinds of plugins, frontend and backend, and recently the new backend system has greatly simplified plugin installation. And what you actually run into in operations is not flashy features but authentication, identity sync, the catalog processing interval, and the database.
Why this was needed
A portal is never "install it and you are done." As the number of attached systems grows, three things become problems.
- Who is who — which User entity in the catalog is the person who logged in, and which Group do they belong to. If this is wrong, the "my services" list is empty and every ownership-based feature becomes meaningless.
- How fresh the information is — when does the catalog re-read the repositories? Too often and you hit API limits; too rarely and people stop trusting the portal.
- What holds state — a portal looks stateless, but the catalog and the scaffolder's task history are in the database.
How it works
Two kinds of plugins and the new backend system
| Frontend plugin | Backend plugin | |
|---|---|---|
| Form | React components | Node.js modules |
| What it provides | Routes, tabs and cards on entity pages, home widgets | HTTP endpoints, catalog processors and entity providers, scaffolder actions |
| Credentials | Must not have any | They live here |
In the old backend you had to wire each plugin's router by hand and pass in dependencies yourself. The new backend system flipped this: the backend only registers plugins and modules, and they receive what they need (logger, configuration, database, authentication, scheduler, and so on) through dependency injection. As a result, installation shrank to "add a package + one line of registration," and the extension points between plugins became clear. CBA asks about this transition (old backend → new backend system).
Authentication and identity sync
You need to distinguish two things.
- Authentication — login. You attach providers such as GitHub, Google, Okta, Microsoft, and OIDC.
- Identity resolution (sign-in resolver) — which User entity in the catalog to map the logged-in person to. Whether to match by email, match by username, or create one if none exists.
Organization data ingestion is added to this. It periodically reads the list of users and teams from places such as GitHub Org, LDAP, and Microsoft Entra and puts them in as User/Group entities. Only when this is in place does spec.owner: group:team-checkout connect to a real list of people.
The most common adoption failure happens here. If you reference a Group as owner that is not in the catalog, the entity ends up with a broken relation, and if the "my services" page is empty, the user does not come back a second time. It is better not to open the portal to the public before the ownership graph is filled in.
The link in the Kubernetes plugin
How workloads are shown on an entity page comes up on the exam.
- Attach the annotation
backstage.io/kubernetes-id: <값>to the entity in the catalog (the placeholder is the shared value). - Attach the same value to the cluster's workloads as the label
backstage.io/kubernetes-id: <값>(the placeholder is the value). - The backend queries the registered clusters with that label selector, gathers the results, and returns them to the frontend.
The direction is easy to confuse — the entity gets an annotation and the workload gets a label. There is also a backstage.io/kubernetes-namespace annotation that makes it look things up by namespace instead of by label. And since the backend makes the queries, the cluster credentials exist only on the server.
Operations — processing interval and database
The catalog produces data in two stages.
- Entity provider — discovers what is where and feeds it in (for example, scanning a GitHub organization).
- Processor — reads those raw entities, validates them, computes relations, and creates derived entities.
This processing repeats periodically. A shorter interval improves freshness but increases SCM API calls and hits the rate limit. At large scale, you use a combination of updating immediately through webhooks or events and keeping the interval itself long.
For development convenience the database can start as in-memory SQLite, but everything disappears on restart. In production PostgreSQL is effectively the standard, and the catalog, the scaffolder's task history, and the search index go into it. This means a portal is not a stateless application, and you need a backup and recovery plan.
And an easily forgotten operations item is upgrades. Backstage changes actively, and your app is a code tree that you own. If you put it off for a few months, the cost of upgrading all at once later grows sharply. Upgrading a little at a time, regularly, is the only sustainable way.
What it looks like in the field
The author's homelab has a verification record of running PostgreSQL 18 with CloudNativePG as a 2-instance streaming replication. If you put a portal on top, that database becomes the home of the catalog. And the most painful lesson of this cluster applies here exactly — with three control plane nodes the etcd quorum is in place, but controlPlaneEndpoint is the first node's physical IP, so when that node dies, the data is alive but nobody can connect to the API. Data availability and access availability are separate.
A portal is the same. Even if you replicate the catalog database, if the portal app cannot come up, nobody gets any information. Conversely, even if the portal is alive, if the catalog processing cycle has stopped, the information on the screen goes stale silently. The latter is more dangerous — because outages are visible but stale data is not.
One more. The "the state Ready and actually working are different claims" confirmed repeatedly on this cluster is the core instinct of portal operations. KubeVirt had every component AllComponentsReady, yet the VM did not start. A portal can likewise have a green health check while the catalog processor is failing silently. That is why you must watch the processing success rate and the time of the last refresh as metrics.
What to check in the next quiz
This module ends with a quiz. If you re-check the pairing of the backstage.io/kubernetes-id annotation (entity) and label (workload) that you created in the earlier module's lab, what the plugin actually does will come into view.