Skip to content

Collecting a Backstage catalog

backstage-pull fetches catalog-info.yaml from the repositories a manifest names, and records which repository and commit each one came from. It produces no RDF; backstage2linkedarchi converts what it writes.

Why a separate tool

A Backstage catalog is not one file in one place. It is one descriptor per service, each in the repository of the service it describes, so a graph of the whole estate is built by fetching all of them. That makes it the only input in this project that is pulled rather than committed beside the model, and it has a consequence that is easy to miss:

Only the request knows the commit. A fetch at a ref resolves to a commit, and that fact exists only at the moment of the request. Once the file is sitting in a working directory it is indistinguishable from any other file, and the repository doing the converting can say nothing about where it came from.

So --git-provenance auto is wrong for a pulled catalog, and wrong in a way that produces confident output. It describes the repository the conversion is running in, so it records paths that repository has never held and composes forge URLs that return 404 by construction:

https://git.example.org/group/models-backstage/-/blob/b292b0ee…/catalogs/orders/order-service.yaml
#                      ^ the pipeline's repo                    ^ a path only the pull ever created

The pull writes the answer down instead, and the converter reads it.

The three artifacts

File Owner Committed Rewritten
sources.yaml you yes never — you maintain it
the descriptors upstream no — gitignore them every pull
catalog-sources.yaml the pull yes every pull
catalog-index.yaml you yes written once, then never

The descriptors are copies of files other repositories own. Committing them would make this repository a second, stale source of truth for them.

The source map and the index are two files rather than one because they have opposite ownership. A commit changes on every pull, so it cannot live in a file a pull must not overwrite; editorial decisions — model id, lifecycle state, extension statements — accumulate across dozens of lines and must not be discarded by a regeneration.

The manifest

defaults:
  ref: HEAD                    # GitLab reads HEAD as "the default branch"
  path: catalog-info.yaml
sources:
  - id: order-service
    group: orders
    repo: https://git.example.org/group/order-service
  - id: payment-service
    group: payments
    repo: https://git.example.org/group/payment-service
    ref: release-2026-08
    path: catalog/catalog-info.yaml

id names the file the pull writes and joins the three artifacts. group is an output directory only and does not reach the graph. Both are held to the diagram-index id pattern, because the id also becomes a view id and therefore an IRI path segment — so a bad one fails the pull, naming the manifest line, rather than failing the converter two steps later.

Pulling

export BACKSTAGE_PULL_TOKEN=…       # read_repository on the source projects

backstage-pull pull \
  --sources catalogs/sources.yaml \
  -o catalogs \
  --write-index --model-id service-catalog
Option Default Description
--sources required The manifest
-o, --output required Where to write the descriptors. Gitignore it
--source-map catalog-sources.yaml in --output Where to write the map. Commit it
--write-index off Also write the catalog index, if absent
--index-file catalog-index.yaml in --output Where --write-index writes
--model-id the manifest filename Model id for the generated index
--allow-missing off Treat a repository with no descriptor as expected
--gitlab-url host of the first source The instance to read from
--token $BACKSTAGE_PULL_TOKEN, then $CI_JOB_TOKEN Access token
--anonymous off Send no token; public repositories only

A repository with no descriptor is reported and the run exits non-zero unless --allow-missing is passed. A partial rollout is the ordinary case, but a catalog silently missing a quarter of the estate looks exactly like a complete one, and the graph cannot show the difference — so it is opted into rather than assumed.

--anonymous is stated rather than inferred from an absent token, for the same reason: a pull that silently went anonymous would report every private repository as having no descriptor.

Converting

backstage2linkedarchi convert catalogs \
  --base-iri https://example.org/la/ \
  --diagrams-root catalogs --diagrams-index catalogs/catalog-index.yaml \
  --source-map catalogs/catalog-sources.yaml \
  --format TRIG -o out/backstage.trig

With the map, each descriptor's provenance names its own repository and commit:

<…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml>
    a                prov:Entity ;
    schema:name      "catalog-info.yaml" ;                    # the path upstream
    dct:isPartOf     <https://git.example.org/group/order-service> ;
    dct:identifier   "4f2c1ab8e0d1c2b3a4958677889900aabbccddee" ;
    schema:sha256    "4c294617b60715c1d218e61164a3abd4808a4284cbc30e6728a01ad9aada4481" ;
    prov:alternateOf <https://git.example.org/group/order-service/-/blob/4f2c1ab8e0/catalog-info.yaml> .

<https://git.example.org/group/order-service>
    a           prov:Collection ;
    schema:name "group/order-service" .

<…/element/component/default/order-service>
    prov:wasDerivedFrom      <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
    prov:qualifiedDerivation <…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04> .

<…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04>
    a                prov:Derivation ;
    prov:entity      <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
    prov:hadActivity <…/provenance/run/9c2f1e04> .

Reading that from the outside in:

The source node sits under {base}provenance/, not inside this model's graph. It is shared by every model and every run, so one descriptor at one commit is one node however many models read it — which is what lets an aggregate ask "everything derived from this file" against a single subject. See ADR 0008 §2.

schema:name is the path and dct:isPartOf is the repository. path: in the map lands on schema:name: it is the path inside the source repository, which for a Backstage catalog is catalog-info.yaml in every repository — true, and on its own it identifies nothing. Repository plus path plus commit is the identity, and all three are asserted rather than left to be parsed out of a URL.

prov:alternateOf is the identity statement. The source entity is a minted node, so that things can be said about it — the commit it was read at, who committed it. The blob URL is the same document at the same commit, which is what prov:alternateOf asserts. Emitted only when the map supplies a blobUrl, so nothing is ever claimed to be the same document as a URL that 404s.

No second link sits beside it. The same URL used to be repeated under rdfs:seeAlso, which asserts no relationship between two IRIs and so said nothing that prov:alternateOf had not already said. And no dct:source: this node is the descriptor, so saying it derived from that same blob would make it its own source. dct:source sits on the model, naming the descriptors it was built from.

schema:sha256 is contentSha256 from the map — which bytes, as distinct from which revision. Two commits sharing a digest are a descriptor that moved commit without changing content, which is what an incremental re-pull produces and what a commit sha alone cannot show.

The repository is typed prov:Collection and not also prov:Entity. PROV-O makes Collection a subclass of Entity, so a reasoner sees both, while ?s a prov:Entity — how you enumerate the files a graph was built from — returns only files. The upstream blob is a prov:Entity, though, since that is what prov:alternateOf relates, so enumerate files by the predicate only a file carries:

GRAPH ?g { ?file a prov:Entity ; schema:name ?path ; dct:isPartOf ?repo }

Querying across the graphs

Everything one descriptor produced, in semantic and provenance alike, from a single named subject:

PREFIX prov: <http://www.w3.org/ns/prov#>
SELECT ?graph ?concept ?p ?o WHERE {
  GRAPH ?prov {
    ?concept prov:wasDerivedFrom
      <https://example.org/la/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> .
  }
  GRAPH ?graph { ?concept ?p ?o }
}

Every revision of one descriptor the aggregate holds — this one drops the commit, so it has to match on properties rather than on an IRI, which is what dct:isPartOf and schema:name are for:

PREFIX dct:    <http://purl.org/dc/terms/>
PREFIX prov:   <http://www.w3.org/ns/prov#>
PREFIX schema: <https://schema.org/>
SELECT ?commit ?when WHERE {
  GRAPH ?prov {
    ?src dct:isPartOf <https://git.example.org/group/order-service> ;
         schema:name  "catalog-info.yaml" ;
         dct:identifier ?commit ;
         prov:wasGeneratedBy ?c .
    ?c prov:endedAtTime ?when .
  }
} ORDER BY DESC(?when)

And which run produced a given derivation, which the unqualified form cannot answer once several runs' graphs are merged:

PREFIX prov: <http://www.w3.org/ns/prov#>
PREFIX dct:  <http://purl.org/dc/terms/>
SELECT ?concept ?job WHERE {
  GRAPH ?g {
    ?concept prov:qualifiedDerivation [ prov:entity ?src ; prov:hadActivity ?run ] .
    ?run dct:identifier ?job .
  }
}

An entry in the map wins over --git-provenance for the file it names. A file the map does not mention falls back to git, and the run reports it — a catalog may legitimately mix pulled descriptors with committed ones, and the fallback is correct for the committed half, but a map that has drifted from the manifest would otherwise attribute the very files it exists to describe to the wrong repository.

The source map

backstageCatalogSources: "1"
pulledAt: "2026-08-21T09:12:33Z"
catalogs:
  - file: "orders/order-service.yaml"
    repo: "https://git.example.org/group/order-service"
    ref: "main"
    commit: "4f2c1ab8e0d1c2b3a4958677889900aabbccddee"
    path: "catalog-info.yaml"
    blobUrl: "https://git.example.org/group/order-service/-/blob/4f2c1ab8e0…/catalog-info.yaml"
    contentSha256: "4c294617b60715c1d218e61164a3abd4808a4284cbc30e6728a01ad9aada4481"
    committedAt: "2026-08-19T14:02:11Z"
    authorName: "A. Author"
    authorEmail: "a.author@example.org"

file is where the pull wrote it and is what the converter matches an input against. path is the path inside repo, and it is what schema:name and the blob URL are built from. repo becomes dct:isPartOf and contentSha256 becomes schema:sha256. commit, committedAt, authorName and authorEmail describe the Git commit activity and its person; the trimmed, lower-cased email keys the person IRI, is percent-encoded as one path segment, and is emitted as canonical schema:email. ref records what the pull requested, while the resolved commit records what it read.

commit is the commit that last modified the descriptor, not the head of ref. A ref head moves on every unrelated push to the source repository, so recording it would rewrite this file — and every provenance IRI derived from it — on a pull where no catalog descriptor had changed. Together with sorted entries and a fixed key order, that is what makes an unchanged upstream produce an unchanged file, and therefore a reviewable diff. pulledAt is the one line that always moves, which is why it sits at the top rather than being repeated per entry.

GitLab only

After a successful repository-files response, the pull reads its last_commit_id and requests GET /api/v4/projects/{encoded-project}/repository/commits/{last_commit_id}. From that response it records committed_date, author_name and author_email as the optional committedAt, authorName and authorEmail source-map fields. Metadata is cached per project and commit, so several descriptors at one commit cause one metadata request.

File retrieval remains strict; commit-metadata enrichment is best-effort. A failed metadata request emits a warning and keeps the descriptor plus its repository, path, commit SHA and checksum. An interrupted request still aborts and restores the thread's interrupted status.

The same limit GitProvenance draws, for the same reason: it is the forge in use, and the API shape lives in one place so a second forge is a second client rather than a redesign. One pull reads one instance; every source in a manifest must be on it, which is checked before any request is made — reading group/app from the wrong host either 404s, which would be reported as "no descriptor", or finds a different project of the same name.