Backstage Catalog Converter¶
Converts Backstage catalog YAML (software catalog entities) to RDF aligned to the Backstage ontology.
Backstage reference¶
This converter implements a reading of Backstage's published specification, and every behaviour below traces to one of these pages. Where the two disagree, the specification is right and this converter has a bug.
| Upstream page | What it governs here |
|---|---|
| Software catalog | what a catalog is, and the entity/relation vocabulary the rest assumes |
| System model | the kinds and how they nest — the model the kind → class table mirrors |
| Descriptor format | every documented metadata and spec field, which are required per kind, and the well-known spec.type / spec.lifecycle values. The source for what is read |
| Entity references | the [<kind>:][<namespace>/]<name> form and what each omitted part defaults to — see entity references |
| Well-known relations | the seven relation pairs, and which direction each spec field states. The basis for the relationship table and for normalising spec.children / spec.members |
| Well-known annotations | the annotation keys lifted to named bs: properties, and the prefix rules — see Annotations |
| Extending the model (source) | custom kinds, spec fields, annotations, labels and relation types. In particular adding new fields to the spec object, which is the paragraph custom spec fields rests on — it permits the field, names the two risks, and gives prod-versus-staging as its example |
| Life of an entity | processing and stitching. Why a relation in Backstage comes from a processor rather than from a spec key, which is why an undeclared house field is inert upstream too |
| Creating the catalog graph | how relations compose into a graph — the shape this converter re-expresses in RDF |
| External integrations | populating a catalog from another system, the case backstage-pull and --source-map serve |
| Catalog configuration | catalog.locations and friends, for locating the descriptors you then convert |
| ADR002 — descriptor format | why YAML, and why one file may hold several ----separated entities, which is what multi-document input supports |
Every documented spec field of the seven modelled kinds is converted; the per-kind tables
say what each becomes. Three deliberate departures are documented where they occur: spec.dependsOn is
read more leniently than Backstage reads it, spec.visibility is read
defensively though no published schema declares it, and spec.lifecycle is read on any kind
rather than only where it is required.
Ontology¶
Types align to the published Backstage Metamodel Ontology
(bs: prefix), version 0.4.0.
The seven kinds and their meanings are Backstage's system model; their fields are the descriptor format.
Both columns link out: the kind to the descriptor format that defines it, the class to its term in the published ontology.
| Backstage kind | Ontology type |
|---|---|
| Component | bs:Component |
| System | bs:System |
| API | bs:API |
| Resource | bs:Resource |
| Domain | bs:Domain |
| Group | bs:Group |
| User | bs:User |
| Relationship (from spec) | Ontology type |
|---|---|
spec.owner |
bs:Ownership |
spec.system |
bs:SystemMembership |
spec.providesApis |
bs:APIProvision |
spec.consumesApis |
bs:APIConsumption |
spec.dependsOn / spec.dependencyOf |
bs:Dependency, plus bs:ResourceUsage when the target is a Resource |
spec.domain |
bs:DomainMembership |
spec.memberOf / spec.members |
bs:GroupMembership |
spec.parent / spec.children |
bs:GroupParentage |
spec.subcomponentOf |
bs:ComponentComposition |
spec.subdomainOf |
bs:DomainHierarchy |
backstage.io/techdocs-entity |
bs:TechDocsDelegation |
Each row restates one of the well-known relations, which is also where the direction of each pair is defined.
Three fields are declared on the far end of the relation — spec.children and spec.members on a
Group, and spec.dependencyOf on a Component or Resource — because Backstage materialises every relation
in both directions while the ontology declares one canonical direction per pair. Each is read with source
and target swapped onto the canonical property rather than minting an inverse: a Group's children
become childOf edges from each child, its members memberOf edges from each member, and a
dependencyOf list becomes dependsOn edges from each dependent.
That swap is also what makes one fact stated from either end converge. A spec.dependsOn: [B] and
B spec.dependencyOf: [A] compose the same relationship notation, so they merge onto a single IRI
instead of producing two edges pointing opposite ways.
spec.type
and
spec.lifecycle
are minted as named individuals of the ontology's per-kind vocabularies
(bs:ComponentType,
bs:APIType,
bs:ResourceType,
bs:GroupType,
bs:SystemType,
bs:DomainType,
bs:LifecycleState, and — where present —
bs:ApiVisibility) rather than as plain
string literals. spec.type's well-known values (service, openapi, database, team, …) resolve
to the published individuals; any other token still mints one, typed as the vocabulary's class, since
every one of these is documented as open to organisation-specific values. The pre-0.3.0
bs:lifecycle string property is deprecated and no
longer emitted — bs:lifecycleState replaces
it, and validating against the old one is what bssh:DeprecatedLifecyclePropertyShape is for.
Usage¶
# Single catalog file → TriG
java -jar backstage2linkedarchi.jar convert \
catalog.yaml \
--base-iri https://example.org/la/ \
--model-id my-catalog \
--format TRIG \
-o out.trig
# Entire catalog directory → Turtle
java -jar backstage2linkedarchi.jar convert \
catalog/ \
--base-iri https://example.org/la/ \
--model-id my-catalog \
--format TURTLE \
-o out.ttl
# With direct relationship triples
java -jar backstage2linkedarchi.jar convert \
catalog.yaml \
--base-iri https://example.org/la/ \
--model-id my-catalog \
--format TRIG \
--emit-direct-rel-triples \
-o out.trig
CLI options¶
| Option | Default | Description |
|---|---|---|
<inputs> |
required | Backstage YAML file(s) or directory |
-o, --output |
required | Output artifact, path[:FORMAT[:PROFILE]]. Repeatable — several artifacts are written from one conversion, which is faster than re-running the converter and is what makes them share one prov:generatedAtTime. PROFILE is full (default), no-geometry, no-views or no-diagrams; see output serialization |
--format |
inferred | Default format for any --output that names none: TRIG / TURTLE / JSONLD / RDFXML / NTRIPLES / NQUADS. A file extension such as ttl is also understood |
--base-iri |
required | Base IRI for minting resource IRIs |
--model-id |
filename | Model identifier for a single input. An index id wins over it, and giving it with several inputs is refused — see one model per --model-id. Unlike an index id, the value is not validated |
--type-mapping |
none | YAML type overrides. Also where a house spec field is declared (spec-relations:, with predicates: and optionally relationships:) and where an annotation key's prefix is bound (namespaces:) — see Custom spec fields |
--emit-direct-rel-triples |
false |
Emit {src} bs:ownedBy {tgt} shortcuts, each bridged back to its relationship with rdf:reifies. The qualified predicate (bs:qualifiedOwnedBy and its twelve siblings) is emitted either way |
--emit-skos-labels |
true |
Emit skos:prefLabel |
--label-language |
en |
BCP-47 tag for skos:prefLabel literals |
--emit-skos-notation |
true |
Emit skos:notation |
--diagrams-index |
none | YAML index assigning model IDs and publish status per catalog file. When given, only indexed files are processed — see the index-driven mode below |
--diagrams-root |
index location | Root directory that index file paths are resolved against |
--exclude-states |
none | Leave index entries in these lifecycle states out of the run. Every state is processed by default, so this is the only option that withholds one. Comma-separated, and it cannot name all five. See Lifecycle states |
--include-states |
none | Deprecated and ignored: every state is processed by default. Use --exclude-states |
--include-drafts |
false |
Deprecated and ignored: drafts are processed by default |
--emit-extension-data |
false |
Map the index's elements:, links: and data: entries into the graph — see Extension data |
--ns-global-id |
none | Base IRI for a link target written as a bare name. Requires --emit-extension-data |
--ns-vocab |
none | Base IRI for terms the published ontology has no name for: metadata.annotations keys, and the predicate of a declared custom spec field that predicates: does not name. A prefixed annotation key such as gitlab.com/project-slug needs its prefix declared under namespaces: in --type-mapping instead |
--source-map |
none | Where each input file came from, written by backstage-pull. Needed for a pulled catalog: --git-provenance describes the repository doing the converting, which for a descriptor fetched from elsewhere is not where it came from |
Input format¶
Standard Backstage catalog YAML,
with multi-document support via --- separators — a shape
ADR002 allows explicitly:
---
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
name: payment-api
title: Payment API
description: REST API for payment operations.
tags: [kotlin, spring-boot]
spec:
type: service
lifecycle: production
owner: team-payments
system: payment-gateway
providesApis: [payments-rest-api]
dependsOn: [resource:default/payment-db]
---
apiVersion: backstage.io/v1alpha1
kind: System
metadata:
name: payment-gateway
title: Payment Gateway
spec:
owner: team-payments
domain: payments
Output example¶
.../ in every example on this page
IRIs are shown abbreviated. .../ stands for {--base-iri}backstage/{--model-id}/, so with
--base-iri https://example.org/la/ --model-id catalog the entity Resource:default/payment-db
is written in full as:
https://example.org/la/ backstage/ catalog/ element/ resource/ default/ payment-db
└── --base-iri ────────┘ └ notation ┘ └ model id ┘ └ segment ┘ └ kind ┘ └ namespace ┘ └ name ┘
Three things follow from that shape, and each is covered below:
- the notation slug is always
backstage, which is what keeps a catalog's IRIs from colliding with an ArchiMate or BPMN model converted under the same base - the model id is a namespace, so the same entity in two models is two IRIs — see the multi-repo caveat
- the last three segments are the entity reference, lower-cased, with
:and/becoming path separators — see entity references
The shape is shared by every converter: {base}{notation}/{modelId}/{segment}/{localId}. See
IRI path shape.
@prefix bs: <https://meta.linked.archi/backstage/onto#> .
@prefix arch: <https://meta.linked.archi/core#> .
<.../graph/semantic/group-payments/catalog-info-yaml> {
<.../element/component/default/payment-api> a bs:Component, arch:Element, arch:ModelConcept ;
arch:inModel <.../backstage/service-catalog> ;
skos:prefLabel "Payment API"@en ;
skos:notation "payment-api" ;
skos:definition "REST API for payment operations." ;
bs:entityRef "Component:default/payment-api" ;
bs:kind "Component" ;
bs:namespace "default" ;
bs:name "payment-api" ;
bs:title "Payment API" ;
bs:componentType bs:ServiceType ;
bs:lifecycleState bs:Production .
<.../element/system/default/payment-gateway> a bs:System, arch:Element ;
skos:prefLabel "Payment Gateway"@en ;
bs:entityRef "System:default/payment-gateway" ;
bs:kind "System" ;
bs:namespace "default" ;
bs:name "payment-gateway" .
<.../element/group/default/team-payments> a bs:Group, arch:Element ;
skos:prefLabel "Payments Team"@en ;
bs:entityRef "Group:default/team-payments" ;
bs:kind "Group" ;
bs:namespace "default" ;
bs:name "team-payments" .
<.../relationship/ownedBy--component-default-payment-api--group-default-team-payments>
a bs:Ownership, arch:QualifiedRelationship ;
arch:source <.../element/component/default/payment-api> ;
arch:target <.../element/group/default/team-payments> ;
skos:notation "ownedBy--component-default-payment-api--group-default-team-payments" .
<.../relationship/partOfSystem--component-default-payment-api--system-default-payment-gateway>
a bs:SystemMembership, arch:QualifiedRelationship ;
arch:source <.../element/component/default/payment-api> ;
arch:target <.../element/system/default/payment-gateway> .
# The qualified predicates, pointing from the entity into each relationship. `arch:source`
# only points outward, so without these the resources above — and everything they carry —
# would be reachable only by scanning every `arch:source` in the graph.
<.../element/component/default/payment-api>
bs:qualifiedOwnedBy
<.../relationship/ownedBy--component-default-payment-api--group-default-team-payments> ;
bs:qualifiedPartOfSystem
<.../relationship/partOfSystem--component-default-payment-api--system-default-payment-gateway> .
}
One semantic graph per descriptor¶
A catalog is normally many files, so this converter puts each descriptor's facts in a graph named after that descriptor — its repository path where the run knows one, else its file name:
{base}backstage/service-catalog/graph/semantic/group-payments/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-models/catalog-index-yaml
{base}backstage/service-catalog/graph/model
{base}backstage/service-catalog/graph/provenance
It is what makes "everything this descriptor produced" one GRAPH clause. It is also the only form in
which a fact two descriptors both assert is attributable to each of them: two files declaring the
same edge mint one relationship resource, and asserting prov:wasDerivedFrom twice on that resource
cannot say which file contributed which triple.
graph/semantic/{slug}— one per descriptor.- The index gets one too.
architectureStateand the index'selements:/links:/data:assertions are declared there, so the index is the input they were lifted from. graph/modelholds the curated model: thearch:Model, its folders and their ordering. Its input is the index, and it is described as derived from it.
A catalog converted as a single file keeps the bare graph/semantic. Partitioning one input would
name a graph after the only file there is. There is no flag: the rule follows from how many inputs
contribute to the model.
Each graph is described in graph/provenance as a prov:Bundle derived from its own input. Name the
graph directly when you have its IRI; go through provenance when what you have is the file:
SELECT ?s ?p ?o WHERE {
?g prov:wasDerivedFrom ?src .
?src schema:name "catalog-info.yaml" ; dct:isPartOf <https://git.example.org/group/payments> .
GRAPH ?g { ?s ?p ?o }
}
Pin dct:isPartOf as well as schema:name: every descriptor in every repository is called
catalog-info.yaml, so the path alone matches one graph per repository. Do not compose the graph IRI from
the path — see Aggregating into a knowledge
graph for why it cannot be
inverted.
Give the run a way to tell your descriptors apart
A graph is named after the descriptor's repository path, and without one it falls back to the
file name. Every Backstage descriptor is called catalog-info.yaml, so a catalog pulled from forty
repositories with neither --git-provenance auto nor --source-map puts all forty in one graph.
Nothing false is published when that happens — the graph is described as derived from every file that fed it — but the attribution is no more precise than it was before the split, and the run says so:
Retained identity¶
Every catalog entity carries all four parts of its identity as literals — bs:kind, bs:namespace,
bs:name and the compact bs:entityRef — retained source data alongside skos:prefLabel, per the same
"keep the notation's own attribute, never replace it" contract the other converters follow
(bpmn:name, uml:...). bs:title, bs:uid and bs:etag are emitted only when the source declares
them.
bs:kind is not a duplicate of rdf:type, and the ontology is explicit that neither is derived
from the other: rdf:type is how the graph classifies the entity, a modelling decision an adopter may
revise, while bs:kind is what the descriptor said. bssh:KindTypeAlignmentShape reports divergence at
sh:Info as a reconciliation aid rather than a failure.
bs:entityRef is the form Backstage itself circulates — in the catalog API's relations array, in
spec.owner / spec.system / spec.dependsOn, and in the techdocs-entity annotation — which makes it
the stable cross-source join key. bs:uid is not: the catalog reassigns it when the identical file is
unregistered and re-registered, so it identifies a registration rather than the thing.
bssh:EntityRefConsistencyShape checks the reference against the other three fields.
All four keep the case the source used. Only the IRI is folded to lowercase — see entity references.
One caveat on bs:entityRef
The ontology describes it as read from the source rather than assembled, which is right for a lift
of the catalog API. An authored catalog-info.yaml has no entityRef field, so this converter
composes it from kind, namespace and name — meaning bssh:EntityRefConsistencyShape is
checking our own arithmetic here, and only becomes real evidence for an API-based lift.
Architecture¶
Backstage catalog YAML (single or multi-doc)
→ BackstageParser (SnakeYAML, multi-document)
→ BackstageModel (entities + derived relationships)
→ LinkedArchiEmitter (extends BaseLinkedArchiEmitter)
→ TriG / Turtle output
How relationships are derived¶
The parser extracts relationships from spec fields automatically:
| Field | Relationship emitted |
|---|---|
spec.owner: team-x |
{entity} → bs:Ownership → {Group:default/team-x} |
spec.system: my-system |
{entity} → bs:SystemMembership → {System:default/my-system} |
spec.domain: my-domain |
{entity} → bs:DomainMembership → {Domain:default/my-domain} |
spec.providesApis: [api-1] |
{entity} → bs:APIProvision → {API:default/api-1} |
spec.consumesApis: [api-2] |
{entity} → bs:APIConsumption → {API:default/api-2} |
spec.dependsOn: [resource:default/db] |
{entity} → bs:ResourceUsage → {Resource:default/db} |
spec.memberOf: [team-x] |
{entity} → bs:GroupMembership → {Group:default/team-x} |
Entity references¶
A reference is read as
Backstage defines it —
[<kind>:][<namespace>/]<name> — and resolved to the kind/namespace/name triplet that identifies the
entity, so every way of writing the same reference reaches the same element:
| Written in the catalog | Resolves to | Element IRI |
|---|---|---|
team-payments |
Group:default/team-payments — kind from the field, namespace from the referring entity |
element/group/default/team-payments |
default/payment-gateway |
System:default/payment-gateway — kind from the field |
element/system/default/payment-gateway |
System:payment-gateway |
System:default/payment-gateway — namespace from the referring entity |
element/system/default/payment-gateway |
resource:default/payment-db |
Resource:default/payment-db — kind and namespace as written |
element/resource/default/payment-db |
The resolved reference is the element IRI, written as a path. One function mints and compares, which is what keeps a reference and the entity it names from minting different addresses.
The comparison ignores case, because Backstage's does: resource:default/payment-db and the
Resource:default/payment-db built from the entity's own kind: field are one reference typed two
ways, and the IRI is lowercased so both reach one node. Mixed case is legal in a Backstage name, so
MyService and myservice are one entity too. The case the author wrote survives in bs:kind,
bs:name and bs:entityRef.
The kind is in the id, because it is part of the identity. A Backstage name is unique per kind
within a namespace, so Component:default/payments-service and API:default/payments-service are two
entities — a service and the interface it publishes, named alike, which is the ordinary case rather than
a corner one. An id of {namespace}--{name} gave them one node carrying both bs:Component and
bs:API, both spec.type vocabularies, and every relationship of both, with nothing in the output
saying so.
Three path segments rather than one, because a Backstage name may itself contain -, _ and ., so
component--default--pay--ments cannot be split back into a triplet. Neither : nor / is legal in
any part of a reference, so the path form is reversible and needs no percent-encoding. A consumer that
writes one file per resource gets one directory level per part — see
IRI path shape.
A reference the model does not declare still mints the id that entity would have, so the two sides join if it arrives later, and the run warns:
[WARN] catalog-info.yaml: 'ownedBy' on System:default/payment-gateway names
Group:default/team-payments, which model 'payments-services' does not declare. …
That is the usual signal of a typo, and the usual signal of a catalog split across files: IRIs carry
the model id, so an entity referenced from another file needs the files converted into one model —
a directory input, or one --model-id — rather than merged afterwards.
One reference is read more leniently than Backstage reads it. spec.dependsOn may name a Component or
a Resource and Backstage defaults neither, so it requires the kind; this converter assumes Resource
when it is missing. A bare name meant as a component therefore resolves to a resource that does not
exist and is reported as unresolved.
Custom spec fields¶
A house field in spec — deployedTo, maintainedBy, anything an organisation invents — is read as a
relationship where --type-mapping declares it, and reported rather than dropped where it does
not. Backstage
permits such a field
while advising against it, and real catalogs carry them, so the converter reads one when told what it
means and never guesses.
The documented keys, read without any configuration:
| From | Keys |
|---|---|
metadata |
name, namespace, title, description, tags, uid, etag, annotations, labels, links |
spec, onto the entity |
type, lifecycle, definition, visibility, profile |
spec, as a relationship |
owner, system, domain, subcomponentOf, subdomainOf, parent, providesApis, consumesApis, dependsOn, dependencyOf, memberOf, children, members |
Which of those apply to a given entity is per kind.
Every other spec key is a custom field. Nothing about one is inferred, because nothing about one is
inferable: deployedTo: prod-cluster and tier: gold are the same YAML shape, and only the author
knows that the first names an entity and the second is a value. Saying which is what switches the field
on at all.
Declaring a field¶
Two sections, because the reference/literal distinction is the whole point:
# type-mapping.yml
spec-relations: # the value names another resource
deployedTo: Resource # …and a bare reference in this field defaults to kind Resource
spec-literals: # the value is a value
- tier
predicates:
deployedTo: https://vocab.example.org/arch#deployedTo # the predicate to write
tier: https://vocab.example.org/arch#tier
relationships:
deployedTo: https://vocab.example.org/arch#Deployment # optional: makes it a qualified edge
| Section | Says | Required |
|---|---|---|
spec-relations: |
that the key is read as a reference, and the kind a bare one defaults to | one of the two — an undeclared key is reported and not read |
spec-literals: |
that the key is read as a literal | one of the two |
predicates: |
the predicate IRI, for either kind | yes, unless --ns-vocab supplies a namespace to mint the field name into |
relationships: |
the class of the qualified relationship | no, and it applies to a relation field only — see What is emitted |
spec-relations: is a map because a reference needs a default kind; spec-literals: is a list because a
literal needs nothing per key — not even a datatype, which the catalog already stated. Declaring one key
in both is contradictory, so nothing is emitted for it and the run says so.
The default kind plays exactly the part the descriptor format plays for a documented field: Group for
spec.owner, System for spec.system. It is what lets deployedTo: prod-cluster resolve without the
author writing resource:default/prod-cluster every time.
java -jar backstage2linkedarchi.jar convert catalog.yaml \
--base-iri https://example.org/la/ --model-id catalog \
--type-mapping type-mapping.yml \
-o out.trig
A declared field that a descriptor does not carry produces nothing and says nothing. The declaration describes the catalog's schema, not every entity in it.
How a value is resolved¶
The three rules every extension route uses, plus one at the end that belongs to Backstage:
| Value | Read as | Result |
|---|---|---|
example:prod-cluster |
a prefixed name against namespaces: in --type-mapping |
that IRI — a target outside the catalog |
https://k8s.example.com/clusters/prod |
an absolute IRI | itself, untouched |
prod-cluster, resource:default/prod-cluster |
a Backstage entity reference, with the declared default kind and the referring entity's namespace | the element IRI of that entity |
The third row is the one that mints an address rather than using one. deployedTo: aws-account-012345678901
with deployedTo: Resource declared resolves to Resource:default/aws-account-012345678901, and that
triplet is the path:
{--base-iri}backstage/{--model-id}/element/resource/default/aws-account-012345678901
└ kind ┘ └ ns ─┘ └ name ───────────────┘
Which is the same IRI the Resource entity's own descriptor mints, whether or not this model contains
it — that is what makes the two sides join, and why a reference to an entity the catalog does not declare
still points somewhere and warns rather than being dropped.
Where the target's kind comes from¶
The kind in that path is read from the reference, never looked up from the target. The value wins
where it states one, and spec-relations: supplies it where the value does not:
| Value | Kind used | From |
|---|---|---|
aws-account-012345678901 |
Resource |
the declaration — deployedTo: Resource |
component:default/legacy-host |
Component |
the value, overriding the declaration |
Resource:aws-account-012345678901 |
Resource |
the value; namespace falls back to the referring entity's |
That is the same rule the documented fields follow — spec.owner: team-payments is a Group because the
descriptor format says so, and spec.owner: user:default/ada is a User because the value says so.
A wrong default kind produces a dangling edge, quietly
Nothing searches the catalog for an entity of that name under a different kind. Declaring
deployedTo: Resource while the targets are actually components gives:
an edge to element/resource/default/prod-cluster — a node nothing describes — while
element/component/default/prod-cluster sits in the same file untouched. The run warns:
[WARN] catalog.yaml: 'deployedTo' on Component:default/order-service names
Resource:default/prod-cluster, which model 'catalog' does not declare. The relationship points at
element/resource/default/prod-cluster, a node nothing in this model describes.
The fix is the declaration, or the kind written into the value. This is why the kind is part of the
identity rather than decoration: Component:default/prod-cluster and Resource:default/prod-cluster
are two entities, and a converter that guessed between them would silently merge a cluster with a
service that happened to share a name.
Where a field's targets are genuinely of mixed kinds, leave the kind out of the declaration's reach by writing it in each value; the declared default is only ever a convenience for the uniform case.
A list yields one statement per item, which is the case this exists for:
Two values, two edges. A scalar and a single-item list read identically.
An entity reference the catalog does not declare still mints the IRI that entity would have and warns,
exactly as spec.dependsOn does — see entity references.
A declared prefix wins, including over a kind name
Resolution tries namespaces: first, because an author's own declaration is better evidence than a
guess — the rule the rest of the codebase follows. So declaring a prefix that happens to be a
Backstage kind, resource: or user:, takes that spelling away from the entity-reference reading
in these fields. Nothing refuses it; it is a configuration nobody writes twice, and failing a whole
run over it would be worse than the ambiguity.
--ns-global-id is deliberately not consulted here. It would claim every bare name for an IRI
base and take prod-cluster away from the reading a catalog almost always means.
When the target is not a Backstage entity¶
Everything above is the entity-reference row of how a value is resolved. A prefixed name and an absolute IRI — the other two rows — are addresses already: there is no triplet to complete, so the declared kind is never consulted. That is the route for a target the catalog does not describe, such as a cluster in a Kubernetes API, a record in a CMDB, or a node in another model sharing the graph.
The entry under spec-relations: is still required, because it is what makes the field readable at
all: a key neither section declares is reported and not read. Where a field's values
are all IRIs its kind is inert, so keep it as the kind a bare value would mean — the declaration then
still says something true if one is ever written.
# type-mapping.yml
namespaces:
example: https://vocab.example.org/id/
spec-relations:
hostedAt: Resource # inert while the values are IRIs, and required for the field to be read
predicates:
hostedAt: https://vocab.example.org/arch#hostedAt
spec:
hostedAt: https://k8s.example.com/clusters/prod # itself, untouched
# or example:prod-cluster → https://vocab.example.org/id/prod-cluster
Nothing is asserted about an external target. It is the object of the direct triple, or the
arch:target of the qualified relationship, and that is all — no rdf:type, no label, no folder
membership. Nothing in this model claims to describe it, so there is no absence to report either, unlike
an entity reference the catalog does not declare. Characterising it belongs to whoever publishes its
namespace; the IRIs are chosen so that graph and this output meet.
core-shapes#QualifiedRelationshipShape is satisfied, since what it requires is one arch:source and
one arch:target — not that either end be described here.
A kind the ontology does not model is still a kind. Where the target is a catalog entity but of a kind your organisation added, declare that kind: nothing checks the value against the seven.
spec-relations:
deployedTo: Cluster
elements:
cluster: https://vocab.example.org/arch#Cluster # keyed on the lowercased kind
deployedTo: prod-cluster then mints element/cluster/default/prod-cluster, the same IRI a
kind: Cluster descriptor mints — so the two sides join exactly as they do for a published kind. The
elements: entry is what gives that kind a class on the entity itself; without it, and without
--ns-vocab, entities of that kind are emitted untyped and the run reports it — see
kinds with no ontology class.
What is emitted¶
Two shapes, and which one you get depends on whether relationships: names a class.
With a class — a full arch:QualifiedRelationship, indistinguishable in structure from the edges the
documented fields produce, so the core relationship-endpoint contract applies to it:
<.../relationship/deployedTo--component-default-order-service--resource-default-aws-account-012345678901>
a arch:QualifiedRelationship, arch:ModelConcept, <https://vocab.example.org/arch#Deployment> ;
arch:source <.../element/component/default/order-service> ;
arch:target <.../element/resource/default/aws-account-012345678901> ;
dct:isPartOf <.../folder/Relationships> .
Without one — the direct triple alone:
<.../element/component/default/order-service>
<https://vocab.example.org/arch#deployedTo> <.../element/resource/default/aws-account-012345678901> .
That is not a lesser fallback, it is the honest output. A qualified relationship needs a class, and
this converter will not invent one: bs: is a published document it reads and does not own, so
bs:Deployment would dress a house term up as part of the Backstage ontology with nothing declaring
its meaning — the same rule that keeps bs:gitlabProjectSlug from being minted for an annotation.
One consequence for --emit-direct-rel-triples: it gates the direct triple where a qualified
relationship already carries the statement, since there the triple is only a shortcut. Where there is no
class the direct triple is the statement's only form, so it is always written and the flag does not
apply.
The bssh: shapes will not validate a house class — they know published terms only — but
core-shapes#QualifiedRelationshipShape still checks that both endpoints exist.
A field under spec-literals: becomes a property of the entity, not an edge from it — the same
shape an annotation produces, and emitted beside them:
<.../element/component/default/order-service>
<https://vocab.example.org/arch#tier> "gold" ;
<https://vocab.example.org/arch#replicas> "3"^^xsd:integer .
A list yields one triple per item here too, so a multi-valued field behaves the same whichever kind it is.
The datatype comes from the catalog¶
Nothing declares it and nothing is guessed: the YAML loader has already resolved the scalar by the time the converter sees it, and that resolution is mapped onto XSD.
| Written in the catalog | YAML resolves to | Emitted as |
|---|---|---|
gold, "42", 012345678901 |
String | a plain literal (xsd:string) |
3, -7, 0x1F (→ 31) |
Integer | xsd:integer |
99999999999999999999 |
BigInteger | xsd:integer |
99.95, 1.0, .inf, .nan |
Double | xsd:double |
true, false |
Boolean | xsd:boolean |
2024-01-15, 2024-01-15T10:30:00Z |
Date | xsd:dateTime, normalised to UTC |
empty, null, ~ |
null | nothing — reported as no readable value |
Integer becomes xsd:integer rather than the narrower xsd:int, because the width of the box the
loader chose is not something the catalog said.
YAML 1.1 implicit typing, which the loader still applies
criticalRegion: no does not produce the string "no". YAML reads yes, no, on and off as
booleans, so that field arrives as "false"^^xsd:boolean — the Norway problem, and it is upstream of
this converter rather than something it can fix.
Likewise version: 1.0 is a double and serialises as 1.0, losing any distinction from 1.00.
Quoting is the remedy, and it belongs in the catalog: "no" and "1.0" are strings. Leading
zeros are already safe — 012345678901 stays a string, so an account id keeps its shape.
No OWL axioms are emitted for either term. The predicate from predicates: is used, and the class
from relationships: appears as an rdf:type object, but the graph does not declare
a owl:ObjectProperty, a owl:DatatypeProperty, rdfs:domain or rdfs:range for them, nor a owl:Class
for the class. That holds for the published bs: terms too — a converted graph states facts about a
catalog and leaves the vocabulary to whoever publishes it. If downstream reasoning or validation needs
example:deployedTo characterised, declare it in your own ontology and load it alongside; the IRIs are
chosen so that ontology and this output meet.
What is reported¶
| Situation | Reported | Why |
|---|---|---|
a spec key neither section declares — whether your organisation added it or it is a Template / Location field |
once per run, grouped by key, in one message | a key the converter discarded and one the catalog never carried look identical in the output |
a declared key with no predicate and no --ns-vocab |
once per run | the statement reached nothing, so saying so is the only alternative to losing it quietly |
| a key declared in both sections | once per run | the two declarations contradict each other, and choosing one silently would hide it |
| a declared key with no readable value — empty, or a nested mapping | once per run | toString() on a YAML node yields {region=eu-central-1}, which is a rendering of a node rather than a fact |
| a value that is not a readable reference | per value | usually a typo |
| a declared key a descriptor does not carry | no | the declaration is about the schema |
[WARN] catalog.yaml: 1 custom spec field(s) were not read: costCentre. Backstage allows a house field
in spec but publishes no term for one, so each is read only once declared: add it under
'spec-relations:' in --type-mapping with the kind a bare reference defaults to — e.g.
'deployedTo: Resource'.
deployedTo is not a Backstage field
Backstage's
well-known relations
documents seven relation pairs — ownedBy/ownerOf, providesApi/apiProvidedBy,
consumesApi/apiConsumedBy, dependsOn/dependencyOf, parentOf/childOf,
memberOf/hasMember, partOf/hasPart — and no deployment relation among them. No kind
declares a spec.deployedTo, and no upstream processor writes or reads one. A catalog carrying it
got it from a house processor or a vendor distribution.
Backstage sanctions the extension —
adding new fields to the spec object
states that a kind's schema does not forbid unknown keys and the catalog stores them, and
extending the model
covers custom kinds and relation types alongside — so the field is
legitimate, and this converter reads it. It is simply not shared vocabulary, which is why you
supply the IRIs: there is no bs: term to map it onto, that namespace being a published document
this converter reads and does not own.
Deployment has a published term elsewhere in this toolchain — Structurizr's
structurizr:deployedOn, from a containerInstance — and the nearest thing a catalog has without
any configuration is a spec.dependsOn naming a Resource, which additionally carries
bs:ResourceUsage.
A kind outside the seven is not invented in bs: either. Its class comes from elements: in
--type-mapping (keyed on the lowercased kind) or from --ns-vocab; failing both, the entity carries no
notation class and the run reports it. It is still an arch:Element with its labels, identity fields and
every relation its spec states — see kinds with no ontology class.
Three routes carry a custom statement into the graph:
| Written in | Emits | Requires | |
|---|---|---|---|
a declared spec field |
the catalog file itself | a link or a typed literal | spec-relations: or spec-literals:, and a predicate |
metadata.annotations |
the catalog file itself | a plain literal only | --ns-vocab, or namespaces: in --type-mapping |
| the diagram index | a file beside the catalog | a literal or a link | --emit-extension-data and --diagrams-index |
Pick on where the statement is authored, not on what it is — all three routes now carry either kind. A
declared spec field keeps the value in the descriptor beside everything else about the component, and
is the only route that preserves a datatype: an annotation value is a string even when it reads 3. The
index keeps the statement out of a catalog you do not own.
Recommended practice¶
For a catalog you author, prefer metadata.annotations under a domain-prefixed key over a new spec
field. The support described above exists because catalogs that already carry house fields have to be
convertible, not as an encouragement to add more.
This follows Backstage's own guidance.
Adding new fields to the spec object of an existing kind
records two objections: the field risks colliding with a later addition to the core model, and the data
ordinarily belongs in labels or annotations, or in a new type value. The example intent that section
gives is a field stating whether a component runs in prod or staging — deployedTo in all but name.
The decisive argument is not advisory. A custom spec field is inert in Backstage as well. Relations
there are emitted by processors, not derived from spec keys, so spec.deployedTo produces no relation,
no graph edge and no view in a stock instance: the catalog stores the key and returns it unchanged. It
acquires meaning only through a custom processor the organisation then maintains. Annotations are the
documented channel for metadata consumed by plugins and for links into external systems, so they are
read by convention rather than only by bespoke code. This converter matches upstream behaviour in both
cases.
Four rules follow.
Require a domain prefix — example.org/deployed-to, not deployedTo. Backstage reserves the
backstage.io prefix, expects a key concerning a third-party system to carry a domain that plausibly
owns it, and treats an unprefixed key as local to one instance and unfit to travel outside the
organisation. The two forms also take different routes here: an unprefixed key requires --ns-vocab,
a prefixed key resolves through namespaces: in --type-mapping. The prefixed form is preferable on
both counts, since it places the namespace decision in a reviewed configuration file rather than on a
command line, where two teams could mint one predicate for two meanings.
Either annotations or labels reaches the graph, so choose on meaning. Both are read, and both become one predicate per key by the same rules — see Labels. Backstage's own split is the one to follow: annotations for metadata a plugin consumes or a link into an external system, labels for something classifying that you expect to filter or select on. A deployment target is arguably the latter.
Prefer a published relation where the statement is an edge. An annotation value is always a
literal, so deployedTo arrives as the string "prod-cluster", not as a link to the Resource of that
name, and nothing traverses it. Where a query must walk component → cluster, spec.dependsOn naming
the resource states the edge with published terms — bs:Dependency and bs:ResourceUsage — and keeps
catalog-info.yaml inside the documented schema. A declared spec field is the
option when the relation genuinely is not a dependency and the distinction matters to your queries;
the cost is a term only your organisation understands.
Model one entity per thing, not one per environment. Backstage advises against separate entities
per environment, preferring a single canonical entity whose variation across environments is presented
by a plugin. Treating deployment as an attribute or a dependency edge is consistent with that;
minting order-service-prod and order-service-staging is not.
The diagram index remains the route where a custom predicate must be an edge and the catalog cannot be changed — a vendor-managed descriptor, for instance. For a catalog you do author it is the weaker of the two link-capable routes, since the value is then stated in a second file with nothing checking it against the descriptor.
Labels¶
metadata.labels is read with no configuration, and each key becomes its own predicate — the same two
routes an annotation with no published term takes:
metadata:
name: order-service
labels:
deployments.example.net/register-srv: "true" # prefix resolved via namespaces:
tier: gold # bare key minted under --ns-vocab
<.../element/component/default/order-service>
<https://vocab.example.org/k8s#register-srv> "true" ;
<https://vocab.example.org/arch#tier> "gold" .
bs:label is not the predicate. The ontology publishes
bs:label as an abstract super-property for an
entity's labels, which describes the shape of the vocabulary rather than naming something to write: a
key/value pair on a super-property loses the key. Per-key predicates are what that super-property is the
parent of, so that is what is emitted.
A label value is always a plain string. Backstage borrows
Kubernetes label semantics
and says outright that both key and value are strings, so unlike a spec literal
nothing is datatype-inferred — "true" stays a string.
The prefix rules are the annotation rules: backstage.io/ is reserved, a key concerning a third-party
system should carry the domain that owns it, and a bare key is local to one Backstage instance. A key whose
local part cannot be an IRI local name is reported rather than mangled.
House metadata keys¶
The metadata object is
open for extension
just as spec is, with the same caution attached. A key outside the documented set is declared under
metadata-literals:, which is deliberately separate from spec-literals: — metadata.tier and
spec.tier are two statements, and one declaration must not answer for both:
# type-mapping.yml
metadata-literals:
- costCentre
predicates:
costCentre: https://vocab.example.org/arch#costCentre
Datatypes come from YAML exactly as for a spec literal, so replicas: 3 arrives as xsd:integer. There
is no relation counterpart: metadata is where Backstage puts descriptive fields, and a reference belongs
in spec or in a label. An undeclared key is reported, with the reminder that a key/value classifier
usually wants metadata.labels, which needs no declaration at all.
Annotations¶
metadata.annotations is Backstage's own open extension point, and unlike an undeclared spec key it
is read with no configuration. The keys with defined meanings are the
well-known annotations, and
those are lifted to named bs: properties. A key the published ontology has no term for needs a namespace
someone owns.
A domain-prefixed key, the recommended form, resolves through namespaces: in --type-mapping:
# catalog-info.yaml
metadata:
name: order-service
annotations:
example.org/deployed-to: prod-cluster
java -jar backstage2linkedarchi.jar convert catalog.yaml \
--base-iri https://example.org/la/ --model-id catalog \
--type-mapping type-mapping.yml \
-o out.trig
<.../element/component/default/order-service>
<https://vocab.example.org/arch#deployed-to> "prod-cluster" .
An unprefixed key is minted onto --ns-vocab instead: deployedTo with
--ns-vocab https://vocab.example.org/arch# gives <https://vocab.example.org/arch#deployedTo>.
Backstage treats such a key as local to one instance, so the prefixed form is preferable for anything
that leaves the organisation.
A prefixed key is deliberately not minted onto --ns-vocab with its prefix dropped, because
github.com/project-slug and gitlab.com/project-slug would then land on one predicate holding two
values that mean different things.
The value is always a literal. namesResource never treats an annotation value as a reference, so
"prod-cluster" is a string and not a link to the Resource entity of that name — a query cannot
traverse it. That is the ceiling of this route, and the reason the index one exists.
Keys that reach no predicate are reported once per run, grouped by key rather than by entity:
[WARN] catalog.yaml: 2 annotation(s) have no published bs: term and reached no predicate:
costCentre, deployedTo. Pass --ns-vocab <your namespace> to publish them under a namespace you own.
Converting a catalog API dump¶
Two root fields exist only in what the catalog API returns, never in an authored descriptor: Backstage
states that a descriptor file
is not supposed to contain relations,
and status is written by the catalog's own ingestion. Both are read, so a lift of a dump keeps what the
dump knows and an authored file is unaffected.
relations — the authoritative array¶
Backstage calls produced relations the authoritative source, and an array carries edges a custom processor
derived that no spec field states. It also carries both directions of every relation, so one
ownership appears as ownedBy on the component and ownerOf on the group — and again in spec.owner.
Every entry is normalised onto the ontology's single canonical direction, so all three converge on one edge rather than three:
| Relation type in the array | Emitted as |
|---|---|
ownedBy / ownerOf |
bs:Ownership |
providesApi / apiProvidedBy |
bs:APIProvision |
consumesApi / apiConsumedBy |
bs:APIConsumption |
dependsOn / dependencyOf |
bs:Dependency |
childOf / parentOf |
bs:GroupParentage |
memberOf / hasMember |
bs:GroupMembership |
partOf / hasPart |
depends on the endpoints — see below |
| anything else | no published class; declarable via relationships: and predicates: |
In each pair the second name is the reverse, and its endpoints are swapped rather than emitted as written.
That is the same treatment spec.children, spec.members and spec.dependencyOf get, and it is what makes
the merge work: identical type and endpoints compose one relationship notation, hence one IRI.
partOf is the one type that cannot be mapped by name. Backstage collapses four containments onto it
while the ontology publishes a class for each, so the kinds at the ends decide — and a targetRef always
states its kind:
| From → to | Emitted as |
|---|---|
Component, API or Resource → System |
bs:SystemMembership |
Component → Component |
bs:ComponentComposition |
System → Domain |
bs:DomainMembership |
Domain → Domain |
bs:DomainHierarchy |
| any other pair | reported, not emitted |
status — the catalog's own report¶
status.items says whether the entity was ingested cleanly. The ontology publishes nothing for it, so
nothing is minted in bs:; with --ns-vocab the items become one node each, the shape metadata.links
already uses, because a level and a message flattened onto the entity lose which belongs to which:
<.../element/component/default/order-service>
<https://vocab.example.org/arch#hasStatusItem> [
<https://vocab.example.org/arch#statusType> "backstage.io/catalog-processing" ;
<https://vocab.example.org/arch#statusLevel> "error" ;
<https://vocab.example.org/arch#statusMessage> "NotFoundError: File not found"
] .
Without --ns-vocab the items are reported and dropped. The nested error object is not read: its shape
is not fixed upstream and a stack trace is not an architectural fact — the message says the same thing
in the form a person reads.
That the ontology has no status vocabulary, and no class for Template or Location, is filed upstream
in todo/UPSTREAM-REQUEST-backstage-kinds-and-status.md.
Fields per kind¶
What the descriptor format
declares for each kind, and what this converter does with it. required and optional are Backstage's
own markers, not this converter's: nothing here is enforced at conversion time, because a missing
required field is the catalog's validation to fail, not a lift's. The
shapes are where per-kind expectations are checked in this pipeline.
Every kind additionally accepts the metadata block — name, namespace, title, description,
tags, annotations, links — all of which are read for every kind.
Component → bs:Component¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.type |
required | bs:componentType → a bs:ComponentType individual |
spec.lifecycle |
required | bs:lifecycleState individual |
spec.owner |
required | bs:Ownership → Group |
spec.system |
optional | bs:SystemMembership → System |
spec.subcomponentOf |
optional | bs:ComponentComposition → Component |
spec.providesApis |
optional | bs:APIProvision → API, one per item |
spec.consumesApis |
optional | bs:APIConsumption → API, one per item |
spec.dependsOn |
optional | bs:Dependency, plus bs:ResourceUsage for a Resource target |
spec.dependencyOf |
optional | the same bs:Dependency, direction swapped: each listed entity → this one |
API → bs:API¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.type |
required | bs:apiType → an bs:APIType individual — openapi, asyncapi, graphql, grpc known |
spec.lifecycle |
required | bs:lifecycleState individual |
spec.owner |
required | bs:Ownership → Group |
spec.definition |
required | bs:definition literal — the interface document itself, copied verbatim |
spec.system |
optional | bs:SystemMembership → System |
spec.visibility |
not in the schema | bs:apiVisibility → a bs:ApiVisibility individual, read defensively — see the note below |
An API is the one kind with a required field carrying a document rather than a reference. spec.definition
is emitted as a literal exactly as written, so an OpenAPI document arrives in the graph as a string; the
converter does not parse it.
spec.visibility is read though no published schema declares it
The descriptor format has no visibility field on an API today, while the system model's prose does
talk about an API having a visibility, and the ontology already publishes
bs:ApiVisibility with public,
restricted and private. It is therefore read if present, so an organisation whose own tooling
populates it loses nothing, and absent otherwise. This is the one field this converter reads that
Backstage does not document.
Resource → bs:Resource¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.type |
required | bs:resourceType → a bs:ResourceType individual — database, s3-bucket, kubernetes-cluster known |
spec.owner |
required | bs:Ownership → Group |
spec.system |
optional | bs:SystemMembership → System |
spec.dependsOn |
optional | bs:Dependency (+ bs:ResourceUsage for a Resource target) |
spec.dependencyOf |
optional | the same bs:Dependency, direction swapped: each listed entity → this one. Since the target is then this Resource, the edge also carries bs:ResourceUsage |
A Resource has no lifecycle in the descriptor format. This converter reads one if present anyway —
see the liberties below.
System → bs:System¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.owner |
required | bs:Ownership → Group |
spec.domain |
optional | bs:DomainMembership → Domain |
spec.type |
optional | bs:systemType → a bs:SystemType individual — product, service, feature-set known |
Domain → bs:Domain¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.owner |
required | bs:Ownership → Group |
spec.subdomainOf |
optional | bs:DomainHierarchy → Domain |
spec.type |
optional | bs:domainType → a bs:DomainType individual |
Group → bs:Group¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.type |
required | bs:groupType → a bs:GroupType individual — team, business-unit, product-area, root known |
spec.children |
required | bs:GroupParentage, one per item, direction swapped: each child → this group |
spec.parent |
optional | bs:GroupParentage: this group → parent |
spec.members |
optional | bs:GroupMembership, one per item, direction swapped: each member → this group |
spec.profile |
optional | bs:displayName, bs:email, bs:picture literals, flattened onto the entity |
spec.children is required and yet lists the other end of the relation, which is why it and
spec.members are normalised onto the canonical child→parent and member→group direction rather than
minting an inverse property. A Group's own spec.parent already runs that way and is read as written.
User → bs:User¶
| Field | Backstage | Emitted as |
|---|---|---|
spec.memberOf |
required | bs:GroupMembership → Group, one per item |
spec.profile |
optional | bs:displayName, bs:email, bs:picture literals |
A User has no spec.type and no spec.owner.
Kinds with no ontology class¶
Template and Location are documented kinds the published ontology does not model, because a scaffolder form and a catalog location are not architecture. A house kind is in the same position.
An entity of such a kind still converts, and carries no bs: class:
<.../element/location/default/all-components>
a arch:Element, arch:ModelConcept ; # and nothing from bs:
skos:prefLabel "all-components"@en ;
bs:kind "Location" ; # what the descriptor said, retained
bs:entityRef "Location:default/all-components" .
bs:Location is not minted, for the reason nothing else is: that namespace is a published document this
converter reads and does not own. A class comes from elements: in --type-mapping, keyed on the
lowercased kind, or from --ns-vocab:
With --ns-vocab https://vocab.example.org/arch# and nothing else, the kind mints
<https://vocab.example.org/arch#Location> under the namespace you own. With neither, the entity keeps
its core types and the run says what happened:
[WARN] catalog.yaml: 1 entity kind(s) carry no notation class: Location. The ontology publishes classes
for Component, System, API, Resource, Domain, Group and User; Template and Location are documented
Backstage kinds it deliberately does not model … These entities are still emitted as arch:Element with
their labels, identity and relations — to type them, name a class under 'elements:' in --type-mapping
(keyed on the lowercased kind) or pass --ns-vocab <your namespace>.
The kind-shared fields (spec.type, spec.owner) are read as for any entity. The kind-specific ones are
not, and are declarable if you want them: spec.parameters and spec.steps describe
a scaffolder form, spec.target, spec.targets and spec.presence where a catalog reads descriptors
from.
Documented fields with no ontology term¶
Every documented field of the seven modelled kinds is converted. Five are not, and all five belong to
Template and Location, the kinds the ontology does not model:
| Field | Kind | What it is |
|---|---|---|
spec.parameters, spec.steps |
Template | a scaffolder form — the shape of a wizard, not an architectural fact |
spec.target, spec.targets, spec.presence |
Location | where a catalog reads descriptors from, not a relation between entities |
They are unmapped by default, not blocked. There is no bs: term for them and this converter will not
mint one, but the declaration mechanism is exactly the escape hatch: list the field
under spec-literals: with a predicate from a namespace you own and it reaches the graph like any other
value.
The converter does not hold a list of these five. They reach the same report as a field your organisation invented, because they are unread for the same reason — no published term names them — and published by the same remedy. A hardcoded copy of another project's schema, existing only to reword a warning, would go stale the moment Backstage added a field to either kind, and would go stale silently:
[WARN] catalog.yaml: 2 spec field(s) reached no predicate and were not read: costCentre, presence. This
converter reads the documented fields of the seven kinds the ontology models; anything else — a field
your organisation added, or a field of the Template and Location kinds — has no published term, and this
converter does not invent one. To publish it, declare the key in --type-mapping: 'spec-relations:' with
the kind a bare reference defaults to (e.g. 'deployedTo: Resource') if its value names another entity, or
'spec-literals:' if its value is a value, together with a predicate under 'predicates:' or a namespace
via --ns-vocab.
So the table above is documentation, where staleness is visible and harmless, rather than code, where it would not be.
Two liberties with the per-kind schema¶
Both deliberate:
spec.lifecycleis read on any kind, not only Component and API where it is required. AResourcecarrying one gets it mapped rather than dropped.spec.dependsOnassumesResourcewhen a reference omits its kind. Backstage requires the kind there precisely because the target may be a Component or a Resource and neither is the default — see entity references.
Directory support¶
Pass a directory path to recursively process all .yaml / .yml files:
java -jar backstage2linkedarchi.jar convert \
/path/to/catalog/ \
--base-iri https://example.org/la/ \
--model-id full-catalog \
--format TRIG \
-o out.trig
Extension data¶
A catalog can carry statements the Backstage schema has no field for — which capability a component
realizes, which LeanIX factsheet it corresponds to, what its cost centre is. --emit-extension-data
maps them onto the entities and the model they annotate.
The index is the only route here. BPMN has extensionElements and PlantUML has comments no
renderer touches; a .yaml owned by Backstage has neither, and adding a key to it would be a claim
on a format this converter only reads. See
Extension data for what the routes share.
prefixes:
am: https://meta.linked.archi/archimate3/onto#
arch: https://meta.linked.archi/core#
kg: https://example.org/graph/
x: https://example.org/vocab#
diagrams:
- id: catalog
file: catalog.yaml
links:
arch:architectureState: arch:Baseline # about the model
data:
x:reviewedBy: Jane Doe
elements:
order-service: # about one entity
links:
am:realizes: kg:CAP-OrderManagement
data:
x:costCentre: CC-4711
java -jar backstage2linkedarchi.jar convert catalog.yaml \
--diagrams-index index.yaml --base-iri https://example.org/la/ \
--emit-extension-data -o out.trig
Both flags are needed. --emit-extension-data on its own does nothing here, because the assertions
arrive with the index and there is no in-file route to fall back to. Note also that an index
restricts the run: only the files it lists are converted.
Worked example: a custom deployedTo¶
spec.deployedTo is a house field, so it is read only where --type-mapping declares it — see
Custom spec fields. Where you would rather not configure the catalog's schema at
all, or cannot because the descriptor belongs to someone else, restating the fact in the index turns it
into an edge between the two entities the catalog already declares.
# catalog.yaml
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
name: order-service
spec:
type: service
lifecycle: production
owner: team-orders
deployedTo: prod-cluster # house field, left undeclared here — reported, not read
---
apiVersion: backstage.io/v1alpha1
kind: Resource
metadata:
name: prod-cluster
spec:
type: kubernetes-cluster
owner: team-platform
# index.yaml
prefixes:
x: https://vocab.example.org/arch#
cat: https://example.org/la/backstage/catalog/element/
diagrams:
- id: catalog
file: catalog.yaml
elements:
component:default/order-service:
links:
x:deployedTo: cat:resource/default/prod-cluster
java -jar backstage2linkedarchi.jar convert catalog.yaml \
--diagrams-index index.yaml \
--diagrams-root . \
--base-iri https://example.org/la/ \
--emit-extension-data \
-o out.trig
<https://example.org/la/backstage/catalog/element/component/default/order-service>
<https://vocab.example.org/arch#deployedTo>
<https://example.org/la/backstage/catalog/element/resource/default/prod-cluster> .
Three things about that are worth knowing before writing one.
A subject and a target resolve by different rules. The subject is looked up among the entity's
names — bare name, entity ref, minted id, per Naming an entity. The target is not:
it goes through the same resolver every route uses, which accepts a CURIE against prefixes:, an
absolute IRI, or a bare name with --ns-global-id, and knows nothing about Backstage refs. Writing
x:deployedTo: resource:default/prod-cluster therefore reads resource as an undeclared prefix, and
the value is kept as a literal and reported:
[WARN] catalog.yaml: 'resource:default/prod-cluster' on 'component:default/order-service' is not a
resolvable IRI — prefix 'resource:' is not declared. Kept as a literal, so it is not a link in the
graph.
Hence the cat: prefix above, bound to this model's element namespace,
{--base-iri}backstage/{model-id}/element/. That does couple the index to the base IRI and the model
id, so an index written this way follows a change to either. Pointing at a system outside the catalog
is simpler, since an absolute IRI needs no prefix at all:
The result is a plain triple, not a qualified relationship. No bs: class, no arch:source /
arch:target, no membership of the Relationships folder — unlike the relations derived from spec.
--emit-direct-rel-triples is unrelated and does not affect it. If the edge needs the qualified form
with attributes of its own, the statement belongs in a graph the aggregating repository builds, not
here.
A term that runs the other way needs no inverse. To say it from the cluster's side while reusing
one predicate, use direction:
elements:
component:default/order-service:
links:
x:hosts: { target: cat:resource/default/prod-cluster, direction: Backward }
That emits {cluster} x:hosts {component}. Both writes the pair. A literal target cannot be the
subject of an inverse, so Backward on one is reported and only the forward triple is written.
For a value that really is just a string, data: puts it on the entity without pretending it is a
reference:
Naming an entity¶
A Backstage entity has several usable names, and any of them resolves:
kind: Component, namespace: default, name: order-service |
|
|---|---|
| entity name | order-service |
| entity ref | Component:default/order-service, or the lowercase form Backstage prints |
| minted id | component/default/order-service |
The ref works in both casings because Backstage's docs write it lowercase while the catalog file
spells kind capitalised — the same reference typed from two places. A bare name shared by two
entities is refused rather than guessed at, and reported; use the ref, which is unique. An annotation
naming an entity the catalog does not hold is reported too, since that is what a rename leaves behind.
An entity ref needs no quoting despite its colon — see colons and quoting for the one case that does:
elements:
component:default/order-service: # fine unquoted
links:
am:realizes: kg:CAP-OrderManagement
About the model¶
links: and data: written directly on the entry describe the model. This converter emits no
arch:View — a catalog is not a diagram — so a view-level block from the model/views schema lands on
the model as well, the same collapse the flat schema makes for every entry-level field.
| Option | Default | Description |
|---|---|---|
--emit-extension-data |
false |
Map the index's elements:, links: and data: entries into the graph |
--ns-global-id |
none | Base IRI for a link target written as a bare name. Requires --emit-extension-data |
Architecture state¶
A catalog can be indexed as describing the current architecture or a planned one:
diagrams:
- id: full-catalog
file: catalog.yaml
architectureState: baseline # or target, transitional
A checked value, so a typo fails the run rather than reaching the graph. This converter declares no
arch:View, so the state lands on the model whichever level of the index declares it — the same
collapse the flat schema makes for every entry-level field.
See Architecture state, including why it is
orthogonal to status:.
Validate¶
Runs SHACL validation on converted output. Shapes and ontologies are fetched from
meta.linked.archi at runtime.
Default shapes are backstage-shapes and core-shapes. The backstage and core ontologies are
both loaded for rdfs:subClassOf reasoning.
bssh:BackstageElementLabelShape is what gives this converter's output a naming rule at all, since
the core label shapes were withdrawn, and the per-kind shapes check the relations a catalog derives
from its spec fields — ownership, system membership, API provision and consumption, resource use,
domain and group membership. Switch the naming rule off for a run with --without-shape labels.
Two shapes cover the identity fields:
| Shape | Severity | Checks |
|---|---|---|
bssh:EntityRefConsistencyShape |
Violation | bs:entityRef agrees with bs:kind, bs:namespace and bs:name, compared case-insensitively. They come from one descriptor, so a mismatch means the lift mangled one of them |
bssh:KindTypeAlignmentShape |
Info | bs:kind matches the skos:notation of some bs: class the entity is typed with. Reported, not failed: retyping a lifted entity under a later convention is legitimate. Switch it off where retyping is standing policy |
A multi-repo run reports its cross-repo references
A single catalog conforms. A run spanning repositories does not, and the shapes are right to
say so. spec.owner: team-payments in one repo mints
{model}/element/group/default/team-payments inside that model's namespace, while the Group
entity declared in another repo is {other-model}/element/group/default/team-payments — two IRIs
for one team. bssh:OwnershipShape reports a target that is not a bs:Group, and merging the
graphs does not reconcile them, because neither IRI is the other.
Entity references resolves the form a reference is written in, within one model. This is the different question of which model owns an entity that a repo references without declaring — and ADR 0001 places cross-repository uniqueness with the aggregating repository, so the converters do not decide it.
It is not a new failure. Those references already failed
core-shapes#QualifiedRelationshipShape, which was in the default set before the Backstage
shapes were: 6 violations before, 12 after, all one root cause. What changed is that the
report now names the problem (Ownership target must be a Group) instead of stating it
generically.
java -jar backstage2linkedarchi.jar validate -i out.trig
# Validate against your own shapes instead
java -jar backstage2linkedarchi.jar validate -i out.trig --shapes ./shapes/catalog-rules.ttl
# Offline, reusing documents fetched by an earlier run
java -jar backstage2linkedarchi.jar validate -i out.trig --asset-dir .assets --offline
Exits 0 when the data conforms, 1 when violations are found, 2 on error.
See Validation for the shared engine, all options, coverage reporting and report format.