Skip to content

Backstage Catalog Converter

Converts Backstage catalog YAML (software catalog entities) to RDF aligned to the Backstage ontology.

Backstage reference

This converter implements a reading of Backstage's published specification, and every behaviour below traces to one of these pages. Where the two disagree, the specification is right and this converter has a bug.

Upstream page What it governs here
Software catalog what a catalog is, and the entity/relation vocabulary the rest assumes
System model the kinds and how they nest — the model the kind → class table mirrors
Descriptor format every documented metadata and spec field, which are required per kind, and the well-known spec.type / spec.lifecycle values. The source for what is read
Entity references the [<kind>:][<namespace>/]<name> form and what each omitted part defaults to — see entity references
Well-known relations the seven relation pairs, and which direction each spec field states. The basis for the relationship table and for normalising spec.children / spec.members
Well-known annotations the annotation keys lifted to named bs: properties, and the prefix rules — see Annotations
Extending the model (source) custom kinds, spec fields, annotations, labels and relation types. In particular adding new fields to the spec object, which is the paragraph custom spec fields rests on — it permits the field, names the two risks, and gives prod-versus-staging as its example
Life of an entity processing and stitching. Why a relation in Backstage comes from a processor rather than from a spec key, which is why an undeclared house field is inert upstream too
Creating the catalog graph how relations compose into a graph — the shape this converter re-expresses in RDF
External integrations populating a catalog from another system, the case backstage-pull and --source-map serve
Catalog configuration catalog.locations and friends, for locating the descriptors you then convert
ADR002 — descriptor format why YAML, and why one file may hold several ----separated entities, which is what multi-document input supports

Every documented spec field of the seven modelled kinds is converted; the per-kind tables say what each becomes. Three deliberate departures are documented where they occur: spec.dependsOn is read more leniently than Backstage reads it, spec.visibility is read defensively though no published schema declares it, and spec.lifecycle is read on any kind rather than only where it is required.

Ontology

Types align to the published Backstage Metamodel Ontology (bs: prefix), version 0.4.0.

The seven kinds and their meanings are Backstage's system model; their fields are the descriptor format.

Both columns link out: the kind to the descriptor format that defines it, the class to its term in the published ontology.

Backstage kind Ontology type
Component bs:Component
System bs:System
API bs:API
Resource bs:Resource
Domain bs:Domain
Group bs:Group
User bs:User
Relationship (from spec) Ontology type
spec.owner bs:Ownership
spec.system bs:SystemMembership
spec.providesApis bs:APIProvision
spec.consumesApis bs:APIConsumption
spec.dependsOn / spec.dependencyOf bs:Dependency, plus bs:ResourceUsage when the target is a Resource
spec.domain bs:DomainMembership
spec.memberOf / spec.members bs:GroupMembership
spec.parent / spec.children bs:GroupParentage
spec.subcomponentOf bs:ComponentComposition
spec.subdomainOf bs:DomainHierarchy
backstage.io/techdocs-entity bs:TechDocsDelegation

Each row restates one of the well-known relations, which is also where the direction of each pair is defined.

Three fields are declared on the far end of the relation — spec.children and spec.members on a Group, and spec.dependencyOf on a Component or Resource — because Backstage materialises every relation in both directions while the ontology declares one canonical direction per pair. Each is read with source and target swapped onto the canonical property rather than minting an inverse: a Group's children become childOf edges from each child, its members memberOf edges from each member, and a dependencyOf list becomes dependsOn edges from each dependent.

That swap is also what makes one fact stated from either end converge. A spec.dependsOn: [B] and B spec.dependencyOf: [A] compose the same relationship notation, so they merge onto a single IRI instead of producing two edges pointing opposite ways.

spec.type and spec.lifecycle are minted as named individuals of the ontology's per-kind vocabularies (bs:ComponentType, bs:APIType, bs:ResourceType, bs:GroupType, bs:SystemType, bs:DomainType, bs:LifecycleState, and — where present — bs:ApiVisibility) rather than as plain string literals. spec.type's well-known values (service, openapi, database, team, …) resolve to the published individuals; any other token still mints one, typed as the vocabulary's class, since every one of these is documented as open to organisation-specific values. The pre-0.3.0 bs:lifecycle string property is deprecated and no longer emitted — bs:lifecycleState replaces it, and validating against the old one is what bssh:DeprecatedLifecyclePropertyShape is for.

Usage

# Single catalog file → TriG
java -jar backstage2linkedarchi.jar convert \
  catalog.yaml \
  --base-iri https://example.org/la/ \
  --model-id my-catalog \
  --format TRIG \
  -o out.trig

# Entire catalog directory → Turtle
java -jar backstage2linkedarchi.jar convert \
  catalog/ \
  --base-iri https://example.org/la/ \
  --model-id my-catalog \
  --format TURTLE \
  -o out.ttl

# With direct relationship triples
java -jar backstage2linkedarchi.jar convert \
  catalog.yaml \
  --base-iri https://example.org/la/ \
  --model-id my-catalog \
  --format TRIG \
  --emit-direct-rel-triples \
  -o out.trig

CLI options

Option Default Description
<inputs> required Backstage YAML file(s) or directory
-o, --output required Output artifact, path[:FORMAT[:PROFILE]]. Repeatable — several artifacts are written from one conversion, which is faster than re-running the converter and is what makes them share one prov:generatedAtTime. PROFILE is full (default), no-geometry, no-views or no-diagrams; see output serialization
--format inferred Default format for any --output that names none: TRIG / TURTLE / JSONLD / RDFXML / NTRIPLES / NQUADS. A file extension such as ttl is also understood
--base-iri required Base IRI for minting resource IRIs
--model-id filename Model identifier for a single input. An index id wins over it, and giving it with several inputs is refused — see one model per --model-id. Unlike an index id, the value is not validated
--type-mapping none YAML type overrides. Also where a house spec field is declared (spec-relations:, with predicates: and optionally relationships:) and where an annotation key's prefix is bound (namespaces:) — see Custom spec fields
--emit-direct-rel-triples false Emit {src} bs:ownedBy {tgt} shortcuts, each bridged back to its relationship with rdf:reifies. The qualified predicate (bs:qualifiedOwnedBy and its twelve siblings) is emitted either way
--emit-skos-labels true Emit skos:prefLabel
--label-language en BCP-47 tag for skos:prefLabel literals
--emit-skos-notation true Emit skos:notation
--diagrams-index none YAML index assigning model IDs and publish status per catalog file. When given, only indexed files are processed — see the index-driven mode below
--diagrams-root index location Root directory that index file paths are resolved against
--exclude-states none Leave index entries in these lifecycle states out of the run. Every state is processed by default, so this is the only option that withholds one. Comma-separated, and it cannot name all five. See Lifecycle states
--include-states none Deprecated and ignored: every state is processed by default. Use --exclude-states
--include-drafts false Deprecated and ignored: drafts are processed by default
--emit-extension-data false Map the index's elements:, links: and data: entries into the graph — see Extension data
--ns-global-id none Base IRI for a link target written as a bare name. Requires --emit-extension-data
--ns-vocab none Base IRI for terms the published ontology has no name for: metadata.annotations keys, and the predicate of a declared custom spec field that predicates: does not name. A prefixed annotation key such as gitlab.com/project-slug needs its prefix declared under namespaces: in --type-mapping instead
--source-map none Where each input file came from, written by backstage-pull. Needed for a pulled catalog: --git-provenance describes the repository doing the converting, which for a descriptor fetched from elsewhere is not where it came from

Input format

Standard Backstage catalog YAML, with multi-document support via --- separators — a shape ADR002 allows explicitly:

---
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
  name: payment-api
  title: Payment API
  description: REST API for payment operations.
  tags: [kotlin, spring-boot]
spec:
  type: service
  lifecycle: production
  owner: team-payments
  system: payment-gateway
  providesApis: [payments-rest-api]
  dependsOn: [resource:default/payment-db]
---
apiVersion: backstage.io/v1alpha1
kind: System
metadata:
  name: payment-gateway
  title: Payment Gateway
spec:
  owner: team-payments
  domain: payments

Output example

.../ in every example on this page

IRIs are shown abbreviated. .../ stands for {--base-iri}backstage/{--model-id}/, so with --base-iri https://example.org/la/ --model-id catalog the entity Resource:default/payment-db is written in full as:

https://example.org/la/     backstage/   catalog/     element/   resource/  default/     payment-db
└── --base-iri ────────┘    └ notation ┘ └ model id ┘ └ segment ┘ └ kind ┘  └ namespace ┘ └ name ┘

Three things follow from that shape, and each is covered below:

  • the notation slug is always backstage, which is what keeps a catalog's IRIs from colliding with an ArchiMate or BPMN model converted under the same base
  • the model id is a namespace, so the same entity in two models is two IRIs — see the multi-repo caveat
  • the last three segments are the entity reference, lower-cased, with : and / becoming path separators — see entity references

The shape is shared by every converter: {base}{notation}/{modelId}/{segment}/{localId}. See IRI path shape.

@prefix bs: <https://meta.linked.archi/backstage/onto#> .
@prefix arch: <https://meta.linked.archi/core#> .

<.../graph/semantic/group-payments/catalog-info-yaml> {
    <.../element/component/default/payment-api> a bs:Component, arch:Element, arch:ModelConcept ;
        arch:inModel <.../backstage/service-catalog> ;
        skos:prefLabel "Payment API"@en ;
        skos:notation "payment-api" ;
        skos:definition "REST API for payment operations." ;
        bs:entityRef "Component:default/payment-api" ;
        bs:kind "Component" ;
        bs:namespace "default" ;
        bs:name "payment-api" ;
        bs:title "Payment API" ;
        bs:componentType bs:ServiceType ;
        bs:lifecycleState bs:Production .

    <.../element/system/default/payment-gateway> a bs:System, arch:Element ;
        skos:prefLabel "Payment Gateway"@en ;
        bs:entityRef "System:default/payment-gateway" ;
        bs:kind "System" ;
        bs:namespace "default" ;
        bs:name "payment-gateway" .

    <.../element/group/default/team-payments> a bs:Group, arch:Element ;
        skos:prefLabel "Payments Team"@en ;
        bs:entityRef "Group:default/team-payments" ;
        bs:kind "Group" ;
        bs:namespace "default" ;
        bs:name "team-payments" .

    <.../relationship/ownedBy--component-default-payment-api--group-default-team-payments>
        a bs:Ownership, arch:QualifiedRelationship ;
        arch:source <.../element/component/default/payment-api> ;
        arch:target <.../element/group/default/team-payments> ;
        skos:notation "ownedBy--component-default-payment-api--group-default-team-payments" .

    <.../relationship/partOfSystem--component-default-payment-api--system-default-payment-gateway>
        a bs:SystemMembership, arch:QualifiedRelationship ;
        arch:source <.../element/component/default/payment-api> ;
        arch:target <.../element/system/default/payment-gateway> .

    # The qualified predicates, pointing from the entity into each relationship. `arch:source`
    # only points outward, so without these the resources above — and everything they carry —
    # would be reachable only by scanning every `arch:source` in the graph.
    <.../element/component/default/payment-api>
        bs:qualifiedOwnedBy
            <.../relationship/ownedBy--component-default-payment-api--group-default-team-payments> ;
        bs:qualifiedPartOfSystem
            <.../relationship/partOfSystem--component-default-payment-api--system-default-payment-gateway> .
}

One semantic graph per descriptor

A catalog is normally many files, so this converter puts each descriptor's facts in a graph named after that descriptor — its repository path where the run knows one, else its file name:

{base}backstage/service-catalog/graph/semantic/group-payments/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-models/catalog-index-yaml
{base}backstage/service-catalog/graph/model
{base}backstage/service-catalog/graph/provenance

It is what makes "everything this descriptor produced" one GRAPH clause. It is also the only form in which a fact two descriptors both assert is attributable to each of them: two files declaring the same edge mint one relationship resource, and asserting prov:wasDerivedFrom twice on that resource cannot say which file contributed which triple.

  • graph/semantic/{slug} — one per descriptor.
  • The index gets one too. architectureState and the index's elements: / links: / data: assertions are declared there, so the index is the input they were lifted from.
  • graph/model holds the curated model: the arch:Model, its folders and their ordering. Its input is the index, and it is described as derived from it.

A catalog converted as a single file keeps the bare graph/semantic. Partitioning one input would name a graph after the only file there is. There is no flag: the rule follows from how many inputs contribute to the model.

Each graph is described in graph/provenance as a prov:Bundle derived from its own input. Name the graph directly when you have its IRI; go through provenance when what you have is the file:

SELECT ?s ?p ?o WHERE {
  ?g prov:wasDerivedFrom ?src .
  ?src schema:name "catalog-info.yaml" ; dct:isPartOf <https://git.example.org/group/payments> .
  GRAPH ?g { ?s ?p ?o }
}

Pin dct:isPartOf as well as schema:name: every descriptor in every repository is called catalog-info.yaml, so the path alone matches one graph per repository. Do not compose the graph IRI from the path — see Aggregating into a knowledge graph for why it cannot be inverted.

Give the run a way to tell your descriptors apart

A graph is named after the descriptor's repository path, and without one it falls back to the file name. Every Backstage descriptor is called catalog-info.yaml, so a catalog pulled from forty repositories with neither --git-provenance auto nor --source-map puts all forty in one graph.

Nothing false is published when that happens — the graph is described as derived from every file that fed it — but the attribution is no more precise than it was before the split, and the run says so:

…/a/catalog-info.yaml, …/b/catalog-info.yaml share one semantic graph, because nothing
distinguishes them but their file name: … Re-run with --git-provenance auto, or with
--source-map for descriptors pulled from elsewhere, so each file's repository path names its graph.

Retained identity

Every catalog entity carries all four parts of its identity as literals — bs:kind, bs:namespace, bs:name and the compact bs:entityRef — retained source data alongside skos:prefLabel, per the same "keep the notation's own attribute, never replace it" contract the other converters follow (bpmn:name, uml:...). bs:title, bs:uid and bs:etag are emitted only when the source declares them.

bs:kind is not a duplicate of rdf:type, and the ontology is explicit that neither is derived from the other: rdf:type is how the graph classifies the entity, a modelling decision an adopter may revise, while bs:kind is what the descriptor said. bssh:KindTypeAlignmentShape reports divergence at sh:Info as a reconciliation aid rather than a failure.

bs:entityRef is the form Backstage itself circulates — in the catalog API's relations array, in spec.owner / spec.system / spec.dependsOn, and in the techdocs-entity annotation — which makes it the stable cross-source join key. bs:uid is not: the catalog reassigns it when the identical file is unregistered and re-registered, so it identifies a registration rather than the thing. bssh:EntityRefConsistencyShape checks the reference against the other three fields.

All four keep the case the source used. Only the IRI is folded to lowercase — see entity references.

One caveat on bs:entityRef

The ontology describes it as read from the source rather than assembled, which is right for a lift of the catalog API. An authored catalog-info.yaml has no entityRef field, so this converter composes it from kind, namespace and name — meaning bssh:EntityRefConsistencyShape is checking our own arithmetic here, and only becomes real evidence for an API-based lift.

Architecture

Backstage catalog YAML (single or multi-doc)
  → BackstageParser (SnakeYAML, multi-document)
  → BackstageModel (entities + derived relationships)
  → LinkedArchiEmitter (extends BaseLinkedArchiEmitter)
  → TriG / Turtle output

How relationships are derived

The parser extracts relationships from spec fields automatically:

Field Relationship emitted
spec.owner: team-x {entity} → bs:Ownership → {Group:default/team-x}
spec.system: my-system {entity} → bs:SystemMembership → {System:default/my-system}
spec.domain: my-domain {entity} → bs:DomainMembership → {Domain:default/my-domain}
spec.providesApis: [api-1] {entity} → bs:APIProvision → {API:default/api-1}
spec.consumesApis: [api-2] {entity} → bs:APIConsumption → {API:default/api-2}
spec.dependsOn: [resource:default/db] {entity} → bs:ResourceUsage → {Resource:default/db}
spec.memberOf: [team-x] {entity} → bs:GroupMembership → {Group:default/team-x}

Entity references

A reference is read as Backstage defines it — [<kind>:][<namespace>/]<name> — and resolved to the kind/namespace/name triplet that identifies the entity, so every way of writing the same reference reaches the same element:

Written in the catalog Resolves to Element IRI
team-payments Group:default/team-payments — kind from the field, namespace from the referring entity element/group/default/team-payments
default/payment-gateway System:default/payment-gateway — kind from the field element/system/default/payment-gateway
System:payment-gateway System:default/payment-gateway — namespace from the referring entity element/system/default/payment-gateway
resource:default/payment-db Resource:default/payment-db — kind and namespace as written element/resource/default/payment-db

The resolved reference is the element IRI, written as a path. One function mints and compares, which is what keeps a reference and the entity it names from minting different addresses.

The comparison ignores case, because Backstage's does: resource:default/payment-db and the Resource:default/payment-db built from the entity's own kind: field are one reference typed two ways, and the IRI is lowercased so both reach one node. Mixed case is legal in a Backstage name, so MyService and myservice are one entity too. The case the author wrote survives in bs:kind, bs:name and bs:entityRef.

The kind is in the id, because it is part of the identity. A Backstage name is unique per kind within a namespace, so Component:default/payments-service and API:default/payments-service are two entities — a service and the interface it publishes, named alike, which is the ordinary case rather than a corner one. An id of {namespace}--{name} gave them one node carrying both bs:Component and bs:API, both spec.type vocabularies, and every relationship of both, with nothing in the output saying so.

Three path segments rather than one, because a Backstage name may itself contain -, _ and ., so component--default--pay--ments cannot be split back into a triplet. Neither : nor / is legal in any part of a reference, so the path form is reversible and needs no percent-encoding. A consumer that writes one file per resource gets one directory level per part — see IRI path shape.

A reference the model does not declare still mints the id that entity would have, so the two sides join if it arrives later, and the run warns:

[WARN] catalog-info.yaml: 'ownedBy' on System:default/payment-gateway names
Group:default/team-payments, which model 'payments-services' does not declare. …

That is the usual signal of a typo, and the usual signal of a catalog split across files: IRIs carry the model id, so an entity referenced from another file needs the files converted into one model — a directory input, or one --model-id — rather than merged afterwards.

One reference is read more leniently than Backstage reads it. spec.dependsOn may name a Component or a Resource and Backstage defaults neither, so it requires the kind; this converter assumes Resource when it is missing. A bare name meant as a component therefore resolves to a resource that does not exist and is reported as unresolved.

Custom spec fields

A house field in spec — deployedTo, maintainedBy, anything an organisation invents — is read as a relationship where --type-mapping declares it, and reported rather than dropped where it does not. Backstage permits such a field while advising against it, and real catalogs carry them, so the converter reads one when told what it means and never guesses.

The documented keys, read without any configuration:

From Keys
metadata name, namespace, title, description, tags, uid, etag, annotations, labels, links
spec, onto the entity type, lifecycle, definition, visibility, profile
spec, as a relationship owner, system, domain, subcomponentOf, subdomainOf, parent, providesApis, consumesApis, dependsOn, dependencyOf, memberOf, children, members

Which of those apply to a given entity is per kind.

Every other spec key is a custom field. Nothing about one is inferred, because nothing about one is inferable: deployedTo: prod-cluster and tier: gold are the same YAML shape, and only the author knows that the first names an entity and the second is a value. Saying which is what switches the field on at all.

Declaring a field

Two sections, because the reference/literal distinction is the whole point:

# type-mapping.yml
spec-relations:               # the value names another resource
  deployedTo: Resource        # …and a bare reference in this field defaults to kind Resource

spec-literals:                # the value is a value
  - tier

predicates:
  deployedTo: https://vocab.example.org/arch#deployedTo    # the predicate to write
  tier: https://vocab.example.org/arch#tier

relationships:
  deployedTo: https://vocab.example.org/arch#Deployment    # optional: makes it a qualified edge
Section Says Required
spec-relations: that the key is read as a reference, and the kind a bare one defaults to one of the two — an undeclared key is reported and not read
spec-literals: that the key is read as a literal one of the two
predicates: the predicate IRI, for either kind yes, unless --ns-vocab supplies a namespace to mint the field name into
relationships: the class of the qualified relationship no, and it applies to a relation field only — see What is emitted

spec-relations: is a map because a reference needs a default kind; spec-literals: is a list because a literal needs nothing per key — not even a datatype, which the catalog already stated. Declaring one key in both is contradictory, so nothing is emitted for it and the run says so.

The default kind plays exactly the part the descriptor format plays for a documented field: Group for spec.owner, System for spec.system. It is what lets deployedTo: prod-cluster resolve without the author writing resource:default/prod-cluster every time.

java -jar backstage2linkedarchi.jar convert catalog.yaml \
  --base-iri https://example.org/la/ --model-id catalog \
  --type-mapping type-mapping.yml \
  -o out.trig

A declared field that a descriptor does not carry produces nothing and says nothing. The declaration describes the catalog's schema, not every entity in it.

How a value is resolved

The three rules every extension route uses, plus one at the end that belongs to Backstage:

Value Read as Result
example:prod-cluster a prefixed name against namespaces: in --type-mapping that IRI — a target outside the catalog
https://k8s.example.com/clusters/prod an absolute IRI itself, untouched
prod-cluster, resource:default/prod-cluster a Backstage entity reference, with the declared default kind and the referring entity's namespace the element IRI of that entity

The third row is the one that mints an address rather than using one. deployedTo: aws-account-012345678901 with deployedTo: Resource declared resolves to Resource:default/aws-account-012345678901, and that triplet is the path:

{--base-iri}backstage/{--model-id}/element/resource/default/aws-account-012345678901
                                           └ kind ┘ └ ns ─┘ └ name ───────────────┘

Which is the same IRI the Resource entity's own descriptor mints, whether or not this model contains it — that is what makes the two sides join, and why a reference to an entity the catalog does not declare still points somewhere and warns rather than being dropped.

Where the target's kind comes from

The kind in that path is read from the reference, never looked up from the target. The value wins where it states one, and spec-relations: supplies it where the value does not:

Value Kind used From
aws-account-012345678901 Resource the declaration — deployedTo: Resource
component:default/legacy-host Component the value, overriding the declaration
Resource:aws-account-012345678901 Resource the value; namespace falls back to the referring entity's

That is the same rule the documented fields follow — spec.owner: team-payments is a Group because the descriptor format says so, and spec.owner: user:default/ada is a User because the value says so.

A wrong default kind produces a dangling edge, quietly

Nothing searches the catalog for an entity of that name under a different kind. Declaring deployedTo: Resource while the targets are actually components gives:

spec:
  deployedTo: prod-cluster        # but prod-cluster is declared `kind: Component`

an edge to element/resource/default/prod-cluster — a node nothing describes — while element/component/default/prod-cluster sits in the same file untouched. The run warns:

[WARN] catalog.yaml: 'deployedTo' on Component:default/order-service names
Resource:default/prod-cluster, which model 'catalog' does not declare. The relationship points at
element/resource/default/prod-cluster, a node nothing in this model describes.

The fix is the declaration, or the kind written into the value. This is why the kind is part of the identity rather than decoration: Component:default/prod-cluster and Resource:default/prod-cluster are two entities, and a converter that guessed between them would silently merge a cluster with a service that happened to share a name.

Where a field's targets are genuinely of mixed kinds, leave the kind out of the declaration's reach by writing it in each value; the declared default is only ever a convenience for the uniform case.

A list yields one statement per item, which is the case this exists for:

spec:
  deployedTo:
    - aws-account-012345678901
    - aws-account-109876543210

Two values, two edges. A scalar and a single-item list read identically.

An entity reference the catalog does not declare still mints the IRI that entity would have and warns, exactly as spec.dependsOn does — see entity references.

A declared prefix wins, including over a kind name

Resolution tries namespaces: first, because an author's own declaration is better evidence than a guess — the rule the rest of the codebase follows. So declaring a prefix that happens to be a Backstage kind, resource: or user:, takes that spelling away from the entity-reference reading in these fields. Nothing refuses it; it is a configuration nobody writes twice, and failing a whole run over it would be worse than the ambiguity.

--ns-global-id is deliberately not consulted here. It would claim every bare name for an IRI base and take prod-cluster away from the reading a catalog almost always means.

When the target is not a Backstage entity

Everything above is the entity-reference row of how a value is resolved. A prefixed name and an absolute IRI — the other two rows — are addresses already: there is no triplet to complete, so the declared kind is never consulted. That is the route for a target the catalog does not describe, such as a cluster in a Kubernetes API, a record in a CMDB, or a node in another model sharing the graph.

The entry under spec-relations: is still required, because it is what makes the field readable at all: a key neither section declares is reported and not read. Where a field's values are all IRIs its kind is inert, so keep it as the kind a bare value would mean — the declaration then still says something true if one is ever written.

# type-mapping.yml
namespaces:
  example: https://vocab.example.org/id/

spec-relations:
  hostedAt: Resource        # inert while the values are IRIs, and required for the field to be read
predicates:
  hostedAt: https://vocab.example.org/arch#hostedAt
spec:
  hostedAt: https://k8s.example.com/clusters/prod    # itself, untouched
  # or example:prod-cluster → https://vocab.example.org/id/prod-cluster

Nothing is asserted about an external target. It is the object of the direct triple, or the arch:target of the qualified relationship, and that is all — no rdf:type, no label, no folder membership. Nothing in this model claims to describe it, so there is no absence to report either, unlike an entity reference the catalog does not declare. Characterising it belongs to whoever publishes its namespace; the IRIs are chosen so that graph and this output meet. core-shapes#QualifiedRelationshipShape is satisfied, since what it requires is one arch:source and one arch:target — not that either end be described here.

A kind the ontology does not model is still a kind. Where the target is a catalog entity but of a kind your organisation added, declare that kind: nothing checks the value against the seven.

spec-relations:
  deployedTo: Cluster
elements:
  cluster: https://vocab.example.org/arch#Cluster    # keyed on the lowercased kind

deployedTo: prod-cluster then mints element/cluster/default/prod-cluster, the same IRI a kind: Cluster descriptor mints — so the two sides join exactly as they do for a published kind. The elements: entry is what gives that kind a class on the entity itself; without it, and without --ns-vocab, entities of that kind are emitted untyped and the run reports it — see kinds with no ontology class.

What is emitted

Two shapes, and which one you get depends on whether relationships: names a class.

With a class — a full arch:QualifiedRelationship, indistinguishable in structure from the edges the documented fields produce, so the core relationship-endpoint contract applies to it:

<.../relationship/deployedTo--component-default-order-service--resource-default-aws-account-012345678901>
    a arch:QualifiedRelationship, arch:ModelConcept, <https://vocab.example.org/arch#Deployment> ;
    arch:source <.../element/component/default/order-service> ;
    arch:target <.../element/resource/default/aws-account-012345678901> ;
    dct:isPartOf <.../folder/Relationships> .

Without one — the direct triple alone:

<.../element/component/default/order-service>
    <https://vocab.example.org/arch#deployedTo> <.../element/resource/default/aws-account-012345678901> .

That is not a lesser fallback, it is the honest output. A qualified relationship needs a class, and this converter will not invent one: bs: is a published document it reads and does not own, so bs:Deployment would dress a house term up as part of the Backstage ontology with nothing declaring its meaning — the same rule that keeps bs:gitlabProjectSlug from being minted for an annotation.

One consequence for --emit-direct-rel-triples: it gates the direct triple where a qualified relationship already carries the statement, since there the triple is only a shortcut. Where there is no class the direct triple is the statement's only form, so it is always written and the flag does not apply.

The bssh: shapes will not validate a house class — they know published terms only — but core-shapes#QualifiedRelationshipShape still checks that both endpoints exist.

A field under spec-literals: becomes a property of the entity, not an edge from it — the same shape an annotation produces, and emitted beside them:

spec:
  tier: gold
  replicas: 3
<.../element/component/default/order-service>
    <https://vocab.example.org/arch#tier> "gold" ;
    <https://vocab.example.org/arch#replicas> "3"^^xsd:integer .

A list yields one triple per item here too, so a multi-valued field behaves the same whichever kind it is.

The datatype comes from the catalog

Nothing declares it and nothing is guessed: the YAML loader has already resolved the scalar by the time the converter sees it, and that resolution is mapped onto XSD.

Written in the catalog YAML resolves to Emitted as
gold, "42", 012345678901 String a plain literal (xsd:string)
3, -7, 0x1F (→ 31) Integer xsd:integer
99999999999999999999 BigInteger xsd:integer
99.95, 1.0, .inf, .nan Double xsd:double
true, false Boolean xsd:boolean
2024-01-15, 2024-01-15T10:30:00Z Date xsd:dateTime, normalised to UTC
empty, null, ~ null nothing — reported as no readable value

Integer becomes xsd:integer rather than the narrower xsd:int, because the width of the box the loader chose is not something the catalog said.

YAML 1.1 implicit typing, which the loader still applies

criticalRegion: no does not produce the string "no". YAML reads yes, no, on and off as booleans, so that field arrives as "false"^^xsd:boolean — the Norway problem, and it is upstream of this converter rather than something it can fix.

Likewise version: 1.0 is a double and serialises as 1.0, losing any distinction from 1.00.

Quoting is the remedy, and it belongs in the catalog: "no" and "1.0" are strings. Leading zeros are already safe — 012345678901 stays a string, so an account id keeps its shape.

No OWL axioms are emitted for either term. The predicate from predicates: is used, and the class from relationships: appears as an rdf:type object, but the graph does not declare a owl:ObjectProperty, a owl:DatatypeProperty, rdfs:domain or rdfs:range for them, nor a owl:Class for the class. That holds for the published bs: terms too — a converted graph states facts about a catalog and leaves the vocabulary to whoever publishes it. If downstream reasoning or validation needs example:deployedTo characterised, declare it in your own ontology and load it alongside; the IRIs are chosen so that ontology and this output meet.

What is reported

Situation Reported Why
a spec key neither section declares — whether your organisation added it or it is a Template / Location field once per run, grouped by key, in one message a key the converter discarded and one the catalog never carried look identical in the output
a declared key with no predicate and no --ns-vocab once per run the statement reached nothing, so saying so is the only alternative to losing it quietly
a key declared in both sections once per run the two declarations contradict each other, and choosing one silently would hide it
a declared key with no readable value — empty, or a nested mapping once per run toString() on a YAML node yields {region=eu-central-1}, which is a rendering of a node rather than a fact
a value that is not a readable reference per value usually a typo
a declared key a descriptor does not carry no the declaration is about the schema
[WARN] catalog.yaml: 1 custom spec field(s) were not read: costCentre. Backstage allows a house field
in spec but publishes no term for one, so each is read only once declared: add it under
'spec-relations:' in --type-mapping with the kind a bare reference defaults to — e.g.
'deployedTo: Resource'.

deployedTo is not a Backstage field

Backstage's well-known relations documents seven relation pairs — ownedBy/ownerOf, providesApi/apiProvidedBy, consumesApi/apiConsumedBy, dependsOn/dependencyOf, parentOf/childOf, memberOf/hasMember, partOf/hasPart — and no deployment relation among them. No kind declares a spec.deployedTo, and no upstream processor writes or reads one. A catalog carrying it got it from a house processor or a vendor distribution.

Backstage sanctions the extension — adding new fields to the spec object states that a kind's schema does not forbid unknown keys and the catalog stores them, and extending the model covers custom kinds and relation types alongside — so the field is legitimate, and this converter reads it. It is simply not shared vocabulary, which is why you supply the IRIs: there is no bs: term to map it onto, that namespace being a published document this converter reads and does not own.

Deployment has a published term elsewhere in this toolchain — Structurizr's structurizr:deployedOn, from a containerInstance — and the nearest thing a catalog has without any configuration is a spec.dependsOn naming a Resource, which additionally carries bs:ResourceUsage.

A kind outside the seven is not invented in bs: either. Its class comes from elements: in --type-mapping (keyed on the lowercased kind) or from --ns-vocab; failing both, the entity carries no notation class and the run reports it. It is still an arch:Element with its labels, identity fields and every relation its spec states — see kinds with no ontology class.

Three routes carry a custom statement into the graph:

Written in Emits Requires
a declared spec field the catalog file itself a link or a typed literal spec-relations: or spec-literals:, and a predicate
metadata.annotations the catalog file itself a plain literal only --ns-vocab, or namespaces: in --type-mapping
the diagram index a file beside the catalog a literal or a link --emit-extension-data and --diagrams-index

Pick on where the statement is authored, not on what it is — all three routes now carry either kind. A declared spec field keeps the value in the descriptor beside everything else about the component, and is the only route that preserves a datatype: an annotation value is a string even when it reads 3. The index keeps the statement out of a catalog you do not own.

For a catalog you author, prefer metadata.annotations under a domain-prefixed key over a new spec field. The support described above exists because catalogs that already carry house fields have to be convertible, not as an encouragement to add more.

This follows Backstage's own guidance. Adding new fields to the spec object of an existing kind records two objections: the field risks colliding with a later addition to the core model, and the data ordinarily belongs in labels or annotations, or in a new type value. The example intent that section gives is a field stating whether a component runs in prod or staging — deployedTo in all but name.

The decisive argument is not advisory. A custom spec field is inert in Backstage as well. Relations there are emitted by processors, not derived from spec keys, so spec.deployedTo produces no relation, no graph edge and no view in a stock instance: the catalog stores the key and returns it unchanged. It acquires meaning only through a custom processor the organisation then maintains. Annotations are the documented channel for metadata consumed by plugins and for links into external systems, so they are read by convention rather than only by bespoke code. This converter matches upstream behaviour in both cases.

Four rules follow.

Require a domain prefix — example.org/deployed-to, not deployedTo. Backstage reserves the backstage.io prefix, expects a key concerning a third-party system to carry a domain that plausibly owns it, and treats an unprefixed key as local to one instance and unfit to travel outside the organisation. The two forms also take different routes here: an unprefixed key requires --ns-vocab, a prefixed key resolves through namespaces: in --type-mapping. The prefixed form is preferable on both counts, since it places the namespace decision in a reviewed configuration file rather than on a command line, where two teams could mint one predicate for two meanings.

Either annotations or labels reaches the graph, so choose on meaning. Both are read, and both become one predicate per key by the same rules — see Labels. Backstage's own split is the one to follow: annotations for metadata a plugin consumes or a link into an external system, labels for something classifying that you expect to filter or select on. A deployment target is arguably the latter.

Prefer a published relation where the statement is an edge. An annotation value is always a literal, so deployedTo arrives as the string "prod-cluster", not as a link to the Resource of that name, and nothing traverses it. Where a query must walk component → cluster, spec.dependsOn naming the resource states the edge with published terms — bs:Dependency and bs:ResourceUsage — and keeps catalog-info.yaml inside the documented schema. A declared spec field is the option when the relation genuinely is not a dependency and the distinction matters to your queries; the cost is a term only your organisation understands.

Model one entity per thing, not one per environment. Backstage advises against separate entities per environment, preferring a single canonical entity whose variation across environments is presented by a plugin. Treating deployment as an attribute or a dependency edge is consistent with that; minting order-service-prod and order-service-staging is not.

The diagram index remains the route where a custom predicate must be an edge and the catalog cannot be changed — a vendor-managed descriptor, for instance. For a catalog you do author it is the weaker of the two link-capable routes, since the value is then stated in a second file with nothing checking it against the descriptor.

Labels

metadata.labels is read with no configuration, and each key becomes its own predicate — the same two routes an annotation with no published term takes:

metadata:
  name: order-service
  labels:
    deployments.example.net/register-srv: "true"   # prefix resolved via namespaces:
    tier: gold                                       # bare key minted under --ns-vocab
<.../element/component/default/order-service>
    <https://vocab.example.org/k8s#register-srv> "true" ;
    <https://vocab.example.org/arch#tier> "gold" .

bs:label is not the predicate. The ontology publishes bs:label as an abstract super-property for an entity's labels, which describes the shape of the vocabulary rather than naming something to write: a key/value pair on a super-property loses the key. Per-key predicates are what that super-property is the parent of, so that is what is emitted.

A label value is always a plain string. Backstage borrows Kubernetes label semantics and says outright that both key and value are strings, so unlike a spec literal nothing is datatype-inferred — "true" stays a string.

The prefix rules are the annotation rules: backstage.io/ is reserved, a key concerning a third-party system should carry the domain that owns it, and a bare key is local to one Backstage instance. A key whose local part cannot be an IRI local name is reported rather than mangled.

House metadata keys

The metadata object is open for extension just as spec is, with the same caution attached. A key outside the documented set is declared under metadata-literals:, which is deliberately separate from spec-literals: — metadata.tier and spec.tier are two statements, and one declaration must not answer for both:

# type-mapping.yml
metadata-literals:
  - costCentre
predicates:
  costCentre: https://vocab.example.org/arch#costCentre

Datatypes come from YAML exactly as for a spec literal, so replicas: 3 arrives as xsd:integer. There is no relation counterpart: metadata is where Backstage puts descriptive fields, and a reference belongs in spec or in a label. An undeclared key is reported, with the reminder that a key/value classifier usually wants metadata.labels, which needs no declaration at all.

Annotations

metadata.annotations is Backstage's own open extension point, and unlike an undeclared spec key it is read with no configuration. The keys with defined meanings are the well-known annotations, and those are lifted to named bs: properties. A key the published ontology has no term for needs a namespace someone owns.

A domain-prefixed key, the recommended form, resolves through namespaces: in --type-mapping:

# catalog-info.yaml
metadata:
  name: order-service
  annotations:
    example.org/deployed-to: prod-cluster
# type-mapping.yml
namespaces:
  example.org: https://vocab.example.org/arch#
java -jar backstage2linkedarchi.jar convert catalog.yaml \
  --base-iri https://example.org/la/ --model-id catalog \
  --type-mapping type-mapping.yml \
  -o out.trig
<.../element/component/default/order-service>
    <https://vocab.example.org/arch#deployed-to> "prod-cluster" .

An unprefixed key is minted onto --ns-vocab instead: deployedTo with --ns-vocab https://vocab.example.org/arch# gives <https://vocab.example.org/arch#deployedTo>. Backstage treats such a key as local to one instance, so the prefixed form is preferable for anything that leaves the organisation.

A prefixed key is deliberately not minted onto --ns-vocab with its prefix dropped, because github.com/project-slug and gitlab.com/project-slug would then land on one predicate holding two values that mean different things.

The value is always a literal. namesResource never treats an annotation value as a reference, so "prod-cluster" is a string and not a link to the Resource entity of that name — a query cannot traverse it. That is the ceiling of this route, and the reason the index one exists.

Keys that reach no predicate are reported once per run, grouped by key rather than by entity:

[WARN] catalog.yaml: 2 annotation(s) have no published bs: term and reached no predicate:
costCentre, deployedTo. Pass --ns-vocab <your namespace> to publish them under a namespace you own.

Converting a catalog API dump

Two root fields exist only in what the catalog API returns, never in an authored descriptor: Backstage states that a descriptor file is not supposed to contain relations, and status is written by the catalog's own ingestion. Both are read, so a lift of a dump keeps what the dump knows and an authored file is unaffected.

relations — the authoritative array

Backstage calls produced relations the authoritative source, and an array carries edges a custom processor derived that no spec field states. It also carries both directions of every relation, so one ownership appears as ownedBy on the component and ownerOf on the group — and again in spec.owner.

Every entry is normalised onto the ontology's single canonical direction, so all three converge on one edge rather than three:

Relation type in the array Emitted as
ownedBy / ownerOf bs:Ownership
providesApi / apiProvidedBy bs:APIProvision
consumesApi / apiConsumedBy bs:APIConsumption
dependsOn / dependencyOf bs:Dependency
childOf / parentOf bs:GroupParentage
memberOf / hasMember bs:GroupMembership
partOf / hasPart depends on the endpoints — see below
anything else no published class; declarable via relationships: and predicates:

In each pair the second name is the reverse, and its endpoints are swapped rather than emitted as written. That is the same treatment spec.children, spec.members and spec.dependencyOf get, and it is what makes the merge work: identical type and endpoints compose one relationship notation, hence one IRI.

partOf is the one type that cannot be mapped by name. Backstage collapses four containments onto it while the ontology publishes a class for each, so the kinds at the ends decide — and a targetRef always states its kind:

From → to Emitted as
Component, API or Resource → System bs:SystemMembership
Component → Component bs:ComponentComposition
System → Domain bs:DomainMembership
Domain → Domain bs:DomainHierarchy
any other pair reported, not emitted

status — the catalog's own report

status.items says whether the entity was ingested cleanly. The ontology publishes nothing for it, so nothing is minted in bs:; with --ns-vocab the items become one node each, the shape metadata.links already uses, because a level and a message flattened onto the entity lose which belongs to which:

<.../element/component/default/order-service>
    <https://vocab.example.org/arch#hasStatusItem> [
        <https://vocab.example.org/arch#statusType> "backstage.io/catalog-processing" ;
        <https://vocab.example.org/arch#statusLevel> "error" ;
        <https://vocab.example.org/arch#statusMessage> "NotFoundError: File not found"
    ] .

Without --ns-vocab the items are reported and dropped. The nested error object is not read: its shape is not fixed upstream and a stack trace is not an architectural fact — the message says the same thing in the form a person reads.

That the ontology has no status vocabulary, and no class for Template or Location, is filed upstream in todo/UPSTREAM-REQUEST-backstage-kinds-and-status.md.

Fields per kind

What the descriptor format declares for each kind, and what this converter does with it. required and optional are Backstage's own markers, not this converter's: nothing here is enforced at conversion time, because a missing required field is the catalog's validation to fail, not a lift's. The shapes are where per-kind expectations are checked in this pipeline.

Every kind additionally accepts the metadata block — name, namespace, title, description, tags, annotations, links — all of which are read for every kind.

Component → bs:Component

Field Backstage Emitted as
spec.type required bs:componentType → a bs:ComponentType individual
spec.lifecycle required bs:lifecycleState individual
spec.owner required bs:Ownership → Group
spec.system optional bs:SystemMembership → System
spec.subcomponentOf optional bs:ComponentComposition → Component
spec.providesApis optional bs:APIProvision → API, one per item
spec.consumesApis optional bs:APIConsumption → API, one per item
spec.dependsOn optional bs:Dependency, plus bs:ResourceUsage for a Resource target
spec.dependencyOf optional the same bs:Dependency, direction swapped: each listed entity → this one

API → bs:API

Field Backstage Emitted as
spec.type required bs:apiType → an bs:APIType individual — openapi, asyncapi, graphql, grpc known
spec.lifecycle required bs:lifecycleState individual
spec.owner required bs:Ownership → Group
spec.definition required bs:definition literal — the interface document itself, copied verbatim
spec.system optional bs:SystemMembership → System
spec.visibility not in the schema bs:apiVisibility → a bs:ApiVisibility individual, read defensively — see the note below

An API is the one kind with a required field carrying a document rather than a reference. spec.definition is emitted as a literal exactly as written, so an OpenAPI document arrives in the graph as a string; the converter does not parse it.

spec.visibility is read though no published schema declares it

The descriptor format has no visibility field on an API today, while the system model's prose does talk about an API having a visibility, and the ontology already publishes bs:ApiVisibility with public, restricted and private. It is therefore read if present, so an organisation whose own tooling populates it loses nothing, and absent otherwise. This is the one field this converter reads that Backstage does not document.

Resource → bs:Resource

Field Backstage Emitted as
spec.type required bs:resourceType → a bs:ResourceType individual — database, s3-bucket, kubernetes-cluster known
spec.owner required bs:Ownership → Group
spec.system optional bs:SystemMembership → System
spec.dependsOn optional bs:Dependency (+ bs:ResourceUsage for a Resource target)
spec.dependencyOf optional the same bs:Dependency, direction swapped: each listed entity → this one. Since the target is then this Resource, the edge also carries bs:ResourceUsage

A Resource has no lifecycle in the descriptor format. This converter reads one if present anyway — see the liberties below.

System → bs:System

Field Backstage Emitted as
spec.owner required bs:Ownership → Group
spec.domain optional bs:DomainMembership → Domain
spec.type optional bs:systemType → a bs:SystemType individual — product, service, feature-set known

Domain → bs:Domain

Field Backstage Emitted as
spec.owner required bs:Ownership → Group
spec.subdomainOf optional bs:DomainHierarchy → Domain
spec.type optional bs:domainType → a bs:DomainType individual

Group → bs:Group

Field Backstage Emitted as
spec.type required bs:groupType → a bs:GroupType individual — team, business-unit, product-area, root known
spec.children required bs:GroupParentage, one per item, direction swapped: each child → this group
spec.parent optional bs:GroupParentage: this group → parent
spec.members optional bs:GroupMembership, one per item, direction swapped: each member → this group
spec.profile optional bs:displayName, bs:email, bs:picture literals, flattened onto the entity

spec.children is required and yet lists the other end of the relation, which is why it and spec.members are normalised onto the canonical child→parent and member→group direction rather than minting an inverse property. A Group's own spec.parent already runs that way and is read as written.

User → bs:User

Field Backstage Emitted as
spec.memberOf required bs:GroupMembership → Group, one per item
spec.profile optional bs:displayName, bs:email, bs:picture literals

A User has no spec.type and no spec.owner.

Kinds with no ontology class

Template and Location are documented kinds the published ontology does not model, because a scaffolder form and a catalog location are not architecture. A house kind is in the same position.

An entity of such a kind still converts, and carries no bs: class:

<.../element/location/default/all-components>
    a arch:Element, arch:ModelConcept ;      # and nothing from bs:
    skos:prefLabel "all-components"@en ;
    bs:kind "Location" ;                     # what the descriptor said, retained
    bs:entityRef "Location:default/all-components" .

bs:Location is not minted, for the reason nothing else is: that namespace is a published document this converter reads and does not own. A class comes from elements: in --type-mapping, keyed on the lowercased kind, or from --ns-vocab:

elements:
  location: https://vocab.example.org/arch#CatalogLocation

With --ns-vocab https://vocab.example.org/arch# and nothing else, the kind mints <https://vocab.example.org/arch#Location> under the namespace you own. With neither, the entity keeps its core types and the run says what happened:

[WARN] catalog.yaml: 1 entity kind(s) carry no notation class: Location. The ontology publishes classes
for Component, System, API, Resource, Domain, Group and User; Template and Location are documented
Backstage kinds it deliberately does not model … These entities are still emitted as arch:Element with
their labels, identity and relations — to type them, name a class under 'elements:' in --type-mapping
(keyed on the lowercased kind) or pass --ns-vocab <your namespace>.

The kind-shared fields (spec.type, spec.owner) are read as for any entity. The kind-specific ones are not, and are declarable if you want them: spec.parameters and spec.steps describe a scaffolder form, spec.target, spec.targets and spec.presence where a catalog reads descriptors from.

Documented fields with no ontology term

Every documented field of the seven modelled kinds is converted. Five are not, and all five belong to Template and Location, the kinds the ontology does not model:

Field Kind What it is
spec.parameters, spec.steps Template a scaffolder form — the shape of a wizard, not an architectural fact
spec.target, spec.targets, spec.presence Location where a catalog reads descriptors from, not a relation between entities

They are unmapped by default, not blocked. There is no bs: term for them and this converter will not mint one, but the declaration mechanism is exactly the escape hatch: list the field under spec-literals: with a predicate from a namespace you own and it reaches the graph like any other value.

The converter does not hold a list of these five. They reach the same report as a field your organisation invented, because they are unread for the same reason — no published term names them — and published by the same remedy. A hardcoded copy of another project's schema, existing only to reword a warning, would go stale the moment Backstage added a field to either kind, and would go stale silently:

[WARN] catalog.yaml: 2 spec field(s) reached no predicate and were not read: costCentre, presence. This
converter reads the documented fields of the seven kinds the ontology models; anything else — a field
your organisation added, or a field of the Template and Location kinds — has no published term, and this
converter does not invent one. To publish it, declare the key in --type-mapping: 'spec-relations:' with
the kind a bare reference defaults to (e.g. 'deployedTo: Resource') if its value names another entity, or
'spec-literals:' if its value is a value, together with a predicate under 'predicates:' or a namespace
via --ns-vocab.

So the table above is documentation, where staleness is visible and harmless, rather than code, where it would not be.

Two liberties with the per-kind schema

Both deliberate:

  • spec.lifecycle is read on any kind, not only Component and API where it is required. A Resource carrying one gets it mapped rather than dropped.
  • spec.dependsOn assumes Resource when a reference omits its kind. Backstage requires the kind there precisely because the target may be a Component or a Resource and neither is the default — see entity references.

Directory support

Pass a directory path to recursively process all .yaml / .yml files:

java -jar backstage2linkedarchi.jar convert \
  /path/to/catalog/ \
  --base-iri https://example.org/la/ \
  --model-id full-catalog \
  --format TRIG \
  -o out.trig

Extension data

A catalog can carry statements the Backstage schema has no field for — which capability a component realizes, which LeanIX factsheet it corresponds to, what its cost centre is. --emit-extension-data maps them onto the entities and the model they annotate.

The index is the only route here. BPMN has extensionElements and PlantUML has comments no renderer touches; a .yaml owned by Backstage has neither, and adding a key to it would be a claim on a format this converter only reads. See Extension data for what the routes share.

prefixes:
  am: https://meta.linked.archi/archimate3/onto#
  arch: https://meta.linked.archi/core#
  kg: https://example.org/graph/
  x: https://example.org/vocab#
diagrams:
  - id: catalog
    file: catalog.yaml
    links:
      arch:architectureState: arch:Baseline   # about the model
    data:
      x:reviewedBy: Jane Doe
    elements:
      order-service:                          # about one entity
        links:
          am:realizes: kg:CAP-OrderManagement
        data:
          x:costCentre: CC-4711
java -jar backstage2linkedarchi.jar convert catalog.yaml \
  --diagrams-index index.yaml --base-iri https://example.org/la/ \
  --emit-extension-data -o out.trig

Both flags are needed. --emit-extension-data on its own does nothing here, because the assertions arrive with the index and there is no in-file route to fall back to. Note also that an index restricts the run: only the files it lists are converted.

Worked example: a custom deployedTo

spec.deployedTo is a house field, so it is read only where --type-mapping declares it — see Custom spec fields. Where you would rather not configure the catalog's schema at all, or cannot because the descriptor belongs to someone else, restating the fact in the index turns it into an edge between the two entities the catalog already declares.

# catalog.yaml
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
  name: order-service
spec:
  type: service
  lifecycle: production
  owner: team-orders
  deployedTo: prod-cluster        # house field, left undeclared here — reported, not read
---
apiVersion: backstage.io/v1alpha1
kind: Resource
metadata:
  name: prod-cluster
spec:
  type: kubernetes-cluster
  owner: team-platform
# index.yaml
prefixes:
  x: https://vocab.example.org/arch#
  cat: https://example.org/la/backstage/catalog/element/
diagrams:
  - id: catalog
    file: catalog.yaml
    elements:
      component:default/order-service:
        links:
          x:deployedTo: cat:resource/default/prod-cluster
java -jar backstage2linkedarchi.jar convert catalog.yaml \
  --diagrams-index index.yaml \
  --diagrams-root . \
  --base-iri https://example.org/la/ \
  --emit-extension-data \
  -o out.trig
<https://example.org/la/backstage/catalog/element/component/default/order-service>
    <https://vocab.example.org/arch#deployedTo>
        <https://example.org/la/backstage/catalog/element/resource/default/prod-cluster> .

Three things about that are worth knowing before writing one.

A subject and a target resolve by different rules. The subject is looked up among the entity's names — bare name, entity ref, minted id, per Naming an entity. The target is not: it goes through the same resolver every route uses, which accepts a CURIE against prefixes:, an absolute IRI, or a bare name with --ns-global-id, and knows nothing about Backstage refs. Writing x:deployedTo: resource:default/prod-cluster therefore reads resource as an undeclared prefix, and the value is kept as a literal and reported:

[WARN] catalog.yaml: 'resource:default/prod-cluster' on 'component:default/order-service' is not a
resolvable IRI — prefix 'resource:' is not declared. Kept as a literal, so it is not a link in the
graph.

Hence the cat: prefix above, bound to this model's element namespace, {--base-iri}backstage/{model-id}/element/. That does couple the index to the base IRI and the model id, so an index written this way follows a change to either. Pointing at a system outside the catalog is simpler, since an absolute IRI needs no prefix at all:

links:
  x:deployedTo: https://k8s.example.com/clusters/prod

The result is a plain triple, not a qualified relationship. No bs: class, no arch:source / arch:target, no membership of the Relationships folder — unlike the relations derived from spec. --emit-direct-rel-triples is unrelated and does not affect it. If the edge needs the qualified form with attributes of its own, the statement belongs in a graph the aggregating repository builds, not here.

A term that runs the other way needs no inverse. To say it from the cluster's side while reusing one predicate, use direction:

elements:
  component:default/order-service:
    links:
      x:hosts: { target: cat:resource/default/prod-cluster, direction: Backward }

That emits {cluster} x:hosts {component}. Both writes the pair. A literal target cannot be the subject of an inverse, so Backward on one is reported and only the forward triple is written.

For a value that really is just a string, data: puts it on the entity without pretending it is a reference:

elements:
  component:default/order-service:
    data:
      x:deploymentTier: gold

Naming an entity

A Backstage entity has several usable names, and any of them resolves:

kind: Component, namespace: default, name: order-service
entity name order-service
entity ref Component:default/order-service, or the lowercase form Backstage prints
minted id component/default/order-service

The ref works in both casings because Backstage's docs write it lowercase while the catalog file spells kind capitalised — the same reference typed from two places. A bare name shared by two entities is refused rather than guessed at, and reported; use the ref, which is unique. An annotation naming an entity the catalog does not hold is reported too, since that is what a rename leaves behind.

An entity ref needs no quoting despite its colon — see colons and quoting for the one case that does:

elements:
  component:default/order-service:        # fine unquoted
    links:
      am:realizes: kg:CAP-OrderManagement

About the model

links: and data: written directly on the entry describe the model. This converter emits no arch:View — a catalog is not a diagram — so a view-level block from the model/views schema lands on the model as well, the same collapse the flat schema makes for every entry-level field.

Option Default Description
--emit-extension-data false Map the index's elements:, links: and data: entries into the graph
--ns-global-id none Base IRI for a link target written as a bare name. Requires --emit-extension-data

Architecture state

A catalog can be indexed as describing the current architecture or a planned one:

diagrams:
  - id: full-catalog
    file: catalog.yaml
    architectureState: baseline      # or target, transitional

A checked value, so a typo fails the run rather than reaching the graph. This converter declares no arch:View, so the state lands on the model whichever level of the index declares it — the same collapse the flat schema makes for every entry-level field.

See Architecture state, including why it is orthogonal to status:.

Validate

Runs SHACL validation on converted output. Shapes and ontologies are fetched from meta.linked.archi at runtime.

Default shapes are backstage-shapes and core-shapes. The backstage and core ontologies are both loaded for rdfs:subClassOf reasoning.

bssh:BackstageElementLabelShape is what gives this converter's output a naming rule at all, since the core label shapes were withdrawn, and the per-kind shapes check the relations a catalog derives from its spec fields — ownership, system membership, API provision and consumption, resource use, domain and group membership. Switch the naming rule off for a run with --without-shape labels.

Two shapes cover the identity fields:

Shape Severity Checks
bssh:EntityRefConsistencyShape Violation bs:entityRef agrees with bs:kind, bs:namespace and bs:name, compared case-insensitively. They come from one descriptor, so a mismatch means the lift mangled one of them
bssh:KindTypeAlignmentShape Info bs:kind matches the skos:notation of some bs: class the entity is typed with. Reported, not failed: retyping a lifted entity under a later convention is legitimate. Switch it off where retyping is standing policy

A multi-repo run reports its cross-repo references

A single catalog conforms. A run spanning repositories does not, and the shapes are right to say so. spec.owner: team-payments in one repo mints {model}/element/group/default/team-payments inside that model's namespace, while the Group entity declared in another repo is {other-model}/element/group/default/team-payments — two IRIs for one team. bssh:OwnershipShape reports a target that is not a bs:Group, and merging the graphs does not reconcile them, because neither IRI is the other.

Entity references resolves the form a reference is written in, within one model. This is the different question of which model owns an entity that a repo references without declaring — and ADR 0001 places cross-repository uniqueness with the aggregating repository, so the converters do not decide it.

It is not a new failure. Those references already failed core-shapes#QualifiedRelationshipShape, which was in the default set before the Backstage shapes were: 6 violations before, 12 after, all one root cause. What changed is that the report now names the problem (Ownership target must be a Group) instead of stating it generically.

java -jar backstage2linkedarchi.jar validate -i out.trig

# Validate against your own shapes instead
java -jar backstage2linkedarchi.jar validate -i out.trig --shapes ./shapes/catalog-rules.ttl

# Offline, reusing documents fetched by an earlier run
java -jar backstage2linkedarchi.jar validate -i out.trig --asset-dir .assets --offline

Exits 0 when the data conforms, 1 when violations are found, 2 on error.

See Validation for the shared engine, all options, coverage reporting and report format.