Skip to content

Extension data

A model file says what the notation can express. It rarely says everything an architecture repository needs: which capability a process realizes, which application serves a task, who owns a component, what its cost centre is. Extension data is how those statements reach the graph without leaving the notation behind.

Every route produces the same thing — a triple on the resource the statement is about — and differs only in where the statement is written and who writes it.

Which route each notation has

Notation In the source file In the diagram index
BPMN ✓ bpmn:extensionElements ✓
PlantUML ✓ '!la-link / '!la-data comments ✓
Backstage ~ metadata.annotations, and a house spec field the pipeline declares ✓
LeanIX ✗ ✓
ArchiMate ✓ element properties (a separate, older mechanism — see ArchiMate) ✗
Structurizr ✗ ✗

The gap is not arbitrary. A notation gets an in-file route only if the format has somewhere to put data it does not understand and tools preserve it. BPMN has extensionElements, which is exactly that. PlantUML has comments, which every renderer ignores and no tool rewrites. Structurizr has neither, and ArchiMate's own <property> mechanism predates this and works differently.

Backstage is the case where the format could take one, and the answer is a qualified no: this converter will not invent a key in a schema it only reads. What it will do is read a key an organisation has already invented, where spec-relations: in --type-mapping declares what that key means — a declaration by the pipeline about a catalog it does not own, rather than a convention these converters ask authors to adopt. Its own open extension point, metadata.annotations, is read with no configuration at all, but annotation values are literals by definition, so it can state a cost centre and not a link.

The index therefore remains the route for a statement that has to be an edge in a catalog nobody here controls. See Backstage → Custom spec fields.

ArchiMate states links a different way. Its ✓ above is the element-property mechanism, which promotes a key to a predicate and by default carries a literal. The link form is not a second subject-level route but a declaration in --type-mapping: listing a key under object-properties: says its values are references, and they are then resolved to IRIs by the same three rules @type uses. So the principle below still holds for ArchiMate — links and literals are stated, not inferred — it is just stated once per key rather than once per statement, which suits a mechanism where the key is authored in a modelling tool and the declaration belongs to the pipeline. See ArchiMate → Making a property value a link.

LeanIX has a stronger version of the same reason. An export is not authored at all — it is regenerated from the workspace on every pull — so anything written into it would be gone by the next one. If a statement belongs to the workspace it goes in the workspace, as a field or a tag, and comes back through the export on its own. If it belongs to the publishing decision, it goes in the index.

The index route requires only that the converter reads a diagram index, which is why it reaches notations that have nowhere in-file to put anything.

What a statement is about

Most extension data is about an element: '!la-link OrderService am:realizes … needs an OrderService in the diagram. But some things are true of the diagram as a whole and of no element in it — which architecture state it depicts, which viewpoint it conforms to, which decision option it articulates — and others are true of the whole model. Those have no element to hang on, so the subject can be the view or the model instead:

Subject In the source file In the diagram index
an element '!la-link / '!la-data elements: on the entry
the view '!la-view-link / '!la-view-data links: / data: on the view entry
the model — links: / data: on the model: entry
'!la-prefix arch: https://meta.linked.archi/core#
'!la-prefix x: https://example.org/vocab#
'!la-view-link arch:architectureState arch:Target
'!la-view-data x:reviewedBy Jane Doe
prefixes:
  am: https://meta.linked.archi/archimate3/onto#
  arch: https://meta.linked.archi/core#
  kg: https://example.org/graph/
  x: https://example.org/vocab#
model:
  id: order-domain
  links:
    arch:architectureState: arch:Baseline     # about the model
views:
  - id: target-payments
    file: target-payments.puml
    links:
      arch:architectureState: arch:Target     # about this view
    data:
      x:reviewedBy: Jane Doe
    elements:
      OrderService:                           # about an element in it
        links: { am:realizes: kg:CAP-Orders }

The view directives take no subject, because a .puml file is one view and there is nothing else they could be about. Everything else is unchanged: the same prefixes, the same direction tokens, the same --emit-extension-data gate.

Where a checked field exists, prefer it

arch:architectureState has one: architectureState: in the index, '!la-architecture-state in a .puml. Its values are a closed set, so a typo fails the run — whereas written as a generic link, arch:architectureState arch:Targt emits a reference to a term nothing defines and nothing complains, because a generic predicate has no set of values to check against. See Architecture state, which sets out every route per notation and level.

The generic route stays right for open-ended vocabulary, which is most of it.

Two gaps in that table are deliberate rather than pending.

There is no in-file route for the model. A model spans several files, so a model-wide claim written in one of them would have no evident owner, and two files disagreeing would need a conflict rule nobody asked for. The index names a model exactly once, which is where the statement goes.

BPMN has no in-file route for the view either. bpmn:extensionElements nests inside the element it annotates, so there is nowhere in a .bpmn to put a statement about the diagram as a whole. The index covers it.

A predicate that points at the view

Some terms run towards the view rather than away from it. arch:inView goes from a concept to the view that shows it, so articulating a decision option needs the triple the other way round — the same Backward token the element route uses, and the reason it exists rather than a second inverse term being invented:

'!la-view-link arch:inView kg:ReadReplicas Backward
links:
  arch:inView: { target: kg:ReadReplicas, direction: Backward }

Both emit kg:ReadReplicas arch:inView <this view>.

Choosing a route

The two routes exist because the work is usually split, and the right choice follows the split rather than any technical merit:

Source file Diagram index
Written by whoever authors the diagram whoever owns the pipeline
Reviewed with the diagram, in the same commit the publishing configuration
Travels when the file is copied yes no
Survives a diagram rename the statement does; the element reference may not same
Needs the author to learn a convention yes no
Can annotate many models at once no yes

Prefer the source file when the statement is a fact about the thing being modelled and the person drawing the diagram is the one who knows it. "This process realizes the Order Management capability" is a modelling statement; it belongs beside the process, gets reviewed with it, and travels with the file.

Prefer the index when the statement belongs to the publishing decision rather than the model, or when the people who know it do not edit the diagrams. Mapping every component in a repository to its LeanIX factsheet is a platform-team concern; putting it in fifty .puml files spreads one decision across fifty reviews.

They are additive. Both routes apply in the same run and neither overwrites the other, so a platform team's index entries and an author's comments coexist. The one thing that has to resolve one way is a prefix declared in both places, and the source file wins there — it is the more specific declaration, and the author editing it can see it.

What is the same everywhere

Whichever route, whichever notation:

  • A published term is preferred to an invented one. am:realizes, arch:refines, arch:conceptOwner, the skos: mapping terms and owl:sameAs already exist. Nothing prevents a consumer pointing at their own ontology, but a published term is one a downstream consumer already understands. See BPMN → terms worth knowing.
  • A correspondence is skos:exactMatch by default, not owl:sameAs. This is the most common reason to write extension data at all, and the two predicates are not interchangeable — see Stating a correspondence below.
  • A prefix must be declared where the host format will not resolve it. The declaration lives with the route: xmlns: in BPMN, '!la-prefix in PlantUML, prefixes: in the index.
  • A declared prefix beats an absolute IRI reading, because lx:APP-1 and doi:10.1000/x have the same shape and the author's own declaration is better evidence.
  • An unresolvable prefix is reported, never guessed. Emitting nowhere:CAP-1 as an IRI would produce a link that resolves nowhere, which is the defect this whole mechanism exists to remove. The value survives as a literal and the run warns.
  • direction is Forward (the default, and what None means), Backward or Both. Backward is what lets an author reuse a term that runs the other way — am:serves runs from the application — instead of inventing an inverse.
  • Links and literals are stated, not inferred. CC-4711 is a cost centre and CAP-1 is probably an element id, and both are bare strings, so each route has an explicit way to say which is meant: rdf:resource in BPMN, -link versus -data in PlantUML, links: versus data: in the index.
  • Off by default. --emit-extension-data gates it, because a file authored in a modelling tool carries that tool's own extension elements and turning those into triples unasked would change the graph of every existing model.
  • --ns-global-id gives a base to a target written as a bare name. Everything else resolves without it.

Stating a correspondence

The commonest reason to write extension data is to say this thing here is that thing over there — the BPMN task and the ArchiMate process, the PlantUML participant and the CMDB record. That statement has two candidate predicates, they say different things, and a reasoner treats them differently.

skos:exactMatch is the default. It says the two nodes describe the same real-world thing without claiming they are one RDF resource. Each keeps its own types and properties, nothing is inferred onto the other, and the link can be withdrawn without touching either source. It carries no OWL entailment, so the everyday query path stays reasoner-free.

owl:sameAs is the documented exception. It asserts the two IRIs denote one resource, and a reasoner may then merge everything known about both, in both directions. Use it where that merge is the feature you want and identity is confirmed — reconciling an id after a rename is the clear case (see formerIds: in the diagram index). A wrong owl:sameAs silently corrupts every query downstream of it, which is why it is not the default.

Choose with the substitutability test. Could a reasoner safely replace one IRI with the other everywhere, forever? If yes, the nodes are one entity → owl:sameAs. If not → skos:exactMatch. Most cross-model links fail the test, because notations describe the same system at different granularities and for different concerns: a diagram's mention of an application is not the application's record in a catalog, even when both are accurate.

For a correspondence that is real but not exact, skos:closeMatch and skos:relatedMatch say so. Note relatedMatch, not skos:related — the *Match terms are the cross-scheme family, and skos:related is defined for use within one concept scheme.

<!-- BPMN: the predicate is the element name -->
<skos:exactMatch rdf:resource="archi:id-abc123-GeoMapping"/>
'!la-link Shop skos:exactMatch cmdb:APP-webshop
elements:
  OrderService:
    links: { skos:exactMatch: kg:APP-order-service }

Whichever route, the converter transcribes the predicate you wrote and makes no judgement about it. Getting the choice backwards is how a graph ends up merging two resources that were only ever meant to correspond, and nothing in the toolchain will warn you.

Instance-level and type-level

Everything above relates instances: this task, that component. Relating one language's types to another's — bpmn:UserTask to am:BusinessProcess, c4:Container to am:ApplicationComponent — is a different job, and extension data is the wrong place for it.

A type-level correspondence is a statement about the two languages, true of every model written in them, so writing it inside one model file would scope a general claim to an accident of where it was typed. It belongs in a *-crossmappings.ttl document registered on the metamodel via arch:crossLanguageMappings, authored once per language pair and reviewed as an editorial judgement. The predicates are the same SKOS family, with skos:closeMatch and skos:relatedMatch doing most of the work.

Neither level is ever emitted by a converter. Mappings are judgements, not derivations — see DD-11 for the policy and Authoring cross-language mappings for how to write the type-level document.

Referring to an element

This section is about element-level statements only. A view- or model-level statement names no subject, so there is nothing to look up and nothing that a rename can leave behind — which is the practical difference between the two, beyond what they are about.

The in-file BPMN route is the only one that contains its statements: extensionElements is nested inside the element, so there is nothing to name. Every other route refers to an element, and that has a consequence worth knowing — a reference can be left behind by a rename, so every route reports an annotation that matches no element rather than dropping it:

[WARN] orders.puml:12: no element 'Nowhere' in this model, so <am:realizes> was not attached to
anything. Check the name against the diagram.

A PlantUML element has two names, and either resolves:

class "Payment Gateway" as PayGw
code PayGw what the source refers to it by — in every arrow, and in the IRI
display name Payment Gateway what the picture shows, and its skos:prefLabel

The code is the one to write. It is what the rest of the file already uses, so an annotation reads like the diagram around it, and it is the name the IRI is built from. The display name resolves too, because it is what a reader has in front of them; quote it when it contains a space: '!la-link "Payment Gateway" arch:refines kg:CAP-Payments. A name matching two elements is refused rather than guessed at.

Prefer the code, because a label gets edited

A display name is prose and changes for the reader's benefit. An annotation keyed on it has to be revisited when someone retitles the box; one keyed on the code does not, and neither does the element's IRI — see ADR 0010.

Where an element has no as clause the two names are the same string, so there is nothing to choose and nothing to maintain.

Reference