Extension data¶
A model file says what the notation can express. It rarely says everything an architecture repository needs: which capability a process realizes, which application serves a task, who owns a component, what its cost centre is. Extension data is how those statements reach the graph without leaving the notation behind.
Every route produces the same thing — a triple on the resource the statement is about — and differs only in where the statement is written and who writes it.
Which route each notation has¶
| Notation | In the source file | In the diagram index |
|---|---|---|
| BPMN | ✓ bpmn:extensionElements |
✓ |
| PlantUML | ✓ '!la-link / '!la-data comments |
✓ |
| Backstage | ~ metadata.annotations, and a house spec field the pipeline declares |
✓ |
| LeanIX | ✗ | ✓ |
| ArchiMate | ✓ element properties (a separate, older mechanism — see ArchiMate) | ✗ |
| Structurizr | ✗ | ✗ |
The gap is not arbitrary. A notation gets an in-file route only if the format has somewhere to put
data it does not understand and tools preserve it. BPMN has extensionElements, which is exactly
that. PlantUML has comments, which every renderer ignores and no tool rewrites. Structurizr has
neither, and ArchiMate's own <property> mechanism predates this and works differently.
Backstage is the case where the format could take one, and the answer is a qualified no: this
converter will not invent a key in a schema it only reads. What it will do is read a key an
organisation has already invented, where spec-relations: in --type-mapping declares what that key
means — a declaration by the pipeline about a catalog it does not own, rather than a convention these
converters ask authors to adopt. Its own open extension point, metadata.annotations, is read with no
configuration at all, but annotation values are literals by definition, so it can state a cost centre
and not a link.
The index therefore remains the route for a statement that has to be an edge in a catalog nobody here controls. See Backstage → Custom spec fields.
ArchiMate states links a different way. Its ✓ above is the element-property mechanism, which promotes
a key to a predicate and by default carries a literal. The link form is not a second subject-level
route but a declaration in --type-mapping: listing a key under object-properties: says its values are
references, and they are then resolved to IRIs by the same three rules @type uses. So the principle
below still holds for ArchiMate — links and literals are stated, not inferred — it is just stated once
per key rather than once per statement, which suits a mechanism where the key is authored in a modelling
tool and the declaration belongs to the pipeline. See
ArchiMate → Making a property value a link.
LeanIX has a stronger version of the same reason. An export is not authored at all — it is regenerated from the workspace on every pull — so anything written into it would be gone by the next one. If a statement belongs to the workspace it goes in the workspace, as a field or a tag, and comes back through the export on its own. If it belongs to the publishing decision, it goes in the index.
The index route requires only that the converter reads a diagram index, which is why it reaches notations that have nowhere in-file to put anything.
What a statement is about¶
Most extension data is about an element: '!la-link OrderService am:realizes … needs an
OrderService in the diagram. But some things are true of the diagram as a whole and of no
element in it — which architecture state it depicts, which viewpoint it conforms to, which decision
option it articulates — and others are true of the whole model. Those have no element to hang on,
so the subject can be the view or the model instead:
| Subject | In the source file | In the diagram index |
|---|---|---|
| an element | '!la-link / '!la-data |
elements: on the entry |
| the view | '!la-view-link / '!la-view-data |
links: / data: on the view entry |
| the model | — | links: / data: on the model: entry |
'!la-prefix arch: https://meta.linked.archi/core#
'!la-prefix x: https://example.org/vocab#
'!la-view-link arch:architectureState arch:Target
'!la-view-data x:reviewedBy Jane Doe
prefixes:
am: https://meta.linked.archi/archimate3/onto#
arch: https://meta.linked.archi/core#
kg: https://example.org/graph/
x: https://example.org/vocab#
model:
id: order-domain
links:
arch:architectureState: arch:Baseline # about the model
views:
- id: target-payments
file: target-payments.puml
links:
arch:architectureState: arch:Target # about this view
data:
x:reviewedBy: Jane Doe
elements:
OrderService: # about an element in it
links: { am:realizes: kg:CAP-Orders }
The view directives take no subject, because a .puml file is one view and there is nothing else
they could be about. Everything else is unchanged: the same prefixes, the same direction tokens,
the same --emit-extension-data gate.
Where a checked field exists, prefer it
arch:architectureState has one: architectureState: in the index, '!la-architecture-state in
a .puml. Its values are a closed set, so a typo fails the run — whereas written as a generic
link, arch:architectureState arch:Targt emits a reference to a term nothing defines and nothing
complains, because a generic predicate has no set of values to check against. See
Architecture state, which sets out every route per notation and level.
The generic route stays right for open-ended vocabulary, which is most of it.
Two gaps in that table are deliberate rather than pending.
There is no in-file route for the model. A model spans several files, so a model-wide claim written in one of them would have no evident owner, and two files disagreeing would need a conflict rule nobody asked for. The index names a model exactly once, which is where the statement goes.
BPMN has no in-file route for the view either. bpmn:extensionElements nests inside the element
it annotates, so there is nowhere in a .bpmn to put a statement about the diagram as a whole. The
index covers it.
A predicate that points at the view¶
Some terms run towards the view rather than away from it. arch:inView goes from a concept to
the view that shows it, so articulating a decision option needs the triple the other way round — the
same Backward token the element route uses, and the reason it exists rather than a second inverse
term being invented:
Both emit kg:ReadReplicas arch:inView <this view>.
Choosing a route¶
The two routes exist because the work is usually split, and the right choice follows the split rather than any technical merit:
| Source file | Diagram index | |
|---|---|---|
| Written by | whoever authors the diagram | whoever owns the pipeline |
| Reviewed with | the diagram, in the same commit | the publishing configuration |
| Travels when the file is copied | yes | no |
| Survives a diagram rename | the statement does; the element reference may not | same |
| Needs the author to learn a convention | yes | no |
| Can annotate many models at once | no | yes |
Prefer the source file when the statement is a fact about the thing being modelled and the person drawing the diagram is the one who knows it. "This process realizes the Order Management capability" is a modelling statement; it belongs beside the process, gets reviewed with it, and travels with the file.
Prefer the index when the statement belongs to the publishing decision rather than the model, or
when the people who know it do not edit the diagrams. Mapping every component in a repository to its
LeanIX factsheet is a platform-team concern; putting it in fifty .puml files spreads one decision
across fifty reviews.
They are additive. Both routes apply in the same run and neither overwrites the other, so a platform team's index entries and an author's comments coexist. The one thing that has to resolve one way is a prefix declared in both places, and the source file wins there — it is the more specific declaration, and the author editing it can see it.
What is the same everywhere¶
Whichever route, whichever notation:
- A published term is preferred to an invented one.
am:realizes,arch:refines,arch:conceptOwner, theskos:mapping terms andowl:sameAsalready exist. Nothing prevents a consumer pointing at their own ontology, but a published term is one a downstream consumer already understands. See BPMN → terms worth knowing. - A correspondence is
skos:exactMatchby default, notowl:sameAs. This is the most common reason to write extension data at all, and the two predicates are not interchangeable — see Stating a correspondence below. - A prefix must be declared where the host format will not resolve it. The declaration lives with
the route:
xmlns:in BPMN,'!la-prefixin PlantUML,prefixes:in the index. - A declared prefix beats an absolute IRI reading, because
lx:APP-1anddoi:10.1000/xhave the same shape and the author's own declaration is better evidence. - An unresolvable prefix is reported, never guessed. Emitting
nowhere:CAP-1as an IRI would produce a link that resolves nowhere, which is the defect this whole mechanism exists to remove. The value survives as a literal and the run warns. directionisForward(the default, and whatNonemeans),BackwardorBoth.Backwardis what lets an author reuse a term that runs the other way —am:servesruns from the application — instead of inventing an inverse.- Links and literals are stated, not inferred.
CC-4711is a cost centre andCAP-1is probably an element id, and both are bare strings, so each route has an explicit way to say which is meant:rdf:resourcein BPMN,-linkversus-datain PlantUML,links:versusdata:in the index. - Off by default.
--emit-extension-datagates it, because a file authored in a modelling tool carries that tool's own extension elements and turning those into triples unasked would change the graph of every existing model. --ns-global-idgives a base to a target written as a bare name. Everything else resolves without it.
Stating a correspondence¶
The commonest reason to write extension data is to say this thing here is that thing over there — the BPMN task and the ArchiMate process, the PlantUML participant and the CMDB record. That statement has two candidate predicates, they say different things, and a reasoner treats them differently.
skos:exactMatch is the default. It says the two nodes describe the same real-world thing
without claiming they are one RDF resource. Each keeps its own types and properties, nothing is
inferred onto the other, and the link can be withdrawn without touching either source. It carries no
OWL entailment, so the everyday query path stays reasoner-free.
owl:sameAs is the documented exception. It asserts the two IRIs denote one resource, and a
reasoner may then merge everything known about both, in both directions. Use it where that merge is
the feature you want and identity is confirmed — reconciling an id after a rename is the clear case
(see formerIds: in the diagram index). A wrong owl:sameAs silently corrupts
every query downstream of it, which is why it is not the default.
Choose with the substitutability test. Could a reasoner safely replace one IRI with the other
everywhere, forever? If yes, the nodes are one entity → owl:sameAs. If not → skos:exactMatch.
Most cross-model links fail the test, because notations describe the same system at different
granularities and for different concerns: a diagram's mention of an application is not the
application's record in a catalog, even when both are accurate.
For a correspondence that is real but not exact, skos:closeMatch and skos:relatedMatch say so.
Note relatedMatch, not skos:related — the *Match terms are the cross-scheme family, and
skos:related is defined for use within one concept scheme.
<!-- BPMN: the predicate is the element name -->
<skos:exactMatch rdf:resource="archi:id-abc123-GeoMapping"/>
Whichever route, the converter transcribes the predicate you wrote and makes no judgement about it. Getting the choice backwards is how a graph ends up merging two resources that were only ever meant to correspond, and nothing in the toolchain will warn you.
Instance-level and type-level¶
Everything above relates instances: this task, that component. Relating one language's
types to another's — bpmn:UserTask to am:BusinessProcess, c4:Container to
am:ApplicationComponent — is a different job, and extension data is the wrong place for it.
A type-level correspondence is a statement about the two languages, true of every model written in
them, so writing it inside one model file would scope a general claim to an accident of where it was
typed. It belongs in a *-crossmappings.ttl document registered on the metamodel via
arch:crossLanguageMappings, authored once per language pair and reviewed as an editorial
judgement. The predicates are the same SKOS family, with skos:closeMatch and skos:relatedMatch
doing most of the work.
Neither level is ever emitted by a converter. Mappings are judgements, not derivations — see DD-11 for the policy and Authoring cross-language mappings for how to write the type-level document.
Referring to an element¶
This section is about element-level statements only. A view- or model-level statement names no subject, so there is nothing to look up and nothing that a rename can leave behind — which is the practical difference between the two, beyond what they are about.
The in-file BPMN route is the only one that contains its statements: extensionElements is nested
inside the element, so there is nothing to name. Every other route refers to an element, and
that has a consequence worth knowing — a reference can be left behind by a rename, so every route
reports an annotation that matches no element rather than dropping it:
[WARN] orders.puml:12: no element 'Nowhere' in this model, so <am:realizes> was not attached to
anything. Check the name against the diagram.
A PlantUML element has two names, and either resolves:
class "Payment Gateway" as PayGw |
||
|---|---|---|
| code | PayGw |
what the source refers to it by — in every arrow, and in the IRI |
| display name | Payment Gateway |
what the picture shows, and its skos:prefLabel |
The code is the one to write. It is what the rest of the file already uses, so an annotation reads
like the diagram around it, and it is the name the IRI is built from. The display name resolves too,
because it is what a reader has in front of them; quote it when it contains a space:
'!la-link "Payment Gateway" arch:refines kg:CAP-Payments. A name matching two elements is refused
rather than guessed at.
Prefer the code, because a label gets edited
A display name is prose and changes for the reader's benefit. An annotation keyed on it has to be revisited when someone retitles the box; one keyed on the code does not, and neither does the element's IRI — see ADR 0010.
Where an element has no as clause the two names are the same string, so there is nothing to
choose and nothing to maintain.
Reference¶
- BPMN → Extension data — the
extensionElementsroute - PlantUML → Extension data — the comment route
- Diagram index → Element entries — the index route for elements
- Diagram index → Statements about a view or a model — the index route for the other two subjects
- ADR 0010 → What names a PlantUML element — why the code, and not the display name, is what an annotation and an IRI are keyed on