Changelog¶
Unreleased¶
Fixed — --emit-direct-rel-triples did nothing unless a --type-mapping declared every predicate¶
The flag is documented as writing the {source} {predicate} {target} shortcut alongside the
qualified relationship resource, plus the rdf:reifies bridge between them. In ArchiMate and
PlantUML the predicate was looked up in exactly one place — the predicates: table of a
--type-mapping file. ConvertCommand defaults that to TypeMapping.EMPTY, so an ordinary run
resolved nothing, skipped the branch that writes both triples, and exited successfully having said
nothing. An operator who passed the flag to make traversal queries work got an empty result at query
time and no route back to the cause.
Nothing caught it because every existing test in those modules pinned the flag to false.
The predicate is now a published fact, read where the vocabulary declares one.
arch:unqualifiedForm states it on eleven ArchiMate relationship classes and fourteen UML ones —
am:Serving arch:unqualifiedForm am:serves, uml:Association arch:unqualifiedForm
uml:associatedWith — so the flag alone is enough. This is what LeanIX, Backstage and Structurizr
already did, which is why the same flag worked there unaided. The pairs are held as a table in each
converter rather than loaded at conversion time, so a conversion still needs no network.
Resolution is per relationship, not per run, which fixes a second defect in the same branch: a
predicates: entry for one type used to silence every other type. Mapping Serving alone now
overrides Serving and leaves the other ten on their published form. A declared predicate still
replaces the published one rather than joining it, so one edge is never stated twice.
PlantUML resolves through the resolved class rather than the arrow name, because the two are not
one-to-one: Uses is uml:Usage. uml:Message and uml:Transition publish no unqualified form, so
a sequence arrow and a state transition get no direct triple — the relationship resource still
carries the statement.
Fixed — BPMN accepted the same flag and silently produced nothing¶
The BPMN ontology declares no arch:unqualifiedForm on any relationship, so there is nothing to fall
back to and the no-op is the correct behaviour. What was wrong was the silence.
The run now reports it once, naming the flag, the type keys that reached no predicate, and the remedy
(predicates: in a --type-mapping file). Once per run rather than once per relationship: two
unmapped sequence flows are one thing to fix. It goes through the same report sink as the
extension-mapping warnings, and states that the relationship resources are unaffected — arch:source
and arch:target still carry every statement.
Fixed — every LeanIX fact sheet subtype was discarded, and the nine published classes for them could not be reached¶
LeanIX does not model a subtype as a fact sheet type. A Business Application is an Application fact
sheet whose category is businessApplication; the API's own subtype filter is a category facet
beside the FactSheetTypes one, and type always reports the parent. The pull was built on the
opposite assumption — FactSheetMapper preferred the node's own type "since a workspace can return a
subtype", which it does not — so category was never requested except on ITComponent, and there it was
requested as an ordinary attribute.
Nothing warned. No field was asked for, so no request failed. Five things followed:
| Before | |
|---|---|
FACT_SHEET_CLASSES |
7 of its 20 entries — BusinessApplication, Deployment, Microservice, Process, Product, ValueStream, OrganizationUnit — were unreachable, because elementClass() keys on sheet.type |
type-mapping-leanix.yml |
6 of its 14 elements: keys were never consulted, for the same reason |
leanIX-deliverable-templates.ttl |
3 of the 4 classes its application-portfolio query matches on could not appear |
ITComponent |
its own subtype reached the graph as vocab:category "software" — a string in the author's private namespace, visible in the shipped out/leanix.ttl |
| everything else | a Microservice and a Deployment were indistinguishable from a logical Application |
The pull now asks. subType: on a type in the pull configuration names the field the subtype rides on,
conventionally category. It is per type and not a base field because category is not on
BaseFactSheet — the same trap level sets, where one over-eager base-level entry fails the whole query
for every type that lacks the field. leanix-pull introspect reads the real schema and writes subType:
wherever the field is actually there, which is the reliable route: Application's three subtypes are
optional upstream and off by default, so whether it has the field at all is a per-workspace fact. The
built-in defaults claim it for ITComponent alone, the one type whose subtypes are in the standard
metamodel. The canonical export carries it as subType on the record, lifted out of fields because a
subtype is a statement of kind rather than a per-tenant attribute.
A subtype now reaches a class, alongside its parent. rdf:type lmm:FactSheet, lmm:Application,
lmm:BusinessApplication. Both, not just the narrower one: the ontology's rdfs:subClassOf axiom entails
the parent anyway, so asserting it means a query for every lmm:Application still finds the fact sheet
with no reasoner in the loop — the difference between adding the distinction and breaking every query that
predates it. Five routes, in order:
| Route | How | Result |
|---|---|---|
| Qualified mapping | elements: keyed ITComponent/software |
your class. Prefer this — a bare token is not unique across types |
| Bare mapping | elements: keyed software or Software |
your class, under any type |
| Published class | nothing to do | lmm:BusinessApplication and the other eight |
| Minted class | --ns-vocab |
vocab:Software, guarded by the same local-name check Backstage applies to an unpublished kind |
| None | — | parent type only, reported once with all three remedies |
Nothing is minted in lmm:, the rule every other term here follows. The raw token is kept on
vocab:subType under --ns-vocab whether or not a class was found, for the reason dct:type keeps the
type token: the class is this converter's reading of it, the literal is what the workspace said.
Deliberately not a second dct:type — two values under one predicate with nothing to tell them apart is
the collision dct:identifier was introduced to avoid — and not schema:additionalType, which schema.org
declares rdfs:subPropertyOf rdf:type, so a literal there would entail a literal in rdf:type position.
mapping-report gained a Fact sheet subtypes section, keyed Type/subtype, so an unmapped subtype is
answerable in CI rather than after the fact. It also says when an export carries no subtype at all: a
workspace with none and a pull that never asked look identical from there, and the second is the likelier.
Fixed — a relationship could not be reached from the element it left, and arch:relPredicate was never published¶
Core describes four representations for every relationship: the resource, a qualified predicate
from the source element to it, an optional direct triple between the endpoints, and rdf:reifies
bridging the direct triple back to the resource. Only the LeanIX converter emitted all four. Each of
the others had settled the question for itself as it was added, and the result had drifted:
| qualified predicate | rdf:reifies |
arch:relPredicate |
|
|---|---|---|---|
| ArchiMate | behind --emit-qualified-rel-triples, off by default |
yes | — |
| LeanIX | yes | yes | — |
| BPMN | only with a --type-mapping |
no | yes |
| PlantUML | no | no | yes, with a mapping |
| Structurizr | no | no | — |
| Backstage | no | no | — |
1,049 of the 1,074 relationship resources across the shipped playground outputs could not be
reached from either endpoint — 948 from ArchiMate with the flag off, 42 from Backstage, 34 from
PlantUML, 21 from BPMN run without a type mapping, 4 from Structurizr. arch:source points out of a
relationship, so with nothing pointing in, a resource — and its class, its labels, its lifecycle, its
provenance — was reachable only by scanning every arch:source in the graph.
Two things are fixed.
arch:relPredicate is gone. Core never published it. The skos:historyNote on
QualifiedRelationship records the rdf:Statement design behind relSource / relTarget /
relPredicate being dropped, and the metamodel linter flags the name — so BPMN and PlantUML were
writing a term no consumer could resolve. In its place is the RDF 1.2 bridge the ontology does
specify, rdf:reifies carrying a triple term, which BPMN, PlantUML, Structurizr and Backstage now
emit under --emit-direct-rel-triples as ArchiMate and LeanIX already did.
The qualified predicate is emitted by every converter, unconditionally. From
qualifiedPredicates: where a mapping names one; from the notation's published term where there is
one — Structurizr gained c4:qualifiedUses and structurizr:qualifiedDeployedOn, Backstage the
twelve bs:qualified*; and from arch:hasQualifiedRelationship otherwise, which core names as the
fallback for exactly this case.
ArchiMate's --emit-qualified-rel-triples is the one behaviour change to an existing flag. Gating it
left ArchiMate the only converter whose relationships were unreachable by default, so the triples are
now always emitted and the flag is accepted and ignored rather than removed. Existing command lines
keep working, and gain triples rather than losing them.
Fixed — retyping a LeanIX relation put the replaced class back by inference¶
A relationships: entry replaces a class. Once relApplicationToITComponent was retyped from
lmm:Requiring to am:Serving, the converter still reached for the published
lmm:qualifiedRequires — whose rdfs:range names lmm:Requiring. Because rdfs:range is an
entailment rather than a constraint, a reasoner would infer lmm:Requiring straight back onto the
resource and quietly undo the retyping. Four relationships in the shipped ArchiMate mapping were
affected.
A published predicate is now used only while the resource still carries the published class it was
declared against. Where relationships: has retyped it, qualifiedPredicates: must supply the
predicate; failing that the relationship carries arch:hasQualifiedRelationship, whose range always
holds, and the run warns naming the key to add. The direct triple is dropped rather than substituted —
there is no generic core predicate to stand in for one, and the resource carries the statement anyway.
Found while fixing that: predicates: and qualifiedPredicates: were only ever matched against the
relation field names an export declares, never against the canonical name the converter resolves,
although relationships: matched both and the mapping file documents both. So a mapping keyed
Requiring retyped an edge while its predicate entries were silently ignored — producing the very
mismatch above. All three sections now resolve the same way.
playground/config/type-mapping-leanix.yml grew from three predicates: and two
qualifiedPredicates: to fifteen of each, one per relationships: entry.
Fixed — an ArchiMate 2.1 exchange file lost its view links and its whole folder tree¶
archimate2linkedarchi read only the ArchiMate 3.x spelling of the Open Exchange format's
cross-references. A 2.1 document — including Archisurance.xml, the canonical example The Open Group
publishes — silently lost three things:
- Every view node's element reference. Each node reached the emitter with no
elementRef, so it was typedarch-vis:Noderather thanarch-vis:ArchNode, carried noarchvis:archElement, and got noarch:inView. Every node in every view became layout-only. - Every view connection's relationship reference, with the same effect on
archvis:archRelationship. - The entire organization tree. Folder membership is carried by
identifierRefon<item>, and the containing element is named<organization>in 2.1 against<organizations>in 3.x. Both were missed, so a 2.1 model produced noarch:Folderand noschema:itemListElementat all.
Two renames between format versions are responsible: the reference attributes went from lowercase
(elementref, relationshipref, identifierref) to camelCase, and the organization element gained
its plural. Both spellings are current in the wild.
Nothing reported this. Every affected attribute is optional in the schema, so a document whose references could not be read was indistinguishable from one that genuinely had none — no warning, no count discrepancy, and a conversion that exited 0.
What ArchiSurance produces now, against what it produced before:
| before | after | |
|---|---|---|
arch-vis:ArchNode |
0 | 222 |
arch-vis:Node (layout-only) |
237 | 15 |
archvis:archElement |
0 | 222 |
arch-vis:Link |
0 | 199 |
archvis:archRelationship |
0 | 199 |
arch:inView |
0 | 419 |
arch:Folder |
0 | 10 |
schema:itemListElement |
0 | 303 |
| total triples | 8,422 | 10,817 |
What to do. Re-convert any ArchiMate model whose source declares archimate_v2p1.xsd, or whose
output has view nodes typed arch-vis:Node with no archvis:archElement, or no folders. Output from
3.x documents is byte-for-byte unchanged — the camelCase spelling is still tried first, and the
regression tests assert the two spellings produce the same graph rather than asserting the 2.1 form
works on its own.
Consequences beyond the converter. Anything reading the diagram layer was blind on these models,
so core/view-contents and core/view-usage in linked-archi-apm returned nothing for them, and the
published ArchiMate viewpoint conformance shapes had no arch:inView to judge. Those answers change
without any query changing.
Changed — breaking: membership is arch:inModel, and view membership is now stated semantically¶
Two changes to the same part of the vocabulary, following core 0.4.0 and DD-30.
arch:partOfModel is now arch:inModel. A pure rename, same subject, same object, same graph. The
old property is retained upstream as owl:deprecated and declared rdfs:subPropertyOf the new one, so a
graph written by an earlier converter still entails the current vocabulary for a consumer running RDFS.
Nothing in this pipeline runs one, so for output this project generates the change is a cutover: a
consumer holding arch:partOfModel re-resolves to arch:inModel or re-converts.
arch:inView is emitted, beside the archvis: node layer rather than instead of it. Every converter
that emits views now also states, on the concept itself and in the semantic graph, which views present
it — elements and qualified relationships alike, one deduplicated triple per concept-view pair.
The node layer is unchanged and still says more: an archvis:ArchNode distinguishes two drawings of one
element and carries geometry where the notation has any. Three reasons the semantic statement is needed
as well:
- Serialization profiles drop the node layer. It lives in
graph/views, which--views-profile no-viewsandno-diagramsremove wholesale. Under those profiles thearch:Viewresources survived ingraph/semanticwhile every trace of what was on them disappeared. For an ArchiMate model the views graph is around 80% of the output, so publishing without it is normal. - Some notations have no geometry to lift. PlantUML computes layout at render time; Structurizr and LeanIX name view membership without authored coordinates.
- The published ArchiMate viewpoint conformance shapes read this property, and until now nothing
emitted it — so
--shapes archimate3-viewpoint-shapesreported conformance for a view drawing elements its viewpoint'sarch:includesConceptpalette forbids. The SPARQL constraint matched nothing and a clean report meant only that. Those shapes now fire.
A relationship carries arch:inView directly. It is already a resource with arch:source and
arch:target, and arch:inView has domain arch:ModelConcept, so no reification or RDF-star annotation
is involved — the hedge to that effect in the retired arch:exposedInView definition was a leftover from
the rdf:Statement design those two properties replaced.
Gating. Both properties declare rdfs:domain arch:ModelConcept, so both are emitted only where the
run actually types the subject that way — --emit-core-triples in ArchiMate, --dual-typing in BPMN.
Asserting either on an untyped subject would be a domain violation, which is also why a LeanIX view
member absent from the inventory export, an ArchiMate node whose elementRef names an element that was
not emitted, and a BPMN shape depicting a document container all get the archvis: node and no
arch:inView. Layout-only nodes — ArchiMate groups and frames, typed archvis:Node rather than
ArchNode — get neither, there being no concept to make the subject.
arch:contains is deprecated upstream and was never emitted here, so nothing in converter output
changes on its account. It ran Model → View, which arch:inModel already said in the other direction.
Changed — breaking: a PlantUML element is addressed by the name the source refers to it by¶
A PlantUML entity has two names. class "Payment Gateway" as PayGw displays Payment Gateway and is
referred to as PayGw — in every arrow, and in every annotation. PlantUML calls the second one the
entity's code, and the element ID is now taken from it: this declaration mints
…/element/PayGw, with Payment Gateway remaining its skos:prefLabel.
Where an entity has no as clause nothing moves. PlantUML sets the code to the display text in that
case, so class OrderService and participant "Book API" mint …/element/OrderService and
…/element/Book_API as before. Only aliased entities are re-addressed.
Why the code. Two properties an identifier needs, and the display name has neither.
- The code is unique, and PlantUML enforces it. An arrow endpoint is resolved by code, so two
entities cannot share one. Nothing constrains the display text, so
class "Order Service" as Primaryandclass "Order Service" as Replicaare two entities showing the same words — keyed on the display name they minted one IRI, merging two resources and turningPrimary --> Replicainto a self-loop whosearch:sourceandarch:targetwere equal. - The code is not edited for presentation. Retitling a box to
"Payment Gateway (EU)"now changes oneskos:prefLabeland nothing else. It previously relocated the element IRI, every relationship composed from it, every view node and every SVGhref.
Full reasoning, with worked examples and the annotation forms, in ADR 0010.
What moves. For aliased entities only: the element IRI; the relationship IRIs and skos:notation
composed from its endpoints (ADR 0005);
archvis view node and link IRIs; and the href values render --base-iri writes into an SVG.
What to do. A consumer holding …/element/Payment_Gateway for an aliased shape re-resolves to
…/element/PayGw. No owl:sameAs bridge is emitted: an alias derived from a rule rather than declared
by an author would claim every such pair was a rename, which is false for a shape whose old IRI was
never published or denoted two entities at once. IdentityLock records model and view IDs only, so
there is nothing to detect or authorise an element-level change — the migration is a republish.
Annotations are unaffected, and gain a recommendation. Both names still resolve, so
'!la-link PayGw … and '!la-link "Payment Gateway" … reach the same element. Prefer the code: it is
what the rest of the source uses, and it does not need revisiting when someone edits the label.
- A duplicate element id is now reported against the file that contains it. PlantUML keeps codes
unique within a diagram, so the remaining case is a nested
namespace a.b, which is reported as one group per segment —a.sharedandc.sharedboth yieldshared. The run warns, naming both display names, rather than leaving it to the cross-view collision check, which aborts the run with a message about views disagreeing. - An over-long element segment now comes from a long code. An author who names a participant after a URL and gives it no alias has that URL as its code. Relationship segments are bounded; element segments are not, so a short alias is the fix — and it shortens the relationship segments too.
1.3.0 — 2026-08-30¶
Added — NQUADS, the line-based format that keeps named graphs¶
--format NQUADS, or any --output ending in .nq. Every converter, and validate reads it too.
Why it was missing. NTRIPLES is the format for streaming, grep and bulk loading, and it silently
drops the graph a statement is in. That was a fair trade while a model's graphs were a filing
convenience. It is not one now: with the semantic graph partitioned per input, the graph is the record
of which input asserted a fact, so flattening loses the attribution and not just the layout. N-Quads is
N-Triples plus that fourth term — same one-statement-per-line shape, same absence of prefixes.
{base}bpmn/demo/element/Task_1 rdf:type arch:Element {base}bpmn/demo/graph/semantic/group-models/a-bpmn .
Details:
OutputFormatnow records which formats keep graphs, askeepsNamedGraphs. Preserving:TRIG,JSONLD,NQUADS. Flattening:TURTLE,RDFXML,NTRIPLES. It was previously stated only in prose, in three docs and a code comment, with nothing to keep them honest. A test now asserts the flag against a real write-and-reparse for all six, so a format cannot claim to preserve graphs without doing it.validaterecognises.nqand--data-format NQUADS. The read path resolves formats through its own extension table, which defaults to Turtle for anything unknown — so before this a.nqfile written byconvertwas parsed as Turtle and failed withExpected '.', found '<'.rdf4j-rio-nquadsis a new dependency incoreand inconverter-archimate, which registers its writer factories programmatically rather than through the service files. N-Quads ships as a separate artifact from N-Triples despite being the same grammar, and the writer is registered inMETA-INF/servicesalongside the others for the reason recorded underUnsupportedRDFormatExceptionbelow: shading discards RDF4J'smodule-infodeclarations, so a writer that is on the classpath but not in a service file fails only at write time.- All the usual routes work:
--format NQUADS,--format nq,-o out.nq, and-o out.dat:nq. Naming a.nqfile while asking for a flattening format still warns, as it does for any other mismatch.
Reference: Output serialization for the format table and when to choose which.
Changed — breaking: the curated model moved into its own named graph¶
Previous behaviour. The arch:Model resource, its metamodel conformance, its folders and their
ordering all sat in graph/semantic, beside the concepts lifted out of the input. Measured on a
two-descriptor catalog, that was 27 of 60 triples — nearly half a semantic graph attributable to no
input at all, since none of it comes from a diagram or a descriptor.
They now live in {base}{notation}/{modelId}/graph/model, in all six converters.
{base}{notation}/{modelId}/graph/semantic facts lifted from an input
{base}{notation}/{modelId}/graph/model the curated model: arch:Model, folders, ordering
{base}{notation}/{modelId}/graph/views geometry
{base}{notation}/{modelId}/graph/provenance PROV-O about all of the above
Details:
- Folder membership now travels with the folders.
<concept> dct:isPartOf <folder>used to be written inline beside each concept's type triples while the orderedschema:itemListElemententry was written with the folder — one relation in two places, and oncesemanticwas split per input, in two graphs. Both halves are ingraph/modelnow. graph/model's input is the diagram index, so it is described in provenance like any other graph.- Separating it is what makes attribution possible at all: curated content has no source to be attributed to, and while it shared the graph a query for "what did this input produce" could not exclude it.
A consumer reading graph/semantic for the model resource or the folder tree must read graph/model
too, or query across the union. A Turtle consumer is unaffected — the graphs are flattened either way.
Reference: Output serialization for the layout and the boundary rule, Architecture overview for what each graph holds, IRI path shape for the named-graph IRIs.
Changed — breaking: a source file's path moved to schema:name; dct:source left the source entity; rdfs:seeAlso became schema:url¶
Previous behaviour. The provenance source entity carried the file's path as dct:source, a literal —
or a bare filename where no repository was known — and linked the upstream blob twice, under both
prov:alternateOf and rdfs:seeAlso.
A bare dct:source "catalog-info.yaml" identifies nothing: it is the same string for every descriptor in a
catalog pulled from many repositories, and a consumer holding it cannot get back to the file. The second
blob link was worse than redundant — rdfs:seeAlso asserts no relationship between two IRIs, so beside
prov:alternateOf, which says they are the same document, it stated nothing.
Now:
<…/backstage/service-catalog>
dct:source <https://git.example.org/group/order-service/-/blob/4f2c1ab8e0…/catalog-info.yaml> ;
prov:wasDerivedFrom <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> .
<…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml>
a prov:Entity ;
schema:name "catalog-info.yaml" ;
dct:isPartOf <https://git.example.org/group/order-service> ;
dct:identifier "4f2c1ab8e0d1c2b3a4958677889900aabbccddee" ;
prov:alternateOf <https://git.example.org/group/order-service/-/blob/4f2c1ab8e0…/catalog-info.yaml> .
Three predicates, three jobs, none of them interchangeable:
schema:namecarries the identifying path, always a literal, always present — the valuedct:sourceused to hold. Recovering it from the blob URL would mean parsing a forge URL, which ADR 0008 §3 forbids.dct:sourceis gone from the source entity, and is not simply repointed at the blob. That node is the file at that commit —prov:alternateOfsays so,schema:sha256digests its bytes and the commit activity generated it — so adct:sourceto the same blob would make the file its own source. The derivation holds of the output, whereprov:wasDerivedFromalready states it.dct:sourceon the model or view — which is where it belongs, since an output really is derived from its input — now carries the commit-pinned blob IRI where the run can build one, and the bare filename where it cannot, which is every run with provenance off. So accept either an IRI or a literal there;FILTER(isIRI(?o))separates them. DCMI sanctions both: rangerdfs:Resource, URI preferred.rdfs:seeAlsois gone from the graph entirely. It was dropped from the source entity, whereprov:alternateOfalready carried the link, and replaced byschema:urlon the two nodes where the URL is the only pointer to a resource: the commit activity (the forge's page for that commit) and the converter agent (the page for the revision it was built from).seeAlsosays only "here is related information";schema:urlsays the object is the URL of the subject, which is what was meant in both places.
What to change.
- ?g prov:wasDerivedFrom ?src . ?src dct:source ?path .
+ ?g prov:wasDerivedFrom ?src . ?src schema:name ?path .
Enumerating source files was documented as filtering on dct:source. It is now:
The type alone still over-matches — the upstream blob is a prov:Entity carrying nothing but its type —
and the repository and commit author carry schema:name without being entities, so both patterns are
needed. A "view source" link reads prov:alternateOf on the source entity, as before.
In an aggregate, pin dct:isPartOf beside schema:name: paths repeat across repositories.
Reference: ADR 0008 §5 for the three predicates side by side, and Generator spec for the query patterns.
Changed — breaking: one Git author keeps one person IRI across display-name changes¶
Previous behaviour. A commit author was minted as {base}provenance/person/{digest-of-name} and only
schema:name was published. Git does not treat a display name as identity: one person can commit as
J. Doe, Jane Doe and Jane A. Doe, and the converter minted three unrelated people. Conversely, two
people sharing a display name collapsed onto one node.
Now the canonical address identifies and the name describes:
<…/provenance/person/a.author%40example.org>
a prov:Person ;
schema:name "A. Author" ;
schema:email "a.author@example.org" .
- The IRI key is the trimmed, lower-cased email address, percent-encoded as one reversible path segment.
Git records no username or numeric person id; the address is its only identity handle. Canonicalising
makes
Author@Example.organdauthor@example.orgone person; reversible encoding keeps distinct valid addresses such asa+b@example.organda-b@example.orgdistinct. - The canonical address is readable rather than hashed because it is published as
schema:emailon the same node; hashing the IRI would conceal nothing and would make it harder to inspect. These graphs are internal, and the suite already makes the same decision for LeanIX stakeholders and Backstage profiles. - Where no address is available, the old digest-of-name IRI remains as a compatibility fallback. Existing
source maps remain readable: a map with
authorNamebut noauthorEmailuses that fallback; a map with neither field emits no author person until refreshed. - Direct CI conversion reads both values from
CI_COMMIT_AUTHOR. Local conversion reads%anand%ae.backstage-pullnow enriches each fetched file from GitLab's commit endpoint and writescommittedAt,authorNameandauthorEmailintocatalog-sources.yaml. schema:emailmatches the predicate LeanIX already uses onarch:Stakeholder, so the same human can be joined across those notations. Backstagebs:Userremains model-scoped and usesbs:email; reconciling it to the provenance person is a separate cross-model identity decision.
Pipeline impact of pulled commit metadata. After a successful file request, backstage-pull makes one
additional GET /api/v4/projects/{project}/repository/commits/{last_commit_id} per distinct project and
commit, cached across descriptors. It sends the same PRIVATE-TOKEN; verify that the token can read the
commit endpoint. HTTP 429 and 5xx responses retry up to four attempts with 2s/4s/8s backoff. Other metadata
failures, including 401/403, warn and continue with committedAt, authorName and authorEmail omitted;
file-fetch authorization failures remain fatal. Without either author field, conversion emits no
prov:Person or author link for that commit.
backstageCatalogSources remains format version "1". authorEmail is a new optional entry key, and a
refreshed map may also newly populate the existing optional committedAt and authorName keys. Old maps
remain valid; update strict external validators or deserializers that reject unknown keys before committing
a refreshed map.
Downstream action. Do not persist or reconstruct the old digest-of-name person IRI. Resolve authors
through the commit activity, or join on the canonical schema:email:
Refreshing a source map and reconverting changes person IRIs where author email becomes available; the commit, source, model, element and relationship IRIs do not change.
Reference: ADR 0008 §5 for the identity rule and fallback.
Added — a fact says which input it came from¶
Previous behaviour. Derivation was asserted per resource. Where two inputs contributed to one
resource, that resource carried two prov:wasDerivedFrom values and nothing paired a value with an
input:
<…/element/component/default/order-service>
bs:lifecycleState bs:Production , bs:Experimental ; # one from each descriptor
skos:prefLabel "Order Service" , "order-service" .
Both values present, both descriptors named, the pairing unrecoverable.
backstage2linkedarchi now puts each input's facts in a named graph of their own, named after that
input — its repository path where the run knows one, else its file name:
{base}backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-payments/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-models/catalog-index-yaml
"Everything this input produced" is then one GRAPH clause, and a fact two inputs both assert is
attributable to each of them — two descriptors declaring the same edge mint one relationship resource,
written into both of their graphs.
Details:
- Only where a model has more than one input. A single-input conversion keeps the bare
graph/semantic: partitioning one input would name a graph after the only file there is. There is no flag — the rule follows from the inputs. Note that a converter is called once per input file, so the count comes from the command; the index-driven, file-list and directory modes all produce the same layout for the same catalog. - What the index declares gets a graph too.
arch:architectureStateand the index'selements:/links:/data:assertions are lifted from the index, so they sit in the graph named after it. The index is an input like any other; it is simply the one whose facts are about the model. - A graph is named after the file, not the file at a commit. The
prov:Entitydescribing an input keeps its sha — it identifies bytes at a revision, and accumulating one node per revision is the history. A graph holds the current facts from a file and a later run replaces it, so keying it on the commit would make a re-pull write a second graph beside the first, and a union would return that file's facts once per revision it had ever been read at. - A run that cannot tell two inputs apart says so. A graph falls back to the file name where no
repository path is known, and every Backstage descriptor is called
catalog-info.yaml— so a catalog pulled from forty repositories with neither--git-provenance autonor--source-mapputs all forty in one graph. Nothing false is published (the graph is derived from every file that fed it) but the attribution is no better than before the split, and the run now names the option that fixes it. - The other five converters still write one
graph/semanticper model. The core helpers are notation-agnostic; what each one needs is per-concept origin tracking, which only Backstage has today. - The per-concept
prov:wasDerivedFromis unchanged and still emitted. It is no longer the only route, though, so it is now recoverable rather than load-bearing — see the CONSTRUCT below.
Reference: Output serialization → One semantic graph per input, Backstage → One semantic graph per descriptor, and ADR 0008 decision 8.
Your store's default graph must be a union of the named graphs
With the semantic graph split, a query with no GRAPH clause no longer sees a whole model — and it
returns zero rows rather than failing, which is the failure mode most easily mistaken for empty
data. Fuseki: tdb2:unionDefaultGraph true. rdflib: Dataset(default_union=True). See
Aggregating into a knowledge graph.
Added — every named graph says what generated it and from what¶
Each graph a conversion writes is described in that model's provenance graph, in all six converters:
<…/graph/semantic/group-orders/catalog-info-yaml>
a prov:Bundle ;
prov:wasGeneratedBy <{base}provenance/run/9c2f1e04> ;
prov:wasDerivedFrom <{base}provenance/source/group-orders/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:generatedAtTime "2026-08-29T09:03:11Z"^^xsd:dateTime .
Details:
prov:Bundleand not alsoprov:Entity. PROV-O makes Bundle a subclass of Entity, so the second type adds nothing a reasoner needs and it puts every graph into?s a prov:Entity— the query that enumerates a run's inputs. The repository node is typedprov:Collectionalone for the same reason.- The diagram index is located and described for the first time. It is the input of
graph/model, and nothing previously said the run had read it at all.ProvenanceOptions.resolvenow takes the index file, and the run carries it asprov:used. - A graph the run wrote but left empty is not described. RDF has no way to state that a graph exists but is empty, so describing one would publish a bundle a consumer cannot find.
- LeanIX's views graph is deliberately left undescribed. Its content comes from
--diagrams-export, a second file the diagram parser hands over by name only, so there is no path to place it in a repository with. Deriving that graph from the fact sheet export instead would name the wrong input. -
Because the graph carries the derivation, the per-concept form is recoverable:
CONSTRUCT { ?c prov:wasDerivedFrom ?src } WHERE { ?g a prov:Bundle ; prov:wasDerivedFrom ?src . GRAPH ?g { ?c a arch:ModelConcept } }Verified against real output: it recovers every derivation the Backstage converter emits, including the cross-file one.
Reference: Output serialization → Every graph says where it came from,
and the generator specification for how a downstream generator should reach
a graph — through ?g prov:wasDerivedFrom ?src, never by composing the IRI.
Added — every concept says which model it belongs to¶
Previous behaviour. Nothing connected a concept to its model. Tracing forward links from an element
reached its folder and stopped; the model resource appeared only as <M> a arch:Model and
<M> arch:modelConformsToMetamodel …. Membership was carried entirely by the named graph, so it was
lost under TURTLE, RDFXML and NTRIPLES, and in a flattened merged-graph.ttl — all supported
outputs. A consumer could recover it only by parsing the IRI, which
ADR 0008 §3 tells consumers
not to rely on.
Every element, relationship and view now carries
arch:partOfModel to its model, in all six converters.
<…/backstage/service-catalog/element/component/default/order-service>
a bs:Component, arch:Element, arch:ModelConcept ;
arch:partOfModel <…/backstage/service-catalog> .
Details:
- The term is published, not minted here.
arch:partOfModelhas domainarch:ModelConceptand rangearch:Model, andElement,QualifiedRelationshipandVieware allModelConceptsubclasses, so one property covers every concept a converter emits. - Not the folder chain. Completing
dct:isPartOffrom a concept through its folder to the model would have cost one triple, but it makes membership depend on folder policy, and a folder is a presentation device — nearer a view than the model's semantics. - Emitted exactly where the concept is typed as a
ModelConcept, which means it follows--emit-core-triplesin ArchiMate and dual typing in BPMN. Those flags decide whetherarch:core types are emitted at all, and asserting a property whose domain isarch:ModelConcepton a subject the run has declined to type that way would be a domain violation. BPMN's document container is excluded for the same reason it is excluded fromarch:Element. - Additive: no IRI changes, no triple removed. A consumer that inferred membership from the graph or the IRI keeps working.
This is the first step of todo/PROPOSAL-source-scoped-graphs.md, and the prerequisite for the rest:
once the semantic graph is partitioned per input, the named graph can no longer imply membership, so
membership has to be a statement first.
Reference: Identifiers for a worked example, and Aggregating into a knowledge graph for the reassembly query. Prefer it over parsing an IRI or relying on a graph name: both are wrong for a partitioned model.
Fixed — a repository reached through a symlink lost its path, repository and blob URL¶
Previous behaviour. GitProvenance compared the input file against the repository root
lexically, with normalize(), which does not resolve symlinks. The two paths arrive from different
places — the root from CI_PROJECT_DIR or git rev-parse --show-toplevel, the file from the command
line — so they need not spell one directory the same way. Where they disagreed, the containment test
reported the file as outside its own repository and the run degraded to a name-only source: dct:source
became the bare filename, and dct:isPartOf, prov:alternateOf, rdfs:seeAlso and the repository
component of the source IRI were all dropped. --git-provenance had been asked for and answered with a
filename.
Reproducible on macOS, where /tmp is a symlink to /private/tmp: the same conversion under /tmp
and under /private/tmp produced different provenance from identical inputs.
Both paths are now resolved before the comparison, and all-or-nothing: toRealPath needs the path
to exist, so resolving whichever side happens to exist while normalising the other would produce two
spellings of one directory that no longer share a prefix — causing the failure rather than fixing it.
Where either side cannot be resolved, both are normalised, which is the previous behaviour exactly.
A genuinely unrelated directory stays unrelated after resolution, so the containment test has not become a rubber stamp; there is a test for each direction.
Changed — breaking: provenance nodes moved out of the per-model graph, into one namespace¶
Previous behaviour. Every provenance node was a fragment of the model's provenance graph:
…/{modelId}/graph/provenance#source-4f2c1ab8-catalog-info-yaml, #commit-4f2c1ab8, #author-1f3a9c02,
and a constant #run. ADR 0008 §2 chose that on the premise that provenance is not accumulated across
runs — which holds for one converter output read on its own, and not for the aggregating consumer these
graphs are published for. A provenance graph IRI is per model, so a file read by two models got two
unrelated nodes, one run converting three models got three activity nodes, and merging two runs'
descriptions of one model put two start times, two end times and two job URLs on the single #run.
A node identified by its content now lives in one base-scoped namespace; only the statements about it are per-model.
| Node | Before | After |
|---|---|---|
| source file | …/graph/provenance#source-{sha8}-{path} |
{base}provenance/source/{repo}/{sha12}/{path} |
| commit | …/graph/provenance#commit-{sha8} |
{base}provenance/commit/{sha12} |
| author | …/graph/provenance#author-{digest} |
{base}provenance/person/{encoded-email} (name-digest fallback) |
| run | …/graph/provenance#run |
{base}provenance/run/{token} |
| agent | …/graph/provenance#agent |
{base}provenance/agent/{digest} |
| derivation | — | {base}provenance/derivation/{repo}/{sha12}/{path}/{run} |
Details:
- One reserved top-level segment, not six.
{base}already holds notation slugs, so sources at{base}source/would reserve a word there and this class of node needs five more — leaving "isleanixa notation or a node kind?" unanswerable from the IRI. Spelledprovenance/rather thanprov/so it is not read as theprov:prefix. A notation may not now be calledprovenance. - The repository is in the source IRI, not only in its triples. Emitting
dct:isPartOfwithout it left a node that could carry two repositories when two descriptors shared a path and a short sha — a false statement rather than an ambiguous one. Keying the IRI on the repository makes it unavailable rather than merely unlikely. - The run is keyed on what distinguishes one run from another: the CI job URL, else the start time and image reference. Where a run carries none of those it is genuinely indistinguishable from another such run, and sharing the node loses nothing, since there is then nothing for the two descriptions to disagree about.
- Shas in an IRI are twelve characters, not eight. Eight was sized when the population at risk was
one model's inputs; these nodes are shared across an aggregate and the sha separates two revisions of
one file, so the bound became "every commit that ever touched this path". The full value is always
dct:identifier. - Statements still land in each model's own provenance graph. A model's description stays self-contained and a TriG consumer loses nothing. What changed is only that a node is no longer named by the graph that mentions it.
Migration. Every provenance node IRI changes, for all six converters. Nothing else does: the statements, their predicates and their literals are unchanged, and no triple is removed. ADR 0008 §3 declares these IRIs opaque and not required to parse or resolve, and a run rewrites its provenance graph in full, so a consumer that followed that contract re-reads and continues. A consumer that hard-coded a fragment does not.
Added — a derivation now says which run it went through¶
prov:wasDerivedFrom cannot name the activity a derivation passed through, and PROV does not let you
infer it — prov:used and prov:wasGeneratedBy follow from a derivation, not the reverse. A provenance
graph has never held only one activity (a 36-descriptor run emits one conversion activity plus one commit
activity per distinct commit), and after a merge across runs there are several candidates rather than one
plausible one.
<…/element/component/default/order-service>
prov:wasDerivedFrom <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:qualifiedDerivation <…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04> .
<…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04>
a prov:Derivation ;
prov:entity <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:hadActivity <…/provenance/run/9c2f1e04> .
- One node per (source, run), shared by every resource derived from that file. A
prov:Derivationstates which entity was used and which activity did it; it names no derived entity, so sharing is what the record says rather than an approximation. A node per resource would cost a 500-entity catalog 500 of them to say one thing 500 times. Cost as emitted: one triple per resource, three per source file. - An IRI, not a blank node. This is the load-bearing half. A blank node is relabelled on every parse, so two runs' descriptions of one derivation could never be recognised as one — which is the entire reason for wanting the qualified form.
- No
prov:hadUsageorprov:hadGeneration. Both together make this a precise-1 derivation under PROV-CONSTRAINTS, with ordering obligations.prov:hadGenerationis per-derived-entity by construction, so adopting it forfeits the shared node, and its only content would restate aprov:wasGeneratedBythe resource already carries;prov:hadUsageis shareable but says what{run} prov:used {source}says. - Emitted by
backstage2linkedarchifor every element and relationship, and bystructurizr2linkedarchifor a lifted ADR, both only under--git-provenanceand friends — the same rule as before, since the target is minted only when a run is described.
Fixed — a pulled descriptor's provenance never said which repository it came from¶
Previous behaviour. A source entity carried dct:source — the path inside its own repository — and
nothing else that identified the file. For a pulled Backstage catalog that path is catalog-info.yaml in
every repository, so a run over many repositories emitted source entities all reporting the same value,
distinguishable only by a truncated commit sha in an opaque IRI. The repository appeared only as a
substring of an rdfs:seeAlso URL, which asserts no relationship.
Four facts now, each on its own predicate.
<…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml>
a prov:Entity ;
dct:source "catalog-info.yaml" ; # where in the repository
dct:isPartOf <https://git.example.org/group/order-service> ; # which repository
dct:identifier "4f2c1ab8e0d1c2b3a4958677889900aabbccddee" ; # which revision
schema:sha256 "4c294617b607…" ; # which bytes
prov:alternateOf <https://git.example.org/group/order-service/-/blob/4f2c1ab8e0/catalog-info.yaml> ;
rdfs:seeAlso <https://git.example.org/group/order-service/-/blob/4f2c1ab8e0/catalog-info.yaml> .
<https://git.example.org/group/order-service>
a prov:Collection ;
schema:name "group/order-service" .
Details:
dct:isPartOf, not a repository-qualifieddct:source. A file is part of a repository, which is the same statement the term already makes about a view and its folder, anddct:sourcekeeps meaning exactly what ADR 0008 §5 says it means. The repository URL is the node IRI directly rather than a minted proxy — unlike a file at a commit, a repository has a stable published IRI that other graphs will cite, so a local proxy would fragment it for nothing. Dropped rather than published when the value is not anhttp(s)URL, since anscp-style remote would be a node nothing can dereference.prov:alternateOfis whatrdfs:seeAlsocould not say. The minted node and the blob URL are the same document at the same commit.seeAlsostays because it is what a human follows. Notowl:sameAs, which licences inferences this block does not intend; notprov:atLocation, which says where to look rather than what a thing is. Emitted only where ablobUrlexists, which is already gated on the commit tracking the path — so nothing is claimed to be the same document as a URL that 404s.schema:sha256readscontentSha256, which the source map has always collected and the converter ignored. Underschema:sha256and not a seconddct:identifier, which carries the commit: two identifiers on one subject is the ambiguity ADR 0008 exists to remove. A digest answers "which bytes" where the commit answers "which revision", and two commits sharing a digest are a descriptor that moved commit without changing — the case an incremental re-pull produces.- Also emitted for a committed input, from
CI_PROJECT_URL, so a model converted in CI and a pulled catalog describe their sources the same way and a consumer merging both needs no per-route knowledge. Silent outside CI, where there is no forge to name.
Enumerating source files needs a predicate, not a type. ?s a prov:Entity matches the upstream blob
as well, because prov:alternateOf has prov:Entity at both ends and an untyped alternate would be a
dangling node. Filter on what only a file carries:
The repository is typed prov:Collection and not also prov:Entity — PROV-O makes Collection a subclass,
so a reasoner sees both while that query returns files.
The repository is also part of the source entity's IRI, not only of its triples; see the namespace change
above. Raised downstream as todo/UPSTREAM-REQUEST-provenance-source-identity.md, with the reasoning and
the options considered in todo/PROPOSAL-provenance-source-identity.md.
Added — a catalog API dump's relations and status are read¶
Previous behaviour. Both root fields were ignored. That cost nothing for an authored descriptor, which
carries neither, and cost a lift of the catalog API two things: the relations a custom processor derived,
which no spec field states and which Backstage calls
the authoritative source,
and any indication of whether the entity had been ingested cleanly.
relations is read and normalised. The array carries both directions of every relation, so one
ownership arrives three times in a dump — ownedBy on the component, ownerOf on the group, and
spec.owner besides. Reading it as written would have trebled the edge and pointed one copy backwards.
The reverse member of each of the seven well-known pairs swaps its endpoints onto the canonical property
instead, so all three compose one relationship notation and merge onto a single IRI.
partOf / hasPart is the one pair that cannot be mapped by name: Backstage collapses four containments
onto it while the ontology publishes a class for each. The kinds at the ends decide, and a targetRef
always states its kind — Component → System is bs:SystemMembership, Component → Component is
bs:ComponentComposition, System → Domain is bs:DomainMembership, Domain → Domain is
bs:DomainHierarchy. Any other pairing is reported rather than guessed.
A relation type outside those fourteen is a house type from a custom processor. It has no published class,
so it follows the rule the rest of this release established: nothing is minted in bs:, and
relationships: / predicates: in --type-mapping name it.
status.items is read under a namespace you own. The ontology publishes no status vocabulary, so
with --ns-vocab each item becomes a node carrying statusType, statusLevel and statusMessage — a
node rather than flat properties, because a level and a message on the entity lose which belongs to which.
Without a namespace the items are reported and dropped. The nested error object is not read: its shape is
not fixed upstream and a stack trace is not an architectural fact.
Filed upstream. todo/UPSTREAM-REQUEST-backstage-kinds-and-status.md asks the bs: ontology for a
status vocabulary, for classes for Template and Location — or for an explicit statement that they are
out of scope, since silence and omission currently look the same — and for bs:label's intended pattern to
be made explicit.
Fixed — metadata.labels was dropped, and an unmodelled kind was invented in bs:¶
Previous behaviour. Two faults on the metadata side, both the same shape as the spec.dependencyOf
one below.
metadata.labels is documented on every kind and the ontology publishes
bs:label for it. The converter did not read it at all,
and said nothing, so a catalog with labels and one without produced identical output.
A kind outside the seven the ontology models — Template, Location, or a house kind — was typed
bs:{Kind}, fabricating bs:Template and bs:Location in a published namespace this converter reads and
does not own. That is the fault removed from the relationship path earlier in this release, still present
for kinds.
Labels are read, one predicate per key, by the same two routes an annotation with no published term takes:
metadata:
labels:
deployments.example.net/register-srv: "true" # prefix via namespaces:
tier: gold # bare key via --ns-vocab
<…/element/component/default/order-service>
<https://vocab.example.org/k8s#register-srv> "true" ;
<https://vocab.example.org/arch#tier> "gold" .
bs:label is deliberately not the emitted predicate: the ontology publishes it as an abstract
super-property, and a key/value pair placed on a super-property loses the key. Values stay plain strings,
since Backstage borrows Kubernetes' label semantics and states that both key and value are strings — so
unlike a spec literal, "true" is not read as a boolean.
An unmodelled kind is no longer invented. Its class comes from elements: in --type-mapping (keyed
on the lowercased kind) or from --ns-vocab. With neither, the entity is emitted as an arch:Element
with no notation class, keeping its labels, identity fields and every relation its spec states, and the
run reports the kind. bs:Template and bs:Location no longer appear in any output.
House metadata keys are declarable, closing the last asymmetry with spec: a new
metadata-literals: section, separate from spec-literals: because metadata.tier and spec.tier are
two statements and one declaration must not answer for both. Datatypes come from YAML as they do for a
spec literal. There is no relation counterpart, metadata being where Backstage puts descriptive fields.
An undeclared key is reported, pointing at metadata.labels as the usual home for a key/value classifier.
Fixed — spec.dependencyOf was dropped, losing documented dependencies¶
Previous behaviour. spec.dependencyOf is part of the descriptor format, on Component and
Resource, and this converter did not read it. A catalog stating its dependencies from the far end —
payment-db declaring dependencyOf: [component:default/payment-service] rather than the service
declaring dependsOn — converted with those edges missing entirely and nothing said about it.
Change. It is read and normalised onto the ontology's canonical direction, the same treatment
spec.children and spec.members already had: the listed entity becomes the source, the entity
declaring the field the target.
<…/relationship/dependsOn--component-default-payment-service--resource-default-payment-db>
a arch:QualifiedRelationship, arch:ModelConcept, bs:Dependency, bs:ResourceUsage ;
arch:source <…/element/component/default/payment-service> ; # the dependent
arch:target <…/element/resource/default/payment-db> . # the entity that declared it
bs:ResourceUsage still applies, because that subproperty is decided by the target's kind and after
the swap the target is the declaring Resource.
One fact stated from either end now converges. A dependsOn B and B dependencyOf A compose the same
relationship notation, so they merge onto one IRI rather than producing two edges pointing opposite ways —
the property that already made one PlantUML arrow drawn in two views a single relationship.
The default kind for a reference that omits one is Component here, against Resource for dependsOn,
because the two fields point at different populations: a dependency is often a resource, while the thing
depending on it is usually a component. Backstage requires the kind in both, so either default is a
leniency.
The remaining five unread documented fields are Template and Location fields — spec.parameters,
spec.steps, spec.target, spec.targets, spec.presence — kinds the published ontology does not
model. They are unmapped rather than blocked: declaring one under spec-literals: publishes it under a
namespace you own. The run now reports them as such instead of describing a documented Backstage field as
a house invention.
Added — a house spec field in a Backstage catalog can become a relationship¶
Previous behaviour. A spec key the descriptor format does not declare was dropped, and nothing
was logged. The parser picked named keys out of spec and never walked the remainder, so a catalog
carrying deployedTo or maintainedBy converted as though the field were not there — and a value that
never arrived looked exactly like a value the catalog never had.
That matters because Backstage permits the field: a kind's schema does not forbid unknown keys and the catalog stores them. Real catalogs have them, whether from a house processor or a vendor distribution.
Change. A new spec-relations: section in --type-mapping names the keys to read and the entity
kind a bare reference in each defaults to:
spec-relations:
deployedTo: Resource
predicates:
deployedTo: https://vocab.example.org/arch#deployedTo
relationships:
deployedTo: https://vocab.example.org/arch#Deployment # optional
A list yields one statement per item, which is the case this exists for. The default kind plays the part
the descriptor format plays for spec.owner (Group), so a bare aws-account-… resolves without the
author writing the kind at every use.
Values resolve as references, the same three ways an extension value does — a prefixed name against
namespaces:, an absolute IRI, and then one rule belonging to Backstage: anything else is an entity
reference, read with the declared default kind. So a field can point at another catalog entity or at a
system outside the catalog (https://k8s.example.com/clusters/prod), and a relationship target is no
longer necessarily an entity of the model. --ns-global-id is deliberately not consulted, since it would
claim every bare name for an IRI base and take prod-cluster away from the reading a catalog means.
The declared kind belongs to that last reading only. An IRI value is an address already, so there is
no triplet to complete and the kind is never consulted — but the entry under spec-relations: is still
what makes the field readable at all, so it stays even where every value of a field is an IRI. Nothing is
asserted about a target outside the catalog: it is the object of the direct triple, or the arch:target
of the qualified relationship, and carries no rdf:type, label or folder membership, nor any report for
the absence, this model not claiming to describe it. Where the target is an entity its kind need not be
one of the seven — a house kind resolves and mints the same IRI that kind's own descriptor mints, taking
its class from elements: like any unmodelled kind.
No class, no qualified relationship. With an entry in relationships: the field produces a full
arch:QualifiedRelationship, structurally identical to the edges the documented fields produce. Without
one it produces the direct triple alone — and that triple is always written, because it is the
statement's only form rather than a shortcut, so --emit-direct-rel-triples does not gate it. The
previous bs:{Type} class fallback is gone: bs: is a published document this converter reads and does
not own, and bs:Deployment would have dressed a house term up as part of the Backstage ontology with
nothing declaring its meaning. Nothing that used to be emitted is lost, since only documented relation
types existed and every one of those has a published class.
A literal-valued house field has its own declaration, spec-literals:, listing keys whose value is a
value rather than a reference:
<…/element/component/default/order-service>
<https://vocab.example.org/arch#tier> "gold" ;
<https://vocab.example.org/arch#replicas> "3"^^xsd:integer .
A list rather than a map, because a literal needs nothing declared per key — including its datatype,
which comes from the catalog. The YAML loader has already resolved 3 to an integer, true to a
boolean and 2024-01-15 to a date by the time the converter sees the value, so that resolution is
mapped onto XSD: xsd:integer (not the narrower xsd:int), xsd:double, xsd:boolean,
xsd:dateTime normalised to UTC, and a plain literal for a string. Declaring the datatype in
configuration would let it disagree with the file, and the file is the fact.
Two consequences of YAML 1.1 implicit typing are documented rather than worked around, since they happen
before this converter is involved: no is a boolean, so criticalRegion: no yields
"false"^^xsd:boolean, and 1.0 is a double. Quoting in the catalog is the remedy. Leading zeros are
already safe — 012345678901 stays a string.
Declaring one key under both sections is contradictory: nothing is emitted for it and the run says so.
An undeclared key is now reported rather than dropped in silence, once per run and grouped by key,
as unmapped annotations already were. So does a declared key that reached no predicate, and one whose
value is a nested mapping — toString() on a YAML node yields {region=eu-central-1}, which is not a
reference and is not worth emitting. A declared key a descriptor does not carry stays silent: the
declaration describes the catalog's schema, not every entity in it.
Guidance is unchanged and now spelled out in the docs: for a catalog you author, a domain-prefixed
metadata.annotations key remains preferable, and spec.dependsOn remains the way to state a
dependency edge with published terms. This exists so a catalog that already carries house fields is
convertible.
No OWL axioms are emitted for a house term. The predicate is used and the class appears as an rdf:type
object, but nothing declares a owl:ObjectProperty, a owl:DatatypeProperty, rdfs:domain or
rdfs:range — as is already true of the published bs: terms. A converted graph states facts about a
catalog; the vocabulary belongs to whoever publishes it.
See Backstage → Custom spec fields and Type mapping → House fields in a Backstage descriptor.
Added — an ArchiMate property value can be a link, not only a string¶
Previous behaviour. A promoted property key carried a literal. Apart from the two reserved keys
@id and @type, there was no way to say "this value is an IRI", so a published owl:ObjectProperty
could not be populated from an ArchiMate model:
That violates the property's rdfs:range arch:ArchitectureState and loses the IRI join, so no query can
follow it into the state vocabulary. The alternatives were to fork the term as a local datatype property
or to leave the fact out of the graph.
Change. A new object-properties: section in --type-mapping lists the keys whose values are
references:
<…/relationship/id-rel> a arch:QualifiedRelationship, arch:ModelConcept, am:Realization ;
arch:architectureState arch:Target . # an IRI
Works on elements, relationships and views alike. Values resolve the three ways @type already does — an
absolute IRI, a prefixed name against namespaces: falling back to --ns-vocab, or a bare local name
against --ns-vocab — and that resolution is now one function used by both routes rather than two copies.
A value that resolves to none of those emits no link. It is kept as a schema:additionalProperty
pair and reported, rather than emitted as a literal that would satisfy the shape of the triple and
contradict the property's range.
Declared rather than detected. Whether a predicate is an owl:ObjectProperty is in the ontology, and
this does not read it: conversion is offline — only validate fetches ontologies — so detection would
make one model convert differently depending on whether a network was reachable; it could not work for an
unpublished in-house vocabulary; and it contradicts the rule every other extension route follows, that
links and literals are stated, not inferred.
TypeMapping.objectProperties lives in core, so the declaration is available to any converter. Wired
into archimate2linkedarchi, which is the one that needed it.
Also fixed. The report added with the @type fix bucketed a leading-colon key as a prefixed name
and told the reader to declare a prefix. :costCentre is not a prefixed name — the colon is the
"promote me" marker — so its remedy is --ns-vocab, and that is what it now says.
Fixed — an ArchiMate @type property is no longer discarded, and pair fallbacks are reported¶
Previous behaviour. An element, relationship or view property takes one of five routes in
archimate2linkedarchi, and four of them fall back to a schema:additionalProperty pair when no
namespace can promote them, so the value survives as a reified key/value blank node. @type had no
fallback: when its value resolved to no class IRI the property was discarded outright — no rdf:type, no
pair, and nothing in the run output.
A value resolves when it is an absolute IRI, a prefixed name whose prefix --type-mapping declares, or a
bare local name with --ns-vocab set. So a model authoring @type: ManagedService and converted without
--ns-vocab lost that property silently. Reproduced against 1.3.0-SNAPSHOT:
value "ManagedService", no --ns-vocab → 0 occurrences in the output
value "ManagedService", --ns-vocab set → 1 occurrence, as `a vocab:ManagedService`
Change. An unresolvable @type keeps its value as a pair, as @id already does when
--ns-global-id is absent. No class is claimed, since inventing one would put an undeclared term in the
graph. The blank-node emission had been written out three times with the third copy missing, so it is now
one function.
Also fixed: the pair fallbacks were silent. undefinedProperties was collected into
ConversionStats and read by nothing, so a property arriving as an unqueryable pair was as invisible as
one that had been dropped. Reported at the end of a run now, bucketed by remedy, because the three have
different fixes:
| Situation | Remedy reported |
|---|---|
a plain key, e.g. gitUrl |
prefix it with : and pass --ns-vocab |
| a prefixed key whose prefix nothing declares | declare it under namespaces: in --type-mapping |
a @type value that resolves to no class |
write it as an IRI, a declared prefixed name, or set --ns-vocab |
ConversionStats gained unresolvedTypes, kept apart from undefinedProperties because a key needs a
namespace or a rename while a @type needs its value written as something resolvable.
docs/converters/archimate.md now documents all five routes. Its --ns-global-id row was wrong and is
corrected: it credited @type to that option, when resolution depends on --ns-vocab or a declared
prefix.
Added — backstage-pull, and --source-map on backstage2linkedarchi¶
A Backstage catalog is one descriptor per service, each in the repository of the service it describes.
It is the only input in this project that is pulled rather than committed beside the model, and that
has a consequence --git-provenance cannot work around.
Previous behaviour. --git-provenance auto asks git about the working tree. For a descriptor fetched
from another repository, that describes the pipeline's own checkout, so the recorded path and forge URL
name a file that repository has never held:
https://git.example.org/group/models-backstage/-/blob/b292b0ee…/catalogs/orders/order-service.yaml
# ^ the pipeline's repo ^ a path only the pull ever created
The commit a fetch resolves to exists only at the moment of the request. Once the file is on disk it is indistinguishable from any other file, so nothing downstream can reconstruct it.
backstage-pull reads a manifest of repositories, fetches each descriptor over GitLab's repository
files API, and records what it learned:
backstage-pull pull --sources catalogs/sources.yaml -o catalogs \
--write-index --model-id service-catalog
backstage2linkedarchi convert catalogs --base-iri https://example.org/la/ \
--source-map catalogs/catalog-sources.yaml --format TRIG -o out.trig
Provenance then describes the repository each descriptor came from:
<…#source-4f2c1ab8-catalog-info-yaml>
a prov:Entity ;
dct:source "catalog-info.yaml" ; # the path upstream
dct:identifier "4f2c1ab8e0d1c2b3a4958677889900aabbccddee" ;
rdfs:seeAlso <https://git.example.org/group/order-service/-/blob/4f2c1ab8e0/catalog-info.yaml> .
Three artifacts, with deliberately different ownership.
| File | Owner | Committed | Rewritten |
|---|---|---|---|
sources.yaml |
you | yes | never |
| the descriptors | upstream | no — gitignore them | every pull |
catalog-sources.yaml |
the pull | yes | every pull |
catalog-index.yaml |
you | yes | written once, then never |
The source map and the index are two files rather than one because a commit changes on every pull, so it cannot live in a file a pull must not overwrite, while editorial decisions accumulate across dozens of lines and must not be discarded by a regeneration. The descriptors are copies of files other repositories own; committing them would make this repository a second, stale source of truth for them.
Decisions worth stating:
- The recorded commit is
last_commit_id, notcommit_id. The head of a ref moves on every unrelated push to a source repository, so recording it would rewrite the map — and every provenance IRI derived from it — on a pull where no descriptor had changed. With sorted entries and a fixed key order, an unchanged upstream now produces an unchanged file, which is what makes the map reviewable as a diff. - A missing descriptor exits non-zero unless
--allow-missing. A partial rollout is ordinary, but a catalog silently missing a quarter of the estate looks exactly like a complete one and the graph cannot show the difference. --anonymousis stated, not inferred from an absent token. A pull that silently went anonymous would report every private repository as having no descriptor.- GitLab only, the same limit
GitProvenancedraws and for the same reason. Every source in a manifest must be on one instance, checked before any request: readinggroup/appfrom the wrong host either 404s — which would be reported as "no descriptor" — or finds a different project of the same name.
In core. SourceMap owns the format, so the puller and the converter share it and neither depends on
the other. ConversionProvenance.SourceRef gained commit, and emit prefers it over the run's: one
run now describes many source commits, which is what a catalog collected from forty-four repositories
is. BaseLinkedArchiEmitter.emitProvenance gained locateSource, so a converter whose input was fetched
from elsewhere can say where it came from instead of asking git about a directory the file was copied
into.
An entry in the map wins over --git-provenance for the file it names. A file the map does not mention
falls back to git and is reported: a catalog may mix pulled descriptors with committed ones, and the
fallback is right for the committed half, but a map that has drifted from its manifest would otherwise
attribute the very files it exists to describe to the wrong repository.
Documented at Collecting a Backstage catalog.
Added — each Backstage entity records which catalog file it came from¶
Previous behaviour. dct:source on the model listed the files a model was built from. Nothing
recorded which file an individual entity came from. For the two input modes this converter is built
for — an index selecting many files into one model, and a directory walk — that is the question a
consumer has, because a catalog entry pulled from another repository is only traceable per entity.
Directory mode could not answer it even in principle: BackstageParser.parseDirectory merged every
file into one model, BackstageModel.sourceFile named the directory, and BackstageEntity had no
path field.
Change. BackstageEntity and BackstageRelation gained sourceFile: File?, populated by
parseEntity from both parse and parseDirectory. The emitter resolves one prov:Entity per
distinct file and links each concept to its own:
<…/graph/provenance> {
<…#source-7cac3422-catalogs-orders-catalog-info-yaml>
a prov:Entity ; dct:source "catalogs/orders/catalog-info.yaml" .
<…#source-7cac3422-catalogs-payments-catalog-info-yaml>
a prov:Entity ; dct:source "catalogs/payments/catalog-info.yaml" .
<…/element/component/default/order-service>
prov:wasDerivedFrom <…#source-7cac3422-catalogs-orders-catalog-info-yaml> .
<…/element/component/default/payment-service>
prov:wasDerivedFrom <…#source-7cac3422-catalogs-payments-catalog-info-yaml> .
}
Details:
prov:wasDerivedFrom, not a per-entitydct:source. The target already carries the path, the commit and the blob URL, so repeating the path per entity would state one fact in two places. No new term is needed.- In the provenance graph, per ADR 0008.
- Relationships too. A relationship is an
arch:QualifiedRelationshipand a model concept, so its origin is a fact about it. Forspec.childrenandspec.membersthe recorded file is the one holding the field, which is the target group's file rather than the source's. - Only under
--git-provenance, since that is what mints the target. An unconditional link would point at a node nothing types or names. - A directory is no longer described as a source file.
emitProvenancegainedprovSources, the files a conversion actually read.opts.inputFileis the directory in directory mode, and describing it as aprov:Entityattached adct:sourceand a blob URL to something nothing was read from. The model's owndct:sourcestill names it, since that is what was converted. - One
prov:Entityper file, however many entities the file holds, and onegit ls-filesper file rather than per entity.
Single-file mode is unchanged: the one entity described is the input file, as before.
Fixed — folder positions no longer repeat when several files build one model¶
Previous behaviour. schema:itemListElement entries carry a schema:position, and the counters
live on the emitter, so BaseLinkedArchiEmitter.resetState has to be told which model is being built.
It continues the counters while the model id is unchanged and clears them when it changes; called with
no argument it clears them on every call.
converter-backstage, converter-leanix and converter-structurizr called it with no argument. Each
loops over its inputs with one emit call per file, so any run where several files share one model
restarted at position 1 per file:
schema:position "1"^^xsd:int ; schema:item <…/element/component/default/alpha-service> .
schema:position "1"^^xsd:int ; schema:item <…/element/component/default/beta-service> .
Affected: backstage2linkedarchi with --diagrams-index selecting several files into one model — the
mode its own --help recommends for multi-repo setups; structurizr2linkedarchi convert *.json
--model-id one-model, which is what the example-architecture-project template runs; and
leanix2linkedarchi with several exports in one model. Backstage's directory mode was correct, being a
single emit call, so two input modes of one converter disagreed.
Change. All three now pass opts.modelId. Positions are unique within a folder across the files of
a model, and still restart for the next model in the same run.
Fixed — a Backstage relationship IRI is derived from the edge, not from its position¶
Every Backstage relationship IRI moves.
before …/backstage/service-catalog/relationship/rel-1
after …/backstage/service-catalog/relationship/ownedBy--component-default-order-service--group-default-team-platform
Previous behaviour. The local id was rel-N from relCounter, a variable local to emit. The
counter is per call and the IRI namespace is per model. An index-driven run is one call per
file, so two files contributing to one model each restarted at rel-1 and their first relationships
collided on one IRI:
<…/service-catalog/relationship/rel-1> a arch:QualifiedRelationship, bs:Ownership ;
arch:source <…/element/component/default/order-service> , # from catalogs/orders/
<…/element/component/default/payment-service> ; # from catalogs/payments/
arch:target <…/element/group/default/team-platform> .
One resource asserting it runs between two different pairs, so a consumer reads an edge nobody wrote. Directory mode was unaffected, being a single call, so the two input modes of one converter disagreed.
Change. The id is composed from the triple Backstage identifies a relation by — type, source,
target — following ADR 0005,
which settled the same question for PlantUML and named this converter in its context. The composed form
is published as skos:notation, which previously carried rel-N.
Bounded with IriSegment.fits / IriSegment.bounded rather than unconditionally, following the LeanIX
converter: Backstage caps metadata.name and metadata.namespace at 63 characters and a kind comes
from a fixed vocabulary, so an ordinary segment lands near eighty characters. Three maximum-length parts
still compose past 255 bytes, and a path segment becomes a filename wherever the graph is written as
documents, so the guard is kept.
Three consequences:
- Ids no longer depend on file order. Adding a file does not renumber the relationships of any other file.
- The same edge declared in two files is one relationship. This corrects a second defect that
BackstageOntology040Testhad been recording as expected: aspec.parenton the child and aspec.childrenon the parent describe one edge, and the graph held it twice. The test's name and comment both said one; only its assertion said two. - A residual ambiguity is accepted.
ns=default, name=a-bandns=default-a, name=bslug alike, so two edges could in principle compose the same notation. Reaching it needs a namespace that is a prefix of another endpoint's name plus a shared type and shared other endpoint, and the check below catches it rather than letting it pass.
Added — a relationship with two endpoints under one predicate is reported¶
ConversionVerifier.checkRelationshipsHaveEndpoints checked for a missing arch:source or
arch:target. It now also reports more than one of either:
produced 1 structurally incomplete resource(s) — this is a converter defect, please report it:
Relationship rel-1 has 2 arch:source values (order-service, payment-service) — two relationships
were minted onto one IRI, so its local id does not tell them apart
Two endpoints under one predicate means two distinct edges were minted onto one IRI. The cause is always a local id that does not distinguish the edges it is derived from. The check runs over the merged model in every converter, so it names the defect on the run that produced it rather than in a shape report days later. It is what surfaced the Backstage defect fixed above.
Added — --ns-vocab on backstage2linkedarchi, for annotations with no published term¶
Previous behaviour. metadata.annotations keys present in
BackstageWellKnown.SIMPLE_ANNOTATION_PROPERTIES were emitted as bs: properties. Every other key
was discarded, with no triple and no log line. Backstage defines the annotation vocabulary as
open-ended, so any organisation-specific key was lost.
Two routes now resolve a key to a predicate.
- A prefixed key resolves through a namespace declared in
--type-mapping. A Backstage annotation key has the formprefix/name; the prefix identifies the vendor and the name becomes the local name.
namespaces:
gitlab.com: https://vocab.example.org/gitlab#
forge.example.org: https://vocab.example.org/forge#
<…/element/component/default/order-service>
<https://vocab.example.org/gitlab#project-slug> "group/order-service" ;
<https://vocab.example.org/forge#project-slug> "platform/order-service" .
- An unprefixed key resolves against
--ns-vocab:costCentrebecomesvocab:costCentre. The option is new here; Backstage was the only one of the six converters without it.
A prefixed key is not resolved against --ns-vocab by discarding its prefix.
github.com/project-slug and gitlab.com/project-slug would both yield vocab:project-slug, giving
one predicate two values on one subject. This is the defect fixed in the Structurizr converter below,
so a prefixed key requires a declared prefix.
No term is minted in bs:. The namespace is a published document this converter reads and does not
own. The live ontology declares bs:gitlabUserId and no GitLab project slug, so gitlab.com/project-slug
has no term to map to. Requested upstream in
todo/UPSTREAM-REQUEST-backstage-annotations.md.
Unresolved keys are reported. Once per run, grouped by cause, since the three causes have different remedies:
| Cause | Remedy reported |
|---|---|
unprefixed key, no --ns-vocab |
pass --ns-vocab |
| prefixed key, prefix not declared | declare it under namespaces: in --type-mapping |
| name is not a legal IRI local name | none; rename the annotation |
Grouped by key rather than by entity, so one annotation on forty components produces one message.
Correction to todo/UPSTREAM-REQUESTS.md §4. backstage.io/source-location is mapped, and was
mapped in the build that report measured. Its absence from the artifact indicates that no converted
entity set the annotation.
Fixed — a malformed diagram-index extension block fails the run instead of being ignored¶
Previous behaviour. elements:, links: and data: are parsed as YAML mappings
(DiagramsIndex.kt). A value of any other type failed the cast and the parser returned an empty
assertion list. The conversion then completed normally, emitting the file's remaining triples with no
error, no warning and no log entry mentioning extension data. Triple counts were identical to a run
with --emit-extension-data and the index entries absent.
The affected shape, a list where a mapping is required:
elements: # ignored: a list, not a mapping keyed by subject
- id: order-service
data:
- predicate: https://vocab.example.org/tam#repo
value: https://git.example.org/group/order-service
New behaviour. A block that is present but not a mapping fails the run with a message naming the index entry and giving the required form:
Malformed 'elements:' in bad-index.yaml (entry 'catalog'): expected a mapping keyed by element,
got a list of 1 item(s). Write:
elements:
OrderService:
links:
am:realizes: kg:CAP-OrderManagement
data:
x:costCentre: CC-4711
Applies to elements:, to an elements: entry body, and to links:/data: at element, view and model
level. An absent block remains silent.
Failing rather than warning follows the existing decision in the same function, which already refuses a malformed value rather than skipping it.
No documentation change. The schema is published in
docs/architecture/extension-data.md, including the Backstage
row, and is in the site navigation.
Fixed — two spellings of one Structurizr property key no longer share a predicate¶
Previous behaviour. :gitUrl and gitUrl are distinct keys in a workspace properties map and
both resolve to vocab:gitUrl. Both were emitted:
One predicate with two values on one subject, and no warning. This contradicts the one-value-per-key
rule stated in models/c4/docs/authoring-properties.md.
New behaviour. emitProperties resolves the target predicate for every key on the subject before
emitting anything, then detects predicates claimed by more than one key.
- The colon-prefixed spelling keeps the predicate. It is the only spelling
archimate2linkedarchipromotes, so a property written both ways yields the same predicate from either converter. Remaining ties are broken alphabetically, so the result does not depend on map iteration order. - Each displaced key is emitted as a
schema:additionalProperty/schema:PropertyValuepair, which carries its key as a literal. No value is dropped. - The clash is reported once per run under
PairReason.SPELLED_TWICE, naming the displaced keys.
An uncontested key is unaffected. Predicate resolution is now a single function
(promotedPredicate) used by both the clash check and the emission.
Unchanged. ArchiMate requires the leading colon and Structurizr accepts either spelling. That
difference is deliberate and documented in StructurizrVocabulary: an Archi property key is free text
entered in a dialogue, a Structurizr key comes from DSL source. Only the merge was a defect.
Fixed — every source entity now describes the file it is named for¶
Previous behaviour. In a multi-file conversion the prov:Entity IRI was minted per file, while the
dct:source and rdfs:seeAlso on it were read from the run's SourceCommit. Every source entity
therefore reported the path and blob URL of the file that provenance was detected from — the first
active input. Measured downstream over a 36-file catalog:
SELECT (COUNT(DISTINCT ?e) AS ?entities) (COUNT(DISTINCT ?src) AS ?paths) WHERE {
GRAPH ?g { ?e a prov:Entity ; dct:source ?src }
}
# entities = 36 paths = 1
One Run is shared by every model a process converts, so the incorrect value also crossed model
boundaries.
Cause. SourceCommit held repoRelativePath and blobUrl. Both are per-file values, and the
record describes a run, which resolves once and is reused for every input.
Change. Those two fields moved to a new ConversionProvenance.SourceRef, constructed per file by
GitProvenance.locate(file, commit). SourceCommit retains repoRoot and projectUrl so a per-file
value can be derived. ConversionProvenance.emit takes the SourceRef and reads both the entity IRI
and its payload from it.
| Holds | Resolved | |
|---|---|---|
SourceCommit |
sha, committedAt, authorName, commitUrl, repoRoot, projectUrl |
once per run |
SourceRef |
fileName, repoRelativePath, blobUrl |
once per file |
A source entity is now identified by its repository path rather than its file name. Under
--git-provenance the fragment changes:
before …/graph/provenance#source-3d4a705a-catalog-info-yaml
after …/graph/provenance#source-3d4a705a-catalogs-one-catalog-info-yaml
Required, not cosmetic: file names repeat across directories. A pulled catalog of 44 repositories each
publishing catalog-info.yaml produced one entity for all 44 when keyed on the name, which — once each
file reported its own path — would have placed 44 dct:source values on one subject. Provenance IRIs
are opaque and are not accumulated across runs, so no consumer dereferences the previous form. A run
without git provenance is unchanged: no path is available and the name remains the identity.
dct:source on the model or view is unchanged. It is written once per conversion and was already
per-file correct.
Fixed — no blob URL for a path its commit does not track¶
Previous behaviour. --git-provenance auto composed {project}/-/blob/{sha}/{path} from the
recorded commit and the file's path without checking that the commit tracks the path. A pipeline that
pulls its inputs into an ignored working tree and converts them there produced URLs for paths that
existed at no commit:
https://git.example.org/group/models-backstage/-/blob/3d4a705a…/catalogs/one/catalog-info.yaml
# ^ .gitignore contains catalogs/*/
New behaviour. GitProvenance.locate runs git ls-files --error-unmatch for the path and omits
rdfs:seeAlso when the path is untracked or ignored. One git invocation per file. The commit sha, the
repository-relative path and the author are still recorded.
Also changed. A file outside the repository root now reports repoRelativePath = null rather than
its bare file name. The previous value composed into a blob URL resolving to the repository root. Its
dct:source is unchanged, since SourceRef.identity falls back to the file name.
Raised downstream as items 1 and 3 of todo/UPSTREAM-REQUESTS.md; see
todo/PROPOSAL-backstage-provenance.md for the full analysis. Regression test: over a multi-file run,
the count of distinct dct:source values equals the count of source entities, and each entity reports
the file its own IRI is minted from.
Added — url, perspectives, !docs and !adrs, all in terms already published¶
Four Structurizr fields were parsed by nothing and emitted by nothing. Together they were the largest
thing still being dropped, and the loss was not proportional to the count: !docs is the prose that
explains why the boxes are arranged as they are, !adrs are the decisions that put them there, and a
perspective is the one place a security or performance judgement about an element is written down. A
converted workspace kept the diagram and lost the argument.
None of the four needed a new term.
url → schema:url, typed xsd:anyURI — the predicate and datatype the LeanIX converter already
uses for a fact sheet's page. Structurizr declares url on ModelItem, so elements and relationships
both carry one.
A perspective is a schema:PropertyValue. Name, description and value are each the author's own
words, and all three survive:
<…/element/2> schema:additionalProperty [
a schema:PropertyValue ;
schema:propertyID "structurizr.perspective" ;
schema:name "Security" ; schema:value "High" ;
schema:description "Handles cardholder data end to end." ] .
Not arch:Perspective. Core publishes that class, but it means a stakeholder perspective —
strategic, operations, physical — and its two properties have domain arch:Stakeholder and
arch:Viewpoint. Nothing relates an arch:Element to one and nothing carries a value for one, so
typing a Structurizr perspective as arch:Perspective would claim membership of a class whose
published examples are a different kind of thing and still leave the value homeless. schema:propertyID
is what separates a perspective from a workspace property, which shares schema:additionalProperty
correctly — schema.org defines it as "an additional characteristic of the entity" and both are that.
A documentation section is a schema:CreativeWork the element is schema:subjectOf — schema.org's
"a CreativeWork about this Thing". Its own addressable resource rather than a literal, because there can
be several, they are ordered, and each has its own media type:
<…/element/2> schema:subjectOf <…/documentation/2/1> .
<…/documentation/2/1> a schema:CreativeWork ;
schema:text "## Sun Deals…" ; schema:encodingFormat "text/markdown" ; schema:position 1 .
Not skos:definition, which already carries the element's one-line description; two values under one
predicate with nothing to tell them apart is a collision this project has had to unpick before, and a
page of Markdown is not a definition. Workspace-level !docs attaches to the model.
An ADR is an ad:Decision — and this is the one that needed no schema.org either, because
arch-decision was published all along:
<…/decision/1> a ad:Decision, arch:Element, arch:ModelConcept ;
skos:prefLabel "Use a micro-frontend for the offer UI"@en ;
ad:decisionState ad:Superseded ;
ad:supersededBy <…/decision/3> ;
ad:relatedConcept <…/element/2> ;
dct:date "2026-03-04T00:00Z"^^xsd:dateTime .
Structurizr's five ADR statuses are exactly the five published ad:DecisionState individuals, so the
mapping is a transcription rather than an interpretation. ad:relatedConcept — published range
arch:ModelConcept — is what makes "which decisions affect this container" answerable.
This is the first converter to emit ad: subjects rather than link to one.
PublishedAssets.ARCH_DECISION_ONTOLOGY recorded that the converters "only ever link to one", which
held while !adrs was unread; a Structurizr workspace carries the records inline. So arch-decision
joins this converter's default ontologies, because ad:Decision rdfs:subClassOf arch:Element has to
resolve for a decision to be reachable from anything targeting arch:Element. Its shapes stay opt-in
as they are everywhere — validate --shapes arch-decision-shapes — since most workspaces hold no
decision records and constraints targeting classes nothing instantiates validate nothing.
The ADR body becomes the rationale document the shapes ask for —
ad:justificationDocument → a schema:CreativeWork — which is what
adsh:AcceptedDecisionRationaleShape requires and the only one of the two rationale properties with
somewhere to record that the prose is Markdown. ad:justification takes a bare xsd:string.
Each lifted ADR carries prov:wasDerivedFrom to the source file's prov:Entity, in the provenance
graph. A decision record is an artifact in its own right, so its origin is a fact about it and not only
about the model: it is what distinguishes "the options are prose in a Markdown body" from "no options
were recorded". Provenance is where a lift marker belongs (DD-15), and keeping it out of the semantic
graph is load-bearing — adsh:AcceptedDecisionSelectionShape relies on a validator over the semantic
graph not seeing it, so conformance never depends on how a consumer assembled the graph. Emitted only
under --git-provenance and friends, since that is what mints the entity being pointed at; a
prov:wasDerivedFrom aimed at an undescribed node would be worse than none.
The target is the workspace file, which is the finest granularity the export offers: Structurizr's
Decision carries an id, title, date, status, links and body and no filename — unlike a
documentation Section — so the individual .md path is not recoverable, and deriving one from the id
would be a guess about a directory layout the converter never saw.
Output conforms to the whole decision shape set, verified against arch-decision-shapes 0.2.0.
adsh:AcceptedDecisionSelectionShape used to report on every accepted ADR — a Structurizr ADR has no
options to name, and minting an ad:Option to satisfy the shape would assert that alternatives were
weighed when the graph has no idea whether they were. That was raised as issue 9 in
todo/UPSTREAM-ISSUES-meta.md and fixed upstream: the shape now fires only on a decision that records
candidate options with ad:hasAlternative, so a record listing alternatives and no outcome still fails
while a lifted prose ADR is out of scope. Both fixes this project proposed were rejected in favour of
that one, for reasons kept in the issue.
Two omissions, both reported rather than silent:
- An unrecognised ADR status gets no state.
adsh:constrainsad:decisionStatewithsh:inover the five individuals, so mintingad:UnderReviewwould turn correct input into a validation failure. The token is kept ondct:type. - Embedded documentation images are not emitted. Structurizr stores an image's bytes base64-encoded in the workspace, so publishing one would put a whole file in a single literal — megabytes per screenshot, not queryable and not dereferenceable. The prose around them is emitted in full.
IriMinting gained decisionIri and documentationIri. A decision gets its own …/decision/<id>
segment rather than sharing element/, even though ad:Decision is an arch:Element: ADR ids are
numbered independently of element ids, so sharing would let ADR 3 and container 3 collide on one IRI.
Added — Structurizr custom elements, typed as honestly as the vocabulary allows¶
model.customElements[] — the DSL's element keyword — was not read at all. The key appeared nowhere
in the module, so an author who modelled an S3 bucket, a regulatory obligation or a partner data feed as
a custom element contributed nothing for it and got no warning. views.customViews was unread for the
same reason, which meant the one place those elements can be drawn was missing too.
The typing is the whole of it. Structurizr defines a custom element as one that "sits outside of the C4
model", so no c4: class is right for it and the deployment extension does not reach it either:
<…/element/30> a arch:Element, arch:ModelConcept ;
skos:prefLabel "Amazon S3"@en ;
dct:type "Infrastructure" . # the metadata token, verbatim
No class is invented, and there is no published root to fall back to. This is where the case differs
from the LeanIX converter's: leanix/onto publishes lmm:FactSheet, so an undeclared fact sheet type
still gets a true, published class that keeps it inside the reach of the naming shape. c4/onto
declares four classes and all four are C4 levels — there is no c4:Element — and
c4sh:C4ElementLabelShape enumerates those four rather than targeting arch:Element, deliberately,
with a comment saying that targeting it "would impose C4's naming policy on every other notation". A
c4:CustomElement minted here would look like part of C4 while nothing declared its meaning and no
shape could constrain it, which is the objection ADR 0006 turns on.
The consequence is stated rather than hidden: an unmapped custom element is covered by no shape, so
nothing checks that it has a name. Each unmapped kind is reported once per run with the option that
fixes it, and the gap is filed upstream as issue 8 in todo/UPSTREAM-ISSUES-meta.md, which proposes
structurizr:CustomElement plus a label shape in the extension's own shape set.
metadata is the type token. element "Amazon S3" "Infrastructure" puts Infrastructure there,
and Structurizr renders it where a C4 element shows [Container: Spring Boot]. It is the only statement
of kind in the file, so it is kept on dct:type whether or not a class was found — the same treatment
LeanIX gives a fact sheet type, and for the same reason. Read only for custom elements, so a metadata
key on another element kind cannot silently become a type.
--type-mapping gives them a real class, keyed on the author's own token or on custom for all of
them at once, with a token entry winning over the blanket one:
elements:
Infrastructure: https://vocab.example.org/arch#Infrastructure
custom: https://vocab.example.org/arch#External
Two things needed no special handling, which is what makes the missing class tolerable rather than
crippling. c4:Using constrains its endpoints to arch:Element and c4sh:UsingShape enforces exactly
that, so a relationship between a container and a custom element is valid as it stands, in either
direction. And a custom view is emitted as an ordinary arch:View claiming no viewpoint — the published
C4 catalogue declares seven and none of them is this, because a custom view is not a C4 level. Same
outcome a filtered view already had, reached for a different reason.
Added — one conversion can write several artifacts¶
--output is now repeatable, and each value is path[:FORMAT[:PROFILE]]:
bpmn2linkedarchi convert models/*.bpmn --base-iri https://example.org/la/ \
-o out/model.trig \
-o out/model.ttl \
-o out/model-only.ttl:TURTLE:no-diagrams
A repository usually publishes the same model more than once — TriG for triplestores, Turtle for tools that cannot read it, and a reduced artifact for consumers that do not render diagrams. That meant running the converter once per file, which was wrong in two ways.
The obvious one is waste: the parse and type mapping ran in full each time. Three artifacts from
archimetal.xml took 1,211 ms as three invocations and take 495 ms as one — about 2.4x, and the
saving grows with the model, since it is the repeated parse that dominates.
The one that matters more is that separate runs disagree about themselves. Each stamps its own
prov:generatedAtTime and its own activity start and end, so three files from one publish claimed
three different generation times. ConversionProvenance.Run is resolved once per run precisely so
every model in a run attributes itself to the same activity, image and commit; invoking the binary
three times defeated that, and an aggregation store loading the artifacts saw three publishes where
there had been one. Now every artifact from a run carries the same run description.
View profiles. The last segment of an --output decides how much of the diagrams that artifact
carries. This is a Turtle problem: a TriG consumer can shed geometry after loading with
DROP GRAPH <…/graph/views>, but Turtle has no graph boundaries, so the cut has to happen when the
file is written.
| Profile | Drops | Archisurance |
|---|---|---|
full (default) |
nothing | 8,094 triples |
no-geometry |
bounds, bendpoints, styles; keeps each view and what is on it | 3,855 |
no-views |
the whole views graph; keeps arch:View with its name and viewpoint |
— |
no-diagrams |
the views graph and every arch:View |
2,038 |
Geometry is about three quarters of that model. no-geometry is the profile that did not previously
exist outside ArchiMate's --skip-visual-triples; it answers "what is drawn on which diagram" without
"where".
Profiles filter the built model rather than changing what the emitters produce, and that is only well
defined because the semantic/views boundary is now identical in all six converters — before the
alignment below, "drop the views graph" would have kept the diagrams for PlantUML and LeanIX, removed
them for ArchiMate, and removed half of each for BPMN. no-diagrams also removes the folder entries
that listed the removed views, so no folder is left asserting a containment whose other half is gone.
Compatibility. --output x.ttl --format TURTLE behaves exactly as before; format resolution runs
inline token → --format → file extension. Two smaller things changed with it:
--formatis optional in every converter. It was required in Backstage, Structurizr and LeanIX while the other three inferred from the extension — the same tool demanding a flag half of it could work out for itself.--format ttlnow means Turtle. It used to be accepted only by accident: the value matched no enum name, so it was discarded and the format inferred from the filename instead. The two agreed in the repository's ownrun-archimate-ttl.sh, which is why nobody noticed — but had the file been named.trig, the run would have quietly written TriG while the command line saidttl. A genuinely unknown--formatis still refused rather than silently replaced.
The ArchiMate emitter gained a build step separate from write, since it had always serialized
inside the emitter; write is unchanged for existing callers.
Changed — ArchiMate emits views by default¶
--include-views now defaults to true on archimate2linkedarchi, matching every other
converter: PlantUML and Structurizr (--include-views) and BPMN (--include-di) have always
defaulted to emitting diagrams. ArchiMate alone dropped every view unless asked, so "convert this
model" meant something different for ArchiMate than for the same command against any other notation —
and a repository converting several notations with one uniform invocation silently published diagrams
for all of them but that one.
The flag is now negatable, so --no-include-views restores the previous behaviour. This is a
behaviour change for anyone relying on the old default: an ArchiMate conversion that took no flag
grows from elements and relationships to include views and their geometry, which is roughly four times
the triples — Archisurance goes from 2,038 to 8,094.
Making it negatable needed fallbackValue = "true" alongside the explicit default. Without it picocli
treats the positive form of a flag whose field is already true as a toggle, so --include-views
would have switched views off. That trap had already been fixed twice in this same file, on
--validate-refs and --emit-core-triples.
Worth being precise about what the negative form does, because it is easy to assume otherwise: it
drops the diagrams entirely — no arch:View, no view nodes — rather than keeping them without
coordinates. The same is true of --no-include-di on BPMN and --no-include-views on PlantUML and
Structurizr, since in every case the flag gates the block that emits the view resource itself.
--skip-visual-triples on ArchiMate is the narrower cut: it keeps the views and their nodes and omits
only geometry and styling.
The example-architecture-project template uses this to publish a third artifact per notation,
out/{notation}-model.ttl — the flat union with the diagrams excluded, for consumers that cannot read
TriG and so cannot simply DROP GRAPH the views graph. Geometry is most of a converted model, so it
is a much smaller load for anything that only queries the architecture. It is skipped where it would
duplicate the plain Turtle: Backstage, which has no diagrams at all; BPMN when no source carries a
<BPMNDiagram> section; and LeanIX without a pulled diagram export.
Fixed — the semantic/views boundary is the same in every converter¶
The converters emit three named graphs per model, and two of them drew the line between semantic
and views differently from the other four. The triples were all present and all correct; they were
filed in the wrong graph, which nothing asserted on, so nothing failed.
The rule, now shared: the boundary falls around geometry, not around diagrams. A view is a model
concept — something the model contains — so the arch:View resource with its notation, label,
viewpoint and folder belongs in semantic. Only archvis: nodes, links, points and their bounds
belong in views. PlantUML, Structurizr and LeanIX already did this.
ArchiMate put the entire view in views — its type, name, documentation, viewpoint,
schema:image and folder. Two consequences. A query for the model's diagrams scoped to the semantic
graph returned nothing for ArchiMate and results for every other notation. And dropping the views
graph to shed geometry — the one operational reason the split exists, worth about 80% of an ArchiMate
artifact — took the diagrams themselves along with it.
The same converter also emitted rdf:type arch:Model, arch:modelConformsToMetamodel, the model
label and its skos:notation into provenance, alone among the six. Which modelling language a model
is expressed in (ADR 0006) is a fact about the model, true regardless of when it was converted; it now
sits in semantic with the rest of the model. prov:generatedAtTime and the PROV description stay
where they were.
BPMN split one view across both graphs: the isBpmnDiagram branch wrote arch:View,
arch:Diagram and the viewpoint into semantic, while the generic DI attribute loop that ran just
before it wrote the same subject's skos:notation and skos:prefLabel into views. A
semantic-only query returned a typed diagram with no name.
Fixing that surfaced two more faults in the same loop, both from it running over every DI subject rather than only the diagram:
- Every shape, edge and label was filed as a member of the
Viewsfolder. A two-diagram file listed 26 members where it should list 2. The folder tree lists elements, relationships and views; geometry is reached througharchvis:view, as it is in ArchiMate. - Each view was filed twice, because a
BPMNDiagramand theBPMNPlanebeneath it remap onto the same view IRI and both were filed — one view under two differentschema:positionvalues, the same two-answers-to-one-question fault ADR 0008 removes elsewhere.
Folder containment is stated from both ends, and the two ends were in different graphs: the member's
dct:isPartOf in views, the folder's ordered schema:itemListElement in semantic. Neither graph
could answer "what is in the Views folder" alone. Both halves are now in semantic, where the folder
tree lives.
Also in this change:
- The three graph suffixes have one definition,
IriMinting.GraphSuffix. The ArchiMate emitter cannot useIriMintingitself — its paths are configurable and--path-modeldefaults tomodel, notarchimate, so routing through it would change every IRI it emits — and it had its own copy of the literals. Its IRIs are unchanged. - Backstage no longer mints a views-graph IRI it never writes to. A catalog has no diagrams, so it has no geometry; the output is unchanged.
No IRI changes and no SPARQL, SHACL or script changes: nothing in the repository queries by graph, and
Turtle output already flattened all three. Consumers that load TriG into separate contexts will see
the view resources move from views to semantic, and BPMN's Views folder listing geometry
disappear.
BpmnViewGraphBoundaryTest and three new tests in ArchiMateViewsTest pin the boundary down; against
the previous code they fail 8 times.
Added — Structurizr workspace properties reach the graph¶
A Structurizr properties map was discarded whole. The key properties did not appear anywhere in
converter-structurizr: not in the parser, not in the emitter, not as a warning. So this, which is what
organisations actually put there —
"properties" : {
"structurizr.dsl.identifier" : "sundeals",
"architect" : "Lead Architect",
"leanIXUrl" : "https://acme.leanix.net/Acme/factsheet/Application/00000000-…",
"runwayUrl" : "https://runway.example.org/catalog/default/domain/offers"
}
— was read out of the file by Gson and thrown away, on elements, relationships and views alike. These are the facts a cross-model graph exists to join on, and a run reported nothing to say they had been dropped.
Every key now reaches the graph. Where it lands depends on how it is written:
| Key | Emitted as | Requires |
|---|---|---|
structurizr.dsl.identifier |
dct:identifier, and optionally the IRI segment |
— |
@id |
owl:sameAs <base + value> |
--ns-global-id |
@type |
an extra rdf:type |
— |
prefix:local |
<namespace><local> |
prefix under namespaces: in --type-mapping |
:local / local |
vocab:local |
--ns-vocab |
| anything else | schema:additionalProperty blank node |
— |
Nothing is dropped for want of an option. Without --ns-vocab a plain key still arrives as a
schema:additionalProperty / schema:PropertyValue pair carrying key and value as literals, so the
choice is between a queryable predicate and a reified pair — not between keeping the value and losing it.
That is the difference from the LeanIX converter's workspace fields, which genuinely are dropped without a
namespace: there the values arrive from a tenant schema with no key to reify.
The dispatch is the ArchiMate converter's, so an organisation converting both notations gets one set
of conventions. One deliberate difference: a plain key becomes vocab:key here, where ArchiMate requires
:key. An Archi property key is free text an author types into a dialogue, often with spaces; a
Structurizr key comes from DSL source or an export script, so architect and leanIXUrl are already
identifier-shaped and asking for a rewrite would be a tax with nothing behind it.
A key that cannot be a local name in an IRI — "Business Owner", anything with a dot or a slash — always
takes the pair route, because no namespace can reach it. The run reports which keys ended up there once
per reason, naming the remedy that actually applies: --ns-vocab for a plain key with nowhere to go,
--type-mapping for an undeclared prefix, a rename for a key no namespace can reach. One message covering
all three would tell most readers to do the wrong thing.
structurizr.dsl.identifier is reserved rather than passed through. It is the only name in a workspace an
author actually chose — the id beside it is a counter the export assigns, which shifts when elements are
added or reordered. It is published as dct:identifier rather than as a second skos:notation, since two
values under one predicate with nothing to tell them apart is the collision this project already had to
unpick from dct:created and dct:creator. --iri-from-dsl-identifier additionally mints IRIs from it,
off by default because switching it on moves every affected address — an
ADR 0003 identity change. The mapping is resolved
for the whole model before anything is minted, so an element and every reference to it move together.
Added — the Structurizr deployment model, in the classes that were already published¶
The parser read only top-level deploymentNodes, with no recursion and no relationships. So a
workspace that models its infrastructure — the ordinary case for anything deployed — contributed one flat
element per root and lost everything beneath it. children, infrastructureNodes and containerInstances
are each one level below a node that was being read, so the deeper the deployment model the less of it
survived. A container instance tagged Container Instance never reached the graph at all.
What did arrive was mistyped. A deployment node came out as c4:SoftwareSystem and an infrastructure node
as c4:Container, under a comment reading // fallback — which is what an untyped element looks like once
a converter has run out of vocabulary. It had not: https://meta.linked.archi/c4/structurizr# declares
both classes, and the C4 ontology's own description names that document and says why the split exists.
<…/element/21> a structurizr:DeploymentNode ;
structurizr:environment "Production" ;
structurizr:instances 3 .
<…/element/20> arch:hasPart <…/element/21> .
<…/relationship/23> a structurizr:Deployment ;
arch:source <…/element/3> ; # the container
arch:target <…/element/21> . # the node it runs on
A container instance is emitted as the relationship it stands for. The extension publishes no
ContainerInstance class; it publishes structurizr:Deployment, and structsh:DeploymentShape constrains
it to run from a c4:Container to a structurizr:DeploymentNode. So the instance's own id becomes the
relationship, its containerId the source, and the node it sits inside the target. --emit-direct-rel-triples
writes structurizr:deployedOn for these rather than c4:uses, which would have said a container uses
the node it runs on.
Nesting uses arch:hasPart. The extension declares no parent/child predicate. c4:hasContainer —
which every non-C4 child used to fall through to — says the child is a c4:Container, which a deployment
node is not; dct:isPartOf is already spoken for by folder and model membership, so reusing it would put
"sits in the Elements folder" and "is nested inside this cluster" under one predicate on one subject.
arch:hasPart is the published generic core composition property, transitive over arch:ModelConcept.
structurizr:instances is xsd:integer upstream, so a value that is not an integer — a workspace can
write "instances": "auto" — is left out rather than emitted as a literal the shapes would reject.
Two constructs are reported rather than emitted, and both are named out loud because a silent omission is indistinguishable from a workspace that had none:
- Software system instances.
structurizr:Deploymentconstrains its source toc4:Container, so emitting one would make correct input produce a graph that fails its own validation. Seetodo/UPSTREAM-ISSUES-meta.mdissue 7. - Relationships drawn between instances have their endpoints redirected onto the containers those
instances instantiate. Since an instance is a relationship here, leaving them alone would point
arch:sourceat astructurizr:Deploymentinstead of anarch:Element, which the core endpoint contract rejects. Which instance is lost; what talks to what is kept.
Container Instance and Software System Instance join the structural tag set. They were absent only
because instances never reached the emitter; publishing Container Instance as a schema:keywords value
would restate the structurizr:Deployment the instance is already emitted as.
Fixed — validate now loads the Structurizr ontology it was already validating against¶
structurizr-shapes was in the default shape set while the ontology it imports was not registered at all.
That combination validated less than it appeared to: structsh:StructurizrElementLabelShape targets
structurizr:DeploymentNode and reaches structurizr:InfrastructureNode only through
structurizr:InfrastructureNode rdfs:subClassOf structurizr:DeploymentNode, declared in the document
nothing loaded. The new structurizr asset is registered and included in the converter's default
ontologies, which is what makes the label rule apply to the deployment elements now being emitted.
Fixed — a Backstage element IRI now carries the whole entity reference¶
Every Backstage element IRI moves. element/{namespace}--{name} becomes
element/{kind}/{namespace}/{name}, lowercased:
before …/backstage/my-catalog/element/default--payments-service
after …/backstage/my-catalog/element/component/default/payments-service
The old id dropped the kind, and a Backstage name is unique only per kind within a namespace. So a
Component and the API it provides, named alike — the ordinary shape of a service and its published
interface, not a corner case — collided on one node. That node carried bs:Component and bs:API,
both spec.type vocabularies, spec.definition from the API, every relationship of both entities, and
two entries in the Elements folder. Nothing in the output said so, and no warning was logged.
The code called the omission deliberate: putting the kind in the id "would mint a second node for every
reference that spells it out". That reasoning does not hold — references are resolved to a normalised
triplet before anything is minted, so how an author wrote one is already irrelevant. The id is now
BackstageEntityRef.elementPath, which is the same key that resolution compares on with : rewritten
to /. One function mints and compares, so the two cannot drift apart again; that, rather than the
extra segment, is the actual fix.
Three path segments, not one. A Backstage name may contain -, _ and ., so
component--default--pay--ments cannot be split back into a triplet — a consumer cannot tell the
separator from a character of the name. Neither : nor / is legal in a kind, namespace or name, so
the path form is reversible and needs no percent-encoding, and the ontology's own bs:name note says as
much. Nesting is not new here: BPMN already mints element/StartEvent_1/messageEventDefinition for
anonymous composites, and ConceptCollisions matches /element/ as a substring, so depth is
immaterial to it.
Lowercased, whole. The specification compares references case-insensitively and mixed case is legal
in a name, so MyService and myservice are one entity and must not become two addresses. The case the
author typed survives in the retained literals — the same split PlantUML already makes when it
lower-cases a label into an ID and leaves skos:prefLabel alone.
This is an identity change under
ADR 0003.
A published Backstage graph needs re-conversion, and IdentityLock will report the move rather than
republishing silently.
Added — Backstage bs:kind and bs:entityRef, from ontology 0.4.0¶
https://meta.linked.archi/backstage/onto 0.4.0 completes entity identity with the two reserved root
fields that were previously implicit, and both are now emitted:
<…/element/component/default/payment-api>
bs:entityRef "Component:default/payment-api" ;
bs:kind "Component" ;
bs:namespace "default" ;
bs:name "payment-api" .
bs:kind is not a duplicate of rdf:type, and neither is derived from the other: rdf:type is how
this graph classifies the entity, a modelling decision an adopter may revise, while bs:kind is what
the descriptor said. It keeps the descriptor's title-case token — API, not the api of the IRI
segment. bssh:KindTypeAlignmentShape reports divergence at sh:Info, as a reconciliation aid rather
than a failure.
bs:entityRef is a dcterms:identifier subproperty and the form Backstage itself circulates — in the
catalog API's relations array, in spec.owner / spec.system / spec.dependsOn, in the
techdocs-entity annotation — which makes it the stable cross-source join key. bs:uid is not: the
catalog reassigns it when an unchanged file is unregistered and re-registered, so it identifies a
registration rather than the thing. bssh:EntityRefConsistencyShape checks the reference against the
other three fields at Violation severity.
One caveat, stated because the ontology's framing does not quite fit this input: it describes
bs:entityRef as read from the source rather than assembled, which is right for a lift of the catalog
API. An authored catalog-info.yaml has no entityRef field, so this converter composes it — meaning
EntityRefConsistencyShape validates our own arithmetic here and becomes real evidence only for an
API-based lift.
Fixed — every fat JAR lost RDF4J's built-in SPARQL functions¶
validate crashed on any shape carrying an sh:sparql constraint that used a string function:
Internal error while trying to validate SHACL Shape bssh:EntityRefConsistencyShape
Caused by: Unknown function 'http://www.w3.org/2005/xpath-functions#lower-case'
RDF4J finds SPARQL functions through the ServiceLoader, and both rdf4j-shacl and
rdf4j-queryalgebra-evaluation ship a
META-INF/services/org.eclipse.rdf4j.query.algebra.evaluation.function.Function. In the shadow JARs
one replaced the other, and the one that won was rdf4j-shacl's single entry — so all 61 built-in
functions, LCASE among them, were absent at runtime. Every one of the seven executable JARs was
affected; duplicatesStrategy = DuplicatesStrategy.INCLUDE alongside the existing
mergeServiceFiles() fixes all of them, and the merged file goes from 56 bytes to 4183.
Latent until now rather than new: no shape in any default set had exercised a SPARQL constraint, so
nothing had asked for a function. Emitting bs:entityRef gives
bssh:EntityRefConsistencyShape a target for the first time, which is what surfaced it. It reproduced
only through a fat JAR — on a plain classpath the two service files are separate resources and both
are read, which is why the test suite never saw it.
Added — docs/architecture/iri-path-shape.md, for graph consumers¶
What anything generating documents, files or routes from the graph needs to know about the IRI path:
that a local ID may span several segments and is not a filename, that byte limits are per segment rather
than per IRI, which converters percent-encode and which do not, and which triples to read instead of
parsing the path. Written for rdf2docs and every other file-per-resource consumer, since the Backstage
change above is the second place a nested element ID appears.
Added — LeanIX diagrams become views; an inventory export still does not¶
A LeanIX diagram is a view; a fact sheet export is not. The converter emitted no arch:View at all,
which was right for the inventory and wrong as a permanent state of affairs: LeanIX has diagrams, and
they were unreachable. Both halves of that are now settled — diagrams reach the graph, and the inventory
still emits nothing view-shaped.
leanix-pull pull-diagrams reads them:
leanix-pull pull-diagrams --subdomain acme -o models/leanix/diagrams.json --write-index
leanix2linkedarchi convert models/leanix/factsheets.json \
--diagrams-export models/leanix/diagrams.json \
--diagram-index models/leanix/diagram-index.yaml ...
bookmarkType — so diagrams are the one thing in a LeanIX workspace that GraphQL does not serve. They also
change on a different clock: a diagram moves when a person drags a box, so folding them into the inventory
export would put cosmetic churn in the same diff as inventory change, and would make two independently
failing pulls look like one artefact. It also means an installation wanting only the inventory never needs
a token permitted to read bookmarks — which are permissioned separately, and a 403 now says so by name
rather than leaving the operator to suspect the token.
Each diagram becomes an arch:View + arch:Diagram in the semantic graph, with an archvis:ArchNode per
fact sheet on its canvas in the views graph, pointing at the element the inventory describes. The view IRI
comes from the LeanIX diagram id, so retitling a diagram does not move a published address. The pay-off is
that "which diagrams show this application" is one notation-agnostic query — the same
archvis:archElement traversal that answers it for a BPMN process or an ArchiMate view.
Three things are deliberately not invented, and all three trace to the same fact: SAP documents the
bookmark envelope but publishes no schema for a diagram's state, saying only that it holds "the layout
and filters" and that its content varies by bookmark type.
- No connectors. A line between two boxes is only meaningful as a relationship, and nothing
recognisable in a layout says which of the relationships joining two fact sheets it stands for — two can
be joined by lmm:Requiring and lmm:DataUsage at once. An arrow nobody drew is worse than no arrow.
- No geometry. Nothing says where the coordinates are, and coordinates are the part of a layout that
moves whenever somebody nudges a box. Structurizr's views carry none either.
- No inferred viewpoint. The published documents rule it out: lmmvp:ApplicationPortfolio and
lmmvp:CapabilityMap declare identical arch:includesConcept sets, so no examination of a canvas can
tell them apart, and a Free Draw canvas says nothing about what was drawn on it.
So membership is the one thing claimed, and even that is discovered rather than parsed: a fact sheet
reference is recognised only when a key that is or ends in factSheetId/factSheetIds carries a value
shaped like a LeanIX id. Requiring both signals matters in both directions — the key alone would accept a
placeholder, and the shape alone would sweep up the node, edge and style ids a layout is full of, minting
view nodes for elements that do not exist. Because a heuristic that finds nothing looks exactly like an
empty diagram, the export records a stateCoverage per diagram naming the keys it did and did not use, the
pull warns per diagram, and the converter warns again with a count. --fail-on-empty turns "nothing found
anywhere" into exit 1 for a pipeline that should notice the format having moved.
--diagram-index is where a viewpoint is declared, keyed by LeanIX diagram id, with status: and
title: alongside. A second index rather than a field on the existing one, because that index is keyed by
file — one entry is one file is one model with at most one view — and a single diagram export holds every
diagram in the workspace. pull-diagrams --write-index scaffolds it with the published catalogue in a
comment and viewpoint: blank; left blank the view claims none, which is honest. A bare name is checked
against the published set and an unrecognised one refused rather than minted, on the rule that governs
every other term here: lmmvp: is a namespace this converter reads and does not own. An absolute IRI is
used as written, and --ns-viewpoints repoints bare names at a catalogue of your own.
--retain-state keeps the raw layout blob, off by default: it is a layout, so keeping it would put every
nudge of a box into the diff of a committed file. Switch it on to investigate a diagram whose contents were
not recognised, or to hold the material for a release that can read layouts.
The inventory side is unchanged and now pinned by a scenario, so "an inventory export emits no view" stays
a decision rather than becoming an oversight somebody helpfully fills in.
Fixed — a LeanIX v3 workspace no longer claims conformance to Meta Model v4¶
leanix2linkedarchi typed every export in lmm: and declared arch:modelConformsToMetamodel
lmmmm:LeanIXv4, whatever the workspace actually ran. For a tenant still on Meta Model v3 that was false on
both counts, and it failed loudly in the wrong place: UserGroup, Project and TechPlatform fell back to
lmm:FactSheet with an "unknown type" warning apiece, which reads as a defect in the export rather than as
the converter having been pointed at the wrong vocabulary.
--meta-model v3 | v4 chooses, defaulting to v4:
v4 (default) | v3 |
|---|---|---|
| Fact sheet classes | lmm:, 20 classes | lmm3:, 12 classes |
| Root on every fact sheet | lmm:FactSheet | lmm3:FactSheet |
| arch:modelConformsToMetamodel | lmmmm:LeanIXv4 | lmm3mm:LeanIXv3 |
| Relationship classes | the sixteen lmm: relationships | none — v3 publishes none |
Nothing is mapped across the two. UserGroup became Organization, Project became Initiative, Process
became a BusinessContext subtype and TechPlatform became Platform: renames with changed semantics, not
aliases. And Process and Product are classes in both documents with different definitions, which is why
the namespace is chosen once for the run rather than inferred per type.
A v3 run emits no LeanIX relationship class, because the v3 manifest publishes none — fact sheet classes
only, deliberately. The v4 classes are not borrowed: every endpoint shape names v4 classes at both ends, so an
lmm:Requiring between two lmm3:Applications would be a validation violation rather than a richer graph. A
v3 relationship carries arch:QualifiedRelationship, the workspace's own field names on dct:type, and
core's arch:hasQualifiedRelationship, and the run says so once for the run rather than once per relation.
--type-mapping remains the way to type them on purpose, and is the natural tool mid-migration: name the v4
class for each type already reworked, and the v3 root stays underneath it.
Two things stay as they are under v4, and both matter more than they look. The two records LeanIX writes per
edge still collapse to one relationship — the v4 relationship table is used as a de-duplication and
direction oracle whatever the meta model, since it asserts nothing and ignoring it would double every edge.
The relationship IRI keeps the published name for the same reason, so an edge holds the same IRI across a
v3 → v4 migration, which makes it the one part of the graph that need not move when a workspace upgrades.
And skos:broader still comes out of the hierarchy, being a SKOS statement rather than a LeanIX term.
Attributes still land on the lmm: terms under v3: status, completion, tags, subscriptions and lifecycle are
platform features rather than meta model versions, and lmm: is the only place they are declared. They attach
through arch:domainIncludes, which entails no type, so this does not quietly make a v3 fact sheet a v4 one.
They go unvalidated, though — LeanIXFactSheetAttributeShape targets lmm:FactSheet only and says so. Labels
are still validated, because LeanIXFactSheetLabelShape targets lmm3:FactSheet as well, which is what makes
a v3 export a supported input rather than an archived one. The asset set does not change: validate already
loads both ontologies.
Pointing the converter at the wrong version now names the right one instead of listing unknown types, in both
directions, and a type in neither version's class list is still offered the --type-mapping remedy — the
messages distinguish "wrong version" from "custom type" because the fixes are different. mapping-report
takes --meta-model too, and under v3 reports relations as n/a rather than FALLBACK: there is no class to
reach and no mapping is missing, so --fail-on-unmapped stays usable for a v3 workspace instead of always
failing.
The version is an option rather than a detection. Nothing in a GraphQL export states which meta model a
workspace runs, and the fact sheet types only suggest it — a v4 workspace that pulled no Organization looks
like neither — so guessing would sometimes be silently wrong where being told is never worse than a warning.
Added — LeanIX: workspace introspection, mapping coverage, and ownership¶
Three follow-ups to the LeanIX pipeline, all aimed at the same gap: the converter was verified against fixtures and the published vocabulary, but nothing had checked it against a real tenant's metamodel.
leanix-pull introspect reads the workspace's own GraphQL schema and writes a pull configuration from
it. Until now the field names in the built-in configuration were the documented ones rather than yours, and a
field that does not exist is a GraphQL validation error that fails the whole query for its type. Introspection
is the right source for three reasons: it is a standard part of the protocol rather than a LeanIX path that
could move, it is served by the same URL with the same credentials, and it describes exactly what the API will
accept — which is the thing that actually breaks. Fact sheet types are the object types implementing
BaseFactSheet; a schema naming that interface differently reports as much rather than writing an empty
configuration that would read as a workspace with no fact sheets. --save-schema and --schema make it
usable offline.
The generated configuration lists every relation field and no scalar field: relations become relationships and a missing one is a missing edge, while requesting every scalar would pull data nobody chose into a committed file. Available scalars go in a comment beside each type, minus the ones already requested at base level.
leanix2linkedarchi mapping-report says what the converter will do with it, without writing a graph. Per
fact sheet type and relation: published, mapped by --type-mapping, or falling back — and for each fall-back,
the key to map it under. It reads an export rather than a configuration, because the export is what gets
converted. --fail-on-unmapped exits 1, so a pipeline can refuse to publish a graph with untyped edges;
without it the command exits 0, since a fall-back is emitted and reachable rather than dropped.
--emit-stakeholders turns subscriptions into people. An arch:Stakeholder per subscriber, lmm:subscriber
linking to it, and arch:conceptOwner from the accountable subscriber — which the ontology names as its
natural source. Responsible keeps data accurate and Observer only follows it, so neither becomes an owner.
The pay-off is that "who owns this application" stops being a LeanIX question: the same query answers it for a
Backstage spec.owner or an ArchiMate assignment. One resource per person, keyed on the lower-cased email, so
Accountable@ and accountable@ are one stakeholder rather than two.
Opt-in, because it creates resources for people and the only identity an export carries is an email
address — which then appears in an IRI as well as in the lmm:subscriberEmail literal. Hashing the IRI would
hide nothing while making the resource unrecognisable, so it stays readable and the flag is the protection. A
subscription with no address gets no stakeholder and says so once.
arch:Stakeholder and arch:conceptOwner are now on LinkedArchiVocab, so the other converters can reach
them without redeclaring — Backstage's spec.owner is the obvious next candidate.
Both example projects now include LeanIX. example-architecture-project gains models/leanix/, a
convert-leanix.sh, a pull-leanix.sh with an introspect mode, a scheduled pull-leanix CI job in its own
fetch stage, and a mapping-report step before conversion so a reshaped workspace shows up in the job log
rather than as untyped edges in the graph. example-archi-graph gains a graph/leanix/ target, a
sources-index.yaml entry, and the lmm:, lmm3: and ap: prefixes — ap: because lifecycle stages are
Architecture Processes individuals, so any query over lifecycle needs it.
Added — LeanIX support, as two tools rather than one¶
A LeanIX workspace can now reach the graph. It arrives as two commands, and the separation is the design rather than a packaging detail:
leanix-pullreads fact sheets over the GraphQL API and writes a canonical export file. It produces no RDF.leanix2linkedarchiconverts that export, and never touches a network.
Why not one command: a pull fails on credentials, rate limits and a workspace someone reshaped, while a conversion fails on the data — bundled together, a token that expired at 03:00 looks like a broken graph. The workspace also changes continuously while the graph is published on a schedule, and a file between them is where that difference is allowed to exist. Reading a file is what makes a conversion reproducible six months later, and committing the export makes a change to the inventory arrive as a diff someone can read before it becomes triples. The same seam is what makes webhooks an addition rather than a rewrite.
Start with a subset. --type BusinessCapability,Application,Interface pulls those types and nothing
else; the built-in configuration covers the four out-of-the-box types an architecture graph usually
starts from, and a type nobody configured is still pullable with the base fields alone. Prefer a scope
closed over its own relations: a pull of applications alone leaves each link to a capability pointing at a
fact sheet nothing describes — supported, warned about, and joined up by a later wider pull, but worth
choosing deliberately.
Output is typed in the published SAP LeanIX Meta Model v4 vocabulary — lmm: at
https://meta.linked.archi/leanix/onto#, with lmm:FactSheet ⊑ arch:Element, twenty fact sheet classes,
sixteen relationships, and arch:modelConformsToMetamodel lmmmm:LeanIXv4. validate loads the published
leanix-shapes — whose LeanIXFactSheetLabelShape is this converter's naming rule — plus core-shapes, and
both the v4 and v3 ontologies, because the shapes target lmm3:FactSheet too. A fully-scoped export
conforms.
LeanIX's metamodel is configured per workspace, which no other source here is. Types can be renamed or
added, fields are defined per tenant, and a relation is a field whose name is derived from the type pair.
So the pull composes its queries from a configuration file, and the converter keys on the names the
workspace uses. A type the published ontology does not declare is not minted as lmm:Whatever — that
would put an undeclared term inside a namespace this converter reads rather than owns — so it keeps
lmm:FactSheet, carries its raw name on dct:type, and is reported with the --type-mapping key that
would give it a real class. print-query prints what would be sent, to be pasted into the workspace's own
GraphiQL — the only authority on whether a field exists. One query per type, so a field this workspace does
not have costs one type instead of the whole pull.
Attributes land on published terms. The ontology declares the attribute set every workspace has, so the
fact sheet id becomes skos:notation, names skos:prefLabel / skos:altLabel, timestamps dct:created /
dct:modified, the URL schema:url, and — new in the 2026-08-16 vocabulary — status becomes
lmm:factSheetStatus lmm:StatusActive, completion lmm:completion as an xsd:decimal, the lifecycle
ap:atLifecycleStage ap:Active reusing the Architecture Processes stage vocabulary, and tags and
subscriptions their own lmm:Tag and lmm:Subscription nodes. States are individuals, not strings: the
shapes constrain them with sh:in and say so, so an unrecognised token is reported rather than passed
through as a literal that would fail validation while looking like data.
Three of those are load-bearing rather than cosmetic, and all three are the shapes' doing: lmm:completion
is emitted from a BigDecimal because a Kotlin Double gives xsd:double and fails sh:datatype; a
lmm:LifecyclePhase is emitted only when both a stage and a start date resolve, because the shape requires
both; and validate now loads arch-processes as well as leanix, leanix-v3 and core, because the
lifecycle values are its individuals and a sh:in set the graph cannot resolve validates nothing while
reporting nothing.
Workspace-defined fields are the one group still needing --ns-vocab, and the ontology is explicit that
they get no term: they are defined per tenant, so there is no set to publish.
The converter's own mapping is now validatable, which it was not before. Each of the sixteen
relationships has a published endpoint shape, so a --type-mapping entry pointing a relation at the wrong
class fails validation by name — "A lmm:Provision must run from an ITComponent." — where it previously
validated clean. One consequence to expect: a dangling endpoint in a scope-limited pull now reports twice,
once from core and once naming the type that failed to arrive, which is a deliberate upstream decision and the
second message is the useful one.
One rule the shapes hand back to the converter. A LeanIX hierarchy never crosses fact sheet types, but no
shape can check it: the two ends need not share a class — an Organization may parent an
OrganizationUnit — so the rule is about the workspace's own type token, which only the export carries. The
converter warns, naming both types, and emits the edge as declared.
All seven requests in UPSTREAM-REQUEST-leanix-attributes.md closed upstream on 2026-08-16; the file is kept as a record of what shipped and how it differed from what was asked. The only fallback path left is a relation a workspace invented, which no published vocabulary could cover.
See the LeanIX converter, pulling a workspace and the export format.
Fixed — the download sources now publish the commit the artifacts were built from¶
Provenance (ADR 0008) has the image tell the converter its own source revision at build time, via the
SOURCE_REVISION and SOURCE_URL build args. That works for a pipeline building from this repository,
where CI_COMMIT_SHA is the converter commit. It did not work for anyone packaging published
JARs into their own image: neither the Pages downloads directory nor the Package Registry published a
commit, VERSION was the only build-identifying file served, and a version cannot stand in for a
revision — every default-branch build publishes the same -SNAPSHOT version, and the revision is
consumed as a bare sha to compose a commit URL from. Downstream fetch scripts were left either
guessing or recording unknown, and every graph produced by such an image named no converter source
commit.
Both publishing jobs now write two one-line files next to the artifacts — SOURCE_REVISION (full sha)
and SOURCE_URL (the project it lives in), the two halves ImageIdentity needs to compose a commit
URL — in the Pages downloads directory and under each Package Registry version path. One value per
file, so reading one is $(curl -fsSL …) with no parser and no sourcing of remote shell.
Pages is still overwritten in place, so a fetch that straddles a pipeline can pair JARs with the next build's sha; a tagged registry version is immutable and is the source to use when the claim has to be exact. See Which build is this?.
Changed — Backstage converter updated for the 0.3.0 ontology and shapes¶
https://meta.linked.archi/backstage/onto and .../backstage/shapes published a substantial revision
(0.2.0 → 0.3.0): well-known relations completed, spec.type and spec.lifecycle promoted from plain
string literals to named individuals of published vocabularies, entity identity and profile fields
added, and metadata.links / well-known annotations given a home. The converter is updated to match:
- New relationship types, derived the same way
ownedBy/partOfSystemalways were:spec.dependsOn→bs:Dependency(Resource-targeted edges also getbs:ResourceUsage, the narrower readingusesResourceused to be the whole story),spec.subcomponentOf→bs:ComponentComposition,spec.subdomainOf→bs:DomainHierarchy,spec.parent/spec.children→bs:GroupParentage,spec.members→bs:GroupMembership, and the one well-known annotation whose value is an entity ref,backstage.io/techdocs-entity, →bs:TechDocsDelegation(withbackstage.io/techdocs-entity-pathriding asbs:techdocsEntityPathon that edge, since the path means nothing without the delegation).spec.membersandspec.childrenare Backstage's reverse-direction fields and are normalised onto the ontology's one canonical direction per pair, source and target swapped, rather than minting inverse properties. spec.typeandspec.lifecycleare minted as named individuals, not plain string literals —bs:ComponentType,bs:APIType,bs:ResourceType,bs:GroupType,bs:SystemType,bs:DomainType,bs:LifecycleState, and (where a source populates it)bs:ApiVisibility. Every one of these vocabularies is documented as open to organisation-specific values, so an unrecognised token still mints an individual — typed as the vocabulary's class sosh:classshapes are satisfied — rather than being dropped. The pre-0.3.0bs:lifecyclestring property is retired; it is now flagged bybssh:DeprecatedLifecyclePropertyShapeif anything still emits it.- Entity identity and profile fields now emitted alongside
skos:prefLabel/skos:notation, per the retained-source-data pattern every converter follows:bs:name,bs:namespace,bs:title,bs:uid,bs:etag,bs:tag(one permetadata.tagsentry),spec.profile'sbs:displayName/bs:email/bs:pictureonUserandGroup, andspec.definitiononAPI. metadata.linksemitted asbs:Linkblank nodes, reached from the entity viabs:hasLink—bs:linkUrl,bs:linkTitle,bs:linkIcon,bs:linkType.- Well-known catalog-provenance annotations lifted to named
bs:properties —backstage.io/managed-by-location,backstage.io/orphan(string"true"→xsd:boolean, absence stays absence, neverfalse),backstage.io/source-location, the GitHub/GitLab/LDAP/Microsoft Graph identity annotations underbs:externalIdentifier, and the rest — seeBackstageWellKnown.SIMPLE_ANNOTATION_PROPERTIESin the converter source for the full key list. Unrecognised annotation keys are left unmapped, since the descriptor format leaves that vocabulary open-ended.
Existing element and relationship IRIs are unchanged — this only changes what is asserted about them,
not their identity. See docs/converters/backstage.md for the current field-to-ontology mapping.
Fixed — ArchiMate names without xml:lang produced an untagged skos:prefLabel¶
Every other converter (PlantUML, BPMN, Structurizr, Backstage) exposes --label-language (default en)
and falls back to it whenever a name has no language tag, because sh:datatype rdf:langString on the
naming shapes rejects a plain xsd:string. The ArchiMate converter had no such option: skos:prefLabel
and skos:definition were minted straight from the Exchange XML's own xml:lang on <name> /
<documentation>, with no fallback, so a source file that omitted it — as opposed to the converter's own
test fixtures, which all declare xml:lang="en" — emitted an untagged literal and failed the naming
shape.
--label-language now exists here too (default en) and is applied only when the source element carries
no xml:lang; an explicit xml:lang in the Exchange XML is unchanged and still wins.
Changed — a notation's own keyword is a label, not a type, and not a UML stereotype¶
uml:stereotype is gone, the second undeclared predicate retired this release. What the PlantUML keyword
says is now split three ways by what it actually means
(ADR 0007):
<…/element/Ui> a …, uml:Class ; schema:keywords "boundary" . # the class cannot say this
<…/element/Shape> a …, uml:Class ; uml:isAbstract true . # a real UML attribute
<…/element/Order> a …, uml:Class . # `class` says nothing more
The predicate was undeclared upstream — grep -rn stereotype --include="*.ttl" finds nothing — and its
name claimed something false, because UML does publish uml:Profile, uml:Stereotype and
uml:ProfileApplication. The values were never stereotypes: they are PlantUML's LeafType /
ParticipantType constants, the keyword the author typed. A real <<stereotype>> is not read by the
parser at all.
abstract→uml:isAbstract true. Published upstream, maps to UML 2.5.1 §9.2Classifier::isAbstract. The class isuml:Classeither way, so the old string was the sole record and "which classes are abstract" required knowing PlantUML's keyword vocabulary.- What outlives its class →
schema:keywords.boundaryandcontrolboth resolve touml:Class, so the keyword is the only record of the robustness stereotype. Published, dereferenceable, and honestly a label — it makes no type claim. Chosen over inventing anarch:term because there is no Linked.Archi property for this and no C4 tag term either. - What restates its class → dropped. Derived rather than tabulated: publish the keyword only when it
does not name the class the element was typed with. A
--type-mappingoverride is followed for free.
Structurizr tags are emitted for the first time. They were parsed into C4Element.tags and discarded,
so a Database tag an author wrote existed in the workspace and nowhere in the graph — and views filter
on tags, so this was the data a FILTERED view is defined by. Structurizr's own structural tags
(Element, Container, Software System, …) are dropped; tags: "Element,Container,Database" publishes
Database. Same predicate as PlantUML's surviving keyword, which makes "everything any notation tagged as
a database" one query.
ArchiMate specialization is deliberately not mapped onto this. It is a metamodel-level specialization that can carry attributes, not a label, and the converter does not read it yet.
Fixed — three element-typing defects with one root cause¶
PlantUML has three entity-diagram classes and distinguishes far more than three things by the
USymbol it draws rather than by LeafType. Reading the leaf type alone produced:
| Source | Was | Now |
|---|---|---|
actor User (use case diagram) |
uml:Component |
uml:Actor |
node Server |
uml:Component |
uml:Node |
artifact App |
uml:Component |
uml:Artifact |
database Store |
uml:Component |
uml:Node + schema:keywords "database" |
actor / database / participant in a sequence diagram |
uml:Actor / uml:Node / uml:Lifeline |
all uml:Lifeline |
Two things worth calling out. uml:Actor was never emitted from any diagram — the mapping entry was
reachable only from a sequence participant, and those are lifelines. And a sequence diagram produced three
unrelated metaclasses on one lifeline axis, one of them uml:Node, a deployment metaclass describing a
message participant. UML 2.5.1 §17.3: every participant in an Interaction is a Lifeline. These are not
two granularities a consumer could reconcile — uml:Lifeline is a uml:NamedElement and uml:Actor is a
Classifier, so they sit on different branches and satisfy different shapes.
Only sequence diagrams are overridden; outside an Interaction actor genuinely is a uml:Actor. What a
lifeline stands for is uml:represents in UML, and it is not emitted — its range is a
ConnectableElement the converter does not synthesise, and fabricating one to hold a shape choice would
be worse than recording the shape as a keyword.
uml:Actor, uml:Device and uml:Artifact are now reachable. Element IRIs are unchanged — identity
is name-derived, so nothing is re-addressed; only rdf:type moves. The USymbol table now serves both
the element's keyword and ADR 0006's deployment-versus-component diagram detection, which were two
substring lists that could disagree.
Changed — a view says what kind of diagram it is, and a model says what language it is in¶
uml:diagramType "SEQUENCE" is gone. In its place every converter now emits the two conformance
properties the core ontology already declared, pointing at vocabularies already published upstream:
<…/plantuml/checkout/view/checkout-sequence>
a arch:View, arch:Diagram ;
arch:viewConformsToViewpoint umlvp:SequenceDiagram .
<…/plantuml/checkout>
a arch:Model ;
arch:modelConformsToMetamodel <https://meta.linked.archi/uml/metamodel#UML2> .
The old predicate was undeclared — grep -rn diagramType across linked-archi-meta returns nothing —
so the converter was minting a term into the published uml: namespace that UML does not define, no
shape could constrain it, and its value was a Kotlin enum name as a plain string. Meanwhile
uml-viewpoints.ttl had published all fourteen UML diagram types as arch:Viewpoint individuals the
whole time. Reasoning and the three rejected alternatives are in
ADR 0006.
What each converter did before, and does now:
| Before | Now | |
|---|---|---|
| PlantUML | uml:diagramType "SEQUENCE" (undeclared) |
umlvp:SequenceDiagram |
| Structurizr | C4ViewType parsed, then discarded |
c4vp:ContainerDiagram and the rest |
| BPMN | never asked | bpmnvp:ProcessFlow / CollaborationDiagram / ChoreographyDiagram |
| ArchiMate | viewpoint= never parsed, though every export carries it |
amvp:ApplicationCooperation and the rest |
| Backstage | no views | metamodel link only |
| every one | no record of the modelling language | arch:modelConformsToMetamodel |
A merged graph can now answer "every C4 model" or "every sequence diagram, whatever produced it" with a
one-hop query and no reasoning. It cannot yet answer "every ArchiMate model regardless of version" —
that needs one shared arch:Framework IRI per language upstream, filed as issue 6 in
todo/UPSTREAM-ISSUES-meta.md.
Four details worth knowing:
- An author's declaration wins. Both annotation routes already wrote
arch:viewConformsToViewpoint— the indexlinks:block and'!la-view-link— and a declared value now suppresses the detected one. Not a nicety: the shapes allowsh:maxCount 1, so emitting both would turn an author's annotation into a validation failure. UNKNOWNemits nothing."UNKNOWN"was a value that looked like a value and meant "the detector did not recognise this file". Absence says the same thing and cannot be matched on.- ArchiMate 2.x viewpoint names are dropped, not translated.
Introductory,Business Function,Actor Co-operationand the rest have no mechanical mapping to ArchiMate 3.2 andIntroductoryhas no successor at all, so guessing one would present the converter's judgement as the author's statement. Reported in one aggregated line per run rather than one per view. - New assets registered. The five viewpoint catalogues and seven metamodel manifests are now in
PublishedAssets, none of them in a default shape set.archimate3-viewpoint-shapesis registered and opt-in: it targetsarch:Viewfrom a notation namespace, so loading it applies ArchiMate's viewpoint policy to every notation in a merged graph.
New option per converter for the viewpoint namespace — --ns-uml-viewpoints and equivalents — kept
separate from the ontology namespace because classes and viewpoints are two published documents.
Fixed — a component, use case or deployment diagram was reported as a class diagram¶
Found while pinning the above. PlantUmlParser.detectEntityDiagramType matched substrings against
PlantUML's own diagram class simple names for Component, UseCase, Deploy and Activity — and
none of those strings appears in any PlantUML diagram class name. PlantUML serves all three from one
DescriptionDiagram, so all four branches were dead code and everything that was not a class or state
diagram fell through else -> DiagramType.CLASS:
| Source | Before | Now |
|---|---|---|
[Payments] --> [Ledger] |
CLASS |
COMPONENT |
actor User / User --> (Checkout) |
CLASS |
USECASE |
node Server / artifact App |
CLASS |
DEPLOYMENT |
start / :Do work; / stop |
CLASS |
UNKNOWN |
Survivable while the kind was a string nobody read; not survivable as a conformance claim SHACL acts on
and a consumer believes. A DescriptionDiagram is now classified by what it contains — a use case leaf,
then a node/artifact/device symbol, then a description leaf — and the else fallback is UNKNOWN,
because a diagram class the parser does not recognise is not evidence of a class diagram. Deployment is
distinguishable only by PlantUML's USymbol, which is the same coupling that caused the original
failure, so every diagram kind is now pinned by a test that fails on a library rename rather than
quietly relabelling.
Fixed — a BPMN label is a label node in its view, not a view-less ArchNode outside the namespace¶
BPMN records where a label is drawn separately from where its owner is drawn, because a tool lets an
author drag the text off the shape. BPMNShape-label and BPMNEdge-label carry that, and both are
composite properties — so a BPMNLabel sits one level below the plane's planeElement list.
The DI walk enumerated only that list. A label was therefore never entered into the remap table at all, and three defects followed from the single omission:
| Before | Now | |
|---|---|---|
| Address | {base}#Label_Start — outside the model namespace |
…/view/{viewId}/node/Label_Start |
| View | absent | archvis:view → the enclosing view |
| Type | archvis:ArchNode |
archvis:LabelNode |
Only the middle one was reported, by the self-check added earlier in this release:
[WARN] … produced 447 structurally incomplete resource(s) — this is a converter defect, please report
it: ArchNode StartEvent_1_di/BPMNLabel missing archvis:view; … (and 437 more)
The address defect is the same class as the lane flowNodeRef one fixed earlier in this release, and
label nodes were the only IRIs the converter still minted outside the model namespace. The typing
defect mattered because archvis:ArchNode is the domain of archvis:archElement, so it asserted the
label depicted a model element — while never carrying the archvis:archElement that would have made
the assertion true. archvis:LabelNode is the published class for a node that positions text.
<…/view/BPMNDiagram_1/node/Shape_Start> a archvis:ArchNode ;
archvis:archElement <…/element/StartEvent_1> ;
archvis:label <…/view/BPMNDiagram_1/node/Label_Start> .
<…/view/BPMNDiagram_1/node/Label_Start> a archvis:LabelNode , bpmndi:BPMNLabel ;
archvis:view <…/view/BPMNDiagram_1> ;
archvis:bounds-x 82.0 ; archvis:bounds-y 140.0 .
The node is kept rather than dropped. Dropping it was the other candidate fix, on the grounds that a label has no independent identity, and it would have lost geometry nothing else records — the label above sits below and wider than the 36×36 start event it belongs to.
An edge's label gets no archvis:label back-link, deliberately. archvis:label is declared
rdfs:domain arch-vis:Node, and core-vis makes Link a sibling of Node under DiagElement rather
than a subclass, so asserting it on an edge would entail that every labelled sequence flow is a Node.
Nothing would have caught that — core-vis declares no disjointness and publishes no shapes — which is
the reason not to emit it rather than a reason to. Edge labels stay reachable through the
notation-native bpmndi:label, emitted for both carriers; an aligned route needs the domain widened to
DiagElement upstream.
Two things this exposed, both worth stating plainly:
- SHACL was never going to catch this. The defective graph conformed to the published shapes, and so
does the fixed one.
core-vispublishes no shapes, so the diagram layer is unvalidated. The converter's own self-check was the only thing that found it, which is the case for wiring it in. - The shape's node was always correct, so every existing assertion about shapes passed. No fixture carried a label until now, which is why a defect affecting every diagram in a 20-diagram corpus was invisible to a green test suite.
Fixed — a synthesised relationship IRI no longer outgrows a filename¶
Every PlantUML relationship IRI changes with this entry. See the migration note below.
Deriving a relationship id from its label fixed identity and broke length. The id composes two endpoint ids and the label, all three author-supplied text with no limit, and while an IRI has no length limit the things built from one do: a path segment becomes a filename wherever a graph is written out as one document per resource, where 255 bytes is the ceiling on every mainstream filesystem.
It is reached by ordinary practice, not by pathological input. An author names a participant after the URL of its API contract — which is what you do when the participant is an API with a published specification — and two of those plus a label runs well past the limit:
https_git.example.org_group-one_spec-registry_alpha-api-contract_-_blob_main_openapi.yaml_alpha-api
__ https_git.example.org_group-two_platform_beta-proxy_-_blob_develop_beta-2.0.yml_beta-endpoint
__ wait_for_confirmation
On one corpus, relationship segments over the limit went from 16 to 116 with the label-derived scheme and the longest from 314 to 644 bytes, and the document generator aborted on every one of them. So a pre-existing problem became seven times more common.
The composed name is now called the notation, and the id is that notation truncated to 180 characters with a 20-character SHA-256 digest of the whole of it appended:
<…/relationship/ClientApp__OrderService__create_order__60f79e38d012a705ddc3>
skos:notation "ClientApp__OrderService__create_order" .
Identity is untouched. The digest is taken over exactly the string that used to be the id, so the same arrow drawn in two views still mints one IRI and merges, a call still separates from a reply, and reordering still renumbers nothing. Only the length changed.
skos:notation used to repeat the local name, which said nothing. It now carries the untruncated
notation, which makes it the way back from an id to the arrow — read it rather than the IRI when you
need the endpoints and label, because past 180 characters the IRI holds only a prefix.
Three choices worth stating, because each had a plausible alternative:
- Every id is hashed, not only the over-long ones. Leaving short ids alone would have kept most IRIs byte-identical and changed only the ones that break. Rejected because it makes the scheme depend on the length, and that boundary is invisible — lengthening a label from 178 to 182 characters would silently move a published IRI from one scheme to the other. One unconditional rule costs a single declared change now instead of an unbounded number of undeclared ones later.
- A digest, not a counter. A counter is shorter but assigned by order of appearance, so inserting an arrow would renumber the rest. A digest is a function of content alone, which is what lets two views agree without coordinating.
- A readable prefix is kept. A bare digest would be shorter and wholly opaque. The first 180 characters cost nothing against the limit and keep an id greppable back to its source.
Known residual: element ids are still unbounded. An element id is a single slugified display name rather than a composition, so no element segment exceeded the limit on the corpus, but nothing enforces it. Bounding it would re-address every element in every model for a problem not yet observed.
Migration: as with the label-derived scheme itself, this is a republish. IdentityLock records model and
view ids only, so nothing detects a relationship id change and --allow-identity-change has nothing to
authorise. Consumers holding a …/relationship/… IRI must re-resolve; querying by skos:notation is
stable across both changes.
Changed — two views spelling one label differently are reconciled, not refused¶
The previous entry on cross-view collisions made any label disagreement fatal. That was too strict in a way that made the check unusable, and this corrects it.
A relationship id is derived from the label through a normalisation that folds case, punctuation and
markup — deliberately, so that an edited capital does not move a published address. The collision check
then compared the raw skos:prefLabel. So every difference the id derivation was built to ignore was
a difference the check treated as a conflict:
[ERROR] Conversion failed: Identity collision in model 'order-domain':
<…/relationship/Adapter__Storage__wait_for_confirmed_or_denied> is claimed by two views:
'flow-ok.puml' says Wait for **CONFIRMED** or DENIED, 'flow-fail.puml' says
Wait for CONFIRMED or **DENIED** — the labels disagree.
Both views mean one message. They differ in where the emphasis markup sits.
On a 153-view corpus: 2,271 relationship IRIs, 735 legitimately claimed by more than one view, and 13 aborting collisions — every one cosmetic. Eight differed only in letter case, four in optional parentheses, one in markup placement. None were two genuinely different messages. Because the error was fatal it stopped at the first, so they were found one failed run at a time.
An identifier that folds case cannot then demand case agreement. The check now compares labels after the
same normalisation the id derivation applies — one function, LabelIdentity.canonical, used by both,
so the two cannot drift apart again — and splits the outcome in two:
- Same normalised form, different spelling. Not a conflict. The first spelling seen is kept, the
other dropped, and a warning names both files and both spellings. Dropping one is necessary rather
than tidy: both views emitted a
skos:prefLabelfor one IRI, and two of those in one language is not a legal SKOS resource, so leaving them would move the defect from a conversion error into an invalid graph. - Genuinely different. Still fatal, as before — different
rdf:typesets, or labels that are different words rather than a different spelling of the same words.
[WARN] <…/element/Task_1> is labelled 'Ship order' in 'a.bpmn' and 'ship order' in 'b.bpmn'. Both mean
the same thing — the id folds case and punctuation — so 'Ship order' was kept and the other dropped,
because a resource can carry only one skos:prefLabel per language. Spell it the same way in both sources
to choose for yourself.
Notation-native labels are reduced the same way. BPMN carries bpmn:name rather than skos:prefLabel,
and reducing only the aligned predicate would leave the duplicate one predicate along.
First-seen wins, so the outcome depends on input order. That is deliberate: the alternative is a rule nobody asked for, such as the longest string, which is equally arbitrary and harder to predict. Every drop is reported with both source files, so the real fix — make the sources agree — stays visible.
A corpus that wants strict agreement between views cannot currently ask for it; a
--fail-on-label-variance switch is noted as an open question in
ADR 0005.
Added — a PlantUML sequence arrow keeps its uml:messageSort¶
->, ->> and the dotted reply all became a bare uml:Message. The parser had already told them
apart — RelType.SEQUENCE_CALL, SEQUENCE_ASYNC, SEQUENCE_RETURN — and emission threw the
distinction away, so a sequence diagram converted to a graph in which every arrow looked alike.
uml:messageSort is where UML 2.5.1 §17.4 records it, and the values are published individuals in
uml/onto:
<…/relationship/rel-1> a uml:Message ;
uml:messageSort <https://meta.linked.archi/uml/onto#SynchCall> .
| PlantUML | uml:messageSort |
|---|---|
A -> B |
SynchCall — UML's own default for an unqualified message |
A ->> B |
AsynchCall |
A --> B (dotted) |
Reply |
asynchCall rather than asynchSignal for ->> because PlantUML does not distinguish a call from
a signal, and a call is the commoner reading of an async arrow between participants.
This also settles umlsh:MessageShape, which the redeployed uml/shapes added: it requires exactly
one messageSort per uml:Message, so out/plantuml-sequence.trig reported 11 violations, one per
message. It now emits 6 SynchCall and 5 Reply and conforms. Emitting the value was the better
fix of the two available — the alternative was arguing the constraint down, since messageSort
carries a spec default and defaulted attributes are exactly what the BPMN shapes strip minCount
from. Here the information genuinely existed and was being discarded.
The values arrive with uml/onto, which every validate run already loads as a supporting
ontology, so nothing extra is fetched to check them. See the entry below on the retired
uml/reference-data namespace for why they are there rather than in an asset of their own.
Fixed — uml:messageSort values pointed at a namespace that no longer defines them¶
Every sequence arrow emitted <https://meta.linked.archi/uml/reference-data#SynchCall> and the two
other sorts. Upstream UML-DD-12 has since declared each UML Enumeration an owl:Class closed with
owl:oneOf, moved its values into uml/onto as owl:NamedIndividuals, and retired the
uml/reference-data namespace. /uml/reference-data now 301s to /uml/onto, but redirecting a
document does not make its fragment IRIs mean anything: reference-data#SynchCall and
onto#SynchCall are different RDF terms, and only the second is defined.
So the converter was emitting a term nothing declared. The same release tightened
umlsh:MessageShape from a cardinality check to sh:class uml:MessageSort plus
sh:nodeKind sh:IRI, which turned that from a quiet dangling reference into a failure — one
sh:ClassConstraintComponent violation per message, on output that had been conforming:
sh:focusNode <…/plantuml/single/relationship/ClientApp__OrderService> ;
sh:value <https://meta.linked.archi/uml/reference-data#SynchCall> ;
sh:sourceConstraintComponent sh:ClassConstraintComponent ;
The values are now minted in UML_NS, alongside the classes. They stay outside --type-mapping:
that option remaps a vocabulary onto a consumer's own classes, and these are individuals whose
identity is the point — a remapped value would not be a member of uml:MessageSort.
The uml-reference-data short name is removed from the asset registry rather than repointed.
Keeping it would give one document two names, so --ontology uml uml-reference-data would fetch
uml/onto twice. Anything passing that name should pass uml, which was already loading the same
file.
Fixed — a PlantUML state diagram converts to a state machine, not to classes and associations¶
Three things were wrong at once, and only one of them was visible.
Every simple state was a uml:Class. A state is a leaf, and the leaf mapping had no branch for
one, so it fell through to class. Only composite states were right, because those are groups and
took the GroupType.STATE path instead. So a machine came out as a handful of classes with two states
among them, by accident of which ones happened to be composite.
Every transition was a uml:Association. PlantUML reports --> in a state diagram with no
decoration and a normal line — the same shape as a plain class-diagram association — so nothing about
the link distinguishes them. The diagram kind is the only thing that can, and the parser had known it
all along without using it.
[*] was dropped. It carries no label, so it was skipped when elements were collected, then
reappeared as a relationship endpoint through the emitter's fallback: an IRI with no rdf:type on it.
Only the third was reported. The first two were invisible because of each other — uml:Class is a
Classifier, so it satisfied the very association shapes the transitions were being checked against.
The six violations in the original report were all [*].
# before # after
element/Idle a uml:Class ; element/Idle a uml:State ;
element/start # untyped element/start a uml:Pseudostate ;
uml:pseudostateKind uml:InitialPseudostate ;
element/end # untyped element/end a uml:FinalState ;
relationship/… a uml:Association relationship/… a uml:Transition ;
uml:Transition is declared rdfs:subClassOf arch:QualifiedRelationship upstream, so retyping keeps
the core alignment and simply takes these nodes out of umlsh:AssociationShape's scope — which is how
the violations go away rather than being argued down.
The two ends of [*] are different things, and PlantUML says which without anything being
inferred from the direction of the transition: CIRCLE_START is an initial uml:Pseudostate, a
transient vertex the machine passes through, and CIRCLE_END is a uml:FinalState, a State the
machine rests in. Nested regions get their own pair, because PlantUML scopes the name to the region —
start and start_Running — so an inner machine's start does not merge with the outer one's.
uml:pseudostateKind is emitted because a bare uml:Pseudostate cannot be told from a join or a
choice. PlantUML can only express the initial kind, so that is the only value ever written.
Both vertices carry a skos:prefLabel of initial / final. They have to: uml:Pseudostate and
uml:FinalState are both uml:Vertex, hence uml:NamedElement, and umlsh:NamedElementShape
requires one. These are descriptions rather than identifiers — the id still comes from PlantUML's
region-scoped name.
Unlabelled transition ids read Idle__Running__transition, following the same rule as every other
relationship: the label when there is one, the kind when there is not.
Fixed — ArchiMate's --emit-core-triples and --validate-refs did the opposite of their names¶
Both were declared as plain booleans whose field already defaulted to true:
Picocli toggles a boolean flag in that shape, so passing the option flipped it to false. Asking
for core triples removed them. On Archisurance that is the difference between
and none of them — a graph with no alignment to the core ontology at all, which every core shape then
rejects. --validate-refs had the same shape and so silently turned reference validation off.
Both now carry defaultValue, negatable and fallbackValue, which is how every other converter
declares a boolean that defaults on. So --emit-core-triples is idempotent, and
--no-emit-core-triples is the way to turn it off:
arch:QualifiedRelationship |
|
|---|---|
| flag absent | 178 |
--emit-core-triples |
178 |
--no-emit-core-triples |
0 |
Added — every conversion checks its own output before writing it¶
ConversionVerifier has existed since the beginning and was called from tests only. It now runs
at the end of every convert, so the three things a converter is wholly answerable for are checked
in the process that built the triples:
- every
arch:QualifiedRelationshiphas anarch:sourceand anarch:target - every
archvis:ArchNodebelongs to a view - every
archvis:Linkhas both of its ends
Not a duplicate of SHACL. The shapes cover the same endpoint rule and much more, but they run later, against a merged graph, with the ontology loaded — by which point "which file caused this" is a search. This runs while the inputs are still nameable, and says so:
[WARN] orders.bpmn produced 2 structurally incomplete resource(s) — this is a converter defect,
please report it: Relationship DataOutputAssociation_1 missing arch:source; ...
Which is not hypothetical: that is the output of this check against the data-association defect fixed below, before it was fixed.
A warning, deliberately. These report a defect in the converter, and the author of the input cannot fix one — so failing would withhold output they can still use, over a problem only we can repair.
Two of the five checks in verify() are deliberately excluded, because each has honest causes:
checkNoGeneratedIris flags _gen_ IRIs, which BPMN mints legitimately for elements a file left
unidentified; checkAllDiagramsHaveNodes flags an empty diagram, which is the author's business.
Warning about either on every run would teach people to ignore the warning. Both remain available as
their own checks.
Wired into all five converters. In ArchiMate it sits inside the emitter rather than the command,
because that is where the model is: write builds the whole graph in a LinkedHashModel, indexes it,
and hands it to Rio.write in one go, so the check costs nothing that was not already materialised.
Fixed — a BPMN reference to the wrong thing, or to nothing, is reported¶
The CMOF declares what every BPMN cross-reference must point at — Lane::flowNodeRefs is a
FlowNode — and nothing in the XML schema enforces it. Neither kind of breakage was noticed.
A reference to the wrong kind of thing is now reported and kept:
[WARN] <Lane_1> flowNodeRefs points at <DataObjectReference_1>, which is a DataObjectReference
rather than a FlowNode. The reference is kept, but a consumer reading flowNodeRefs will find
something that cannot be one.
Kept because it does resolve and the author may have meant something by it. This is the case the shapes report as "not a FlowNode", far downstream and without naming the file.
A reference to something absent is now reported and dropped:
[WARN] <Lane_1> flowNodeRefs points at 'Task_Absent', which is not in this file. The reference was
dropped — keeping it would publish an edge to an IRI nothing defines.
Dropped because keeping it was worse. Every IRI in the notation layer is minted as {base}#{id}, and
the emitter re-mints them into the model namespace from a table built out of the document's
subjects — so a target that is never a subject was not in the table and the raw {base}#{id}
survived into the output. A dangling edge, in a namespace the model does not own, which the emitter's
own comment says the re-minting exists to prevent.
Both are warnings: they are defects in the file, its author can fix them, and a usable graph can still be written.
Only references read from the file are checked. The endpoint inferred for a data association
below is deliberately not among them — it points a DataAssociation's sourceRef at the enclosing
activity, where the CMOF declares ItemAwareElement. That is a knowing departure, and checking it
here would report our own inference as the author's mistake.
Fixed — a BPMN data association keeps both of its ends¶
dataOutputAssociation was published as an arch:QualifiedRelationship with no arch:source,
and dataInputAssociation with no arch:target.
A data association is the one BPMN relationship whose ends are not both written down. One is a child element; the other is understood from where the association sits:
<bpmn:userTask id="UserTask_1">
<bpmn:dataOutputAssociation id="DataOutputAssociation_1">
<bpmn:targetRef>DataObjectReference_1</bpmn:targetRef> <!-- stated -->
</bpmn:dataOutputAssociation> <!-- source: UserTask_1 -->
</bpmn:userTask>
Nothing in the metamodel expresses that, so the generic reference-resolving path read the child and
left the other end empty. Because bpmn/onto declares DataAssociation an
arch:QualifiedRelationship, and core-shapes#QualifiedRelationshipShape requires exactly one
arch:source and one arch:target, every data association was published as a relationship missing
an endpoint — reported by validation on a merged graph, long after the file that produced it was out
of reach.
The enclosing activity is now taken as the missing end: the source of a dataOutputAssociation, the
target of a dataInputAssociation.
Resolved in the parser rather than the emitter, because the containment is a parse fact — known at
the moment the association's element closes, and otherwise only recoverable by inverting
bpmn:dataOutputAssociations after the event. Doing it there also means the native BPMN layer states
what the specification leaves implicit, so the graph is complete before alignment rather than only
after it, and the emitter keeps needing no per-element knowledge.
Inferred only when the file is silent. The specification permits sourceRef on a
dataOutputAssociation, pointing at one of the activity's dataOutputs, which is more precise than
the activity; an end stated explicitly is left as written, since adding to it would give the
relationship two sources and break the same shape from the other side. A plain <bpmn:association>
carries no direction and is untouched.
ConversionVerifier.checkRelationshipsHaveEndpoints — which existed already and would have caught
this — now runs against a fixture that has data associations in it. It passed before because the
only BPMN fixture in the repository had none.
Changed — a PlantUML relationship is named after its label, not just its endpoints¶
Breaking: every published PlantUML relationship IRI changes. See ADR 0005.
A relationship ID was {source}__{target}, with _2, _3 … for repeats. The endpoints name a
pair, not an edge between them — two participants exchange many messages — and the suffix that was
supposed to separate them was counted per file while the IRI namespace is per model. Two
views of one model connecting the same pair therefore each restarted at the bare ID and collided:
# before — one resource, two messages
<…/plantuml/order-domain/relationship/ClientApp__OrderService> a uml:Message ;
skos:prefLabel "create order" , "order created" ;
uml:messageSort uml:SynchCall , uml:Reply .
The ID now carries what actually distinguishes one arrow from another, which is what the author wrote on it — the same rule PlantUML already used for elements, whose ID is the slugified display name:
# after — two resources
<…/relationship/ClientApp__OrderService__create_order> uml:messageSort uml:SynchCall .
<…/relationship/ClientApp__OrderService__order_created> uml:messageSort uml:Reply .
| Source | ID |
|---|---|
ClientApp -> OrderService: create order |
ClientApp__OrderService__create_order |
A --> B: owns |
A__B__owns |
A -> B unlabelled |
A__B__call |
A --> B unlabelled |
A__B__reply |
A ->> B unlabelled |
A__B__async |
A --> B unlabelled, class diagram |
A__B__association |
| arrows identical in all of the above | …, then …_2 |
The label is lower-cased into the ID and left alone in the graph, so fixing a capital does not move a
published address while skos:prefLabel keeps the author's text. Order of appearance survives only
for arrows a reader could not tell apart either — so inserting an arrow no longer renumbers the ones
after it.
Three things follow. One arrow drawn in two views is now one relationship with a link in each
view, which is the merge that scoping IDs to the view would have lost. Two views spelling one message
differently now fail the run: they agree on identity and disagree on the label, and SKOS allows
one skos:prefLabel per language, so the sources have to agree. And the rule applies to structural
relationships too — A --> B: owns and A --> B: manages were A__B and A__B_2.
Superseded, twice, later in this release. The composed ID above is now the
skos:notationand the IRI segment is that name bounded with a digest — see a synthesised relationship IRI no longer outgrows a filename. And a spelling difference no longer fails the run; it is reconciled with a warning — see two views spelling one label differently are reconciled, not refused.
The migration is a republish. IdentityLock records model and view IDs only, so nothing detects a
relationship ID change and --allow-identity-change has nothing to authorise; no owl:sameAs bridge
is emitted either, because the old IRI often denoted two messages at once and there is nothing single
to point it at.
Fixed — a relationship label is no longer wrapped in the punctuation of the list it was stored in¶
ClientApp -> OrderService: create order published skos:prefLabel "[create order]", brackets included, and an
unlabelled arrow published "[]". A PlantUML Display is a list of lines and its toString() is
the list's, so reading it as a string produced the list's punctuation — the same defect fixed for
diagram titles previously, never fixed for labels.
The second half mattered more than it looks: "[]" is not empty, so an unlabelled arrow counted as
labelled, which is why the new ID scheme above needed this fixed before it could fall back to the
arrow's kind.
Added — uml:messageKind distinguishes a lost message from an ordinary one¶
A ->x B and A -> B converted to identical triples. PlantUML's cross marks a message that was sent
and never arrived, and UML records that as messageKind rather than as a different sort — a lost
synchronous call is still a synchCall — so it is emitted alongside:
<…/relationship/A__B__ping> a uml:Message ;
uml:messageSort uml:SynchCall ;
uml:messageKind uml:Lost .
Only Lost is ever emitted. Complete is derived in UML from a message having both a sendEvent and
a receiveEvent, and this converter emits no uml:MessageEnd resources at all, so asserting it would
claim a fact about a structure that is not modelled. Found needs the mirror notation, a message
arriving from outside the diagram, which PlantUML writes with a gate ([-> B) — and the library
reports gates as MessageExo rather than Message, so the parser does not read them at all. Those
arrows are missing from the graph entirely, which is a gap in its own right.
umlsh:MessageShape already allowed this: messageKind carries sh:maxCount 1 with no minCount.
Changed — a colliding relationship id now fails the run, as a colliding element id already did¶
ElementCollisions is now ConceptCollisions, and it inspects /relationship/ subjects on the
same terms as /element/ ones: same IRI claimed by two views, disagreeing on label or type set,
run fails naming both files. It previously skipped them outright —
— which left relationships merging silently whenever two views claimed one id, so a wrong graph was published and only some of it was caught later, by SHACL, on a merged graph with no route back to the file responsible.
The relationship-identity change above removes the largest source of those collisions rather than merely reporting them, so what this check now guards is the cases identity cannot prevent:
- BPMN ids are tool-assigned and only unique per file.
Flow_1in two processes grouped into one model is two different flows on one IRI. Nothing in the id can tell them apart. - Two views disagreeing about the type. An
Associationin one and aDependencyin another between the same pair, or aClassand anInterfaceon one element id. - Two views spelling one PlantUML message differently.
create orderandCreate Orderagree on identity, by design, and disagree on the label — and SKOS allows oneskos:prefLabelper language, so this has to be refused rather than resolved by picking one.
The failure now says which of those it is. Advice that fits a type disagreement — give them distinct ids, or separate the models — is wrong for a spelling difference, where the answer is to make the sources agree, and the message would otherwise have sent authors to split one message into two.
This can surface pre-existing collisions. A corpus that has been converting cleanly may now fail, because the graphs it produced were already wrong.
The third bullet is superseded later in this release. Refusing a spelling difference was too strict to be usable — on a 153-view corpus it aborted 13 runs and every one was cosmetic. It is now reconciled with a warning; see two views spelling one label differently are reconciled, not refused. The first two bullets still fail the run.
Added — backstage-shapes is now a Backstage default, and every converter conforms¶
The two problems that held the Backstage shapes out of the default set are both resolved — the
Component-only SystemMembership source is corrected in the published backstage/shapes, and a
kind-prefixed entity ref now resolves to the entity it names — so the default set becomes
backstage-shapes + core-shapes. That restores a naming rule for Backstage output
(bssh:BackstageElementLabelShape) and adds the per-kind relation checks a catalog derives from its
spec fields.
With that, every converter conforms against its own default shape set on the playground output:
| Converter | Default shapes | Result | Coverage |
|---|---|---|---|
archimate2linkedarchi |
relationships, elements, core-shapes |
conforms | 15/82 |
bpmn2linkedarchi |
bpmn-shapes, bpmn-infra-shapes, core-shapes |
conforms | 11/139 |
plantuml2linkedarchi |
uml-shapes, core-shapes |
conforms | 5/26 |
structurizr2linkedarchi |
c4-shapes, structurizr-shapes, core-shapes |
conforms | 7/11 |
backstage2linkedarchi |
backstage-shapes, core-shapes |
conforms | 14/16 |
Twelve playground outputs across the five converters, .ttl and .trig alike, all conform.
They were regenerated in the process, which is what two of them needed anyway. out/bpmn-* predated
the label change and carried no skos:prefLabel, so they failed the rule that release added —
out/bpmn-full.ttl goes from 930 to 956 triples, the 26 mirrored labels. out/plantuml-sequence.trig
needed the messageSort entry above.
A multi-repo Backstage run still reports violations, and the shapes are right to. A
spec.owner: team-payments reference in one repository mints
{model}/element/default--team-payments inside that model's namespace, while the Group entity
declared in another repository is {other-model}/element/default--team-payments. Two IRIs for one
team, so bssh:OwnershipShape reports a target that is not a bs:Group — and merging the graphs
does not reconcile them, because neither IRI is the other.
This is distinct from the entity-ref fix below, which resolves the form a reference is written in within one model. This is the question of which model owns an entity that a repository references without declaring, and ADR 0001 places cross-repository uniqueness with the aggregating repository, so the converters do not decide it.
It is not a new failure: those references already failed
core-shapes#QualifiedRelationshipShape, which was in the Backstage default set before the
Backstage shapes were. out/backstage-multi-repo.trig reported 6 violations before this change and
reports 12 after, every one of them the same root cause. What changed is that the report now names
the problem (Ownership target must be a Group) rather than stating it generically. Single-catalog
runs — out/backstage.ttl, out/backstage-indexed.trig — conform.
Fixed — the ArchiMate playground fixture modelled deployment in a way ArchiMate does not permit¶
archimate2linkedarchi validate reported two archimate3:AssignmentShape violations on
playground/out/archimate.ttl, and they were the fixture's fault, not the shape's. It modelled
SystemSoftware (Kubernetes Cluster) --Assignment--> ApplicationComponent (Payment Gateway)
SystemSoftware (Kubernetes Cluster) --Assignment--> ApplicationComponent (Ledger Service)
ArchiMate does not permit an Assignment targeting an Application Component. Deployment goes through the
artifact: the system software is assigned the deployable, and the deployable realizes the component.
playground/archimate/payment-platform.xml now carries two Artifacts and says so:
SystemSoftware --Assignment--> Artifact (payment-gateway.jar) --Realization--> ApplicationComponent
SystemSoftware --Assignment--> Artifact (ledger-service.jar) --Realization--> ApplicationComponent
This was already fixed once. The identical correction was made to the test fixture
converter-archimate/src/test/resources/payment-platform.xml in an earlier release; the playground copy
of the same model was missed, so the two drifted. They are byte-identical again, which is the point —
the file the tests assert on and the file the playground publishes should not be able to disagree about
whether the model is valid.
The fixture grows from 9 elements / 8 relationships to 11 / 10, and playground/out/archimate.ttl and
.trig from 232 to 268 triples. ArchiMate output now conforms against its full default shape set,
coverage 15/82.
The rule was never the problem and still bites: a graph carrying the removed pair reports exactly one
AssignmentShape violation, while the artifact-based replacement passes.
ArchiSurance and ArchiMetal still report violations, for an unrelated reason. The other two
playground models are the Open Group's published case studies, and both are ArchiMate 2.1 exchange
files — archimate_v2p1.xsd, with 2.x-only element types (InfrastructureFunction,
InfrastructureService, Network) and 2.x-only relationship pairs
(ApplicationComponent --Assignment--> BusinessFunction, which 3.x replaced with realization from
application behavior to business behavior). Validating them against ArchiMate 3.2 shapes is a version
mismatch, not a defect in the shapes or in the models: 9 violations for ArchiSurance, 23 for ArchiMetal.
Left alone deliberately — they are somebody else's data, and editing it would remove the signal.
Fixed — a kind-prefixed Backstage entity ref now resolves to the entity it names¶
dependsOn: resource:default/payment-db minted element/resource--default--payment-db while the
entity itself was element/default--payment-db. The relationship pointed at a node nothing in the
graph described, so backstage2linkedarchi validate failed on the playground catalog with
core-shapes#QualifiedRelationshipShape — the converter's own sample output did not conform to the
converter's own default shapes.
The cause was string comparison. References were matched as written, against a key built from the
entity's kind: field, and Backstage lets the same reference be typed several ways: the kind may be
lowercase or capitalised, the namespace may be omitted, the kind may be omitted. resource:default/
payment-db and Resource:default/payment-db are one reference in two capitalisations, and the
lookup treated them as two entities.
References are now read as Backstage defines them — [<kind>:][<namespace>/]<name> — and resolved to
the kind/namespace/name triplet, compared case-insensitively as Backstage compares it. All four forms
below reach element/default--payment-db:
spec:
system: default/payment-gateway # namespace written out, kind from the field
dependsOn:
- resource:default/payment-db # fully qualified, lowercase kind
- Resource:default/payment-db # the same reference, kind as the catalog spells it
- resource:payment-db # namespace from the referring entity
The kind stays out of the minted id: an entity is identified by namespace and name within its kind, so including it would mint a second node for every reference that spelled it out.
A reference the model does not declare is now reported. It still mints the id that entity would have, so the two sides join if it arrives later, but the run says so once per reference rather than leaving an untyped node for SHACL to find:
[WARN] catalog-info.yaml: 'ownedBy' on System:default/payment-gateway names
Group:default/team-payments, which model 'payments-services' does not declare. The relationship
points at element/default--team-payments, a node nothing in this model describes. If that entity
lives in another catalog file, convert the files into one model.
The multi-repo playground run prints exactly that twice, and the warnings are correct: IRIs carry the
model id, and run-backstage-multi-repo.sh gives each file its own, so the payments catalog's
references to the platform catalog's team-payments and payments cannot resolve. Converting the two
files as one model — a directory input, or one --model-id — is what joins them. That was already
true; it was simply silent before.
spec.dependsOn is still read leniently, and this is the one place the converter is wider than the
format: Backstage requires a kind there, because a dependency may be a Component or a Resource and
neither is the default, while this converter assumes Resource. A bare name meant as a component
resolves to a resource that does not exist, and is reported as unresolved.
Fixed — bssh:SystemMembershipShape was narrower than the descriptor format (upstream)¶
The second reason backstage-shapes was not a default: the shape required the source of a
bs:SystemMembership to be a bs:Component, and Backstage declares spec.system on Component,
API and Resource. The playground catalog puts an API and a Resource in a system — ordinary
Backstage — and the shape reported both as violations.
Corrected on meta.linked.archi, in backstage/shapes and backstage/onto:
bssh:SystemMembershipShapeadmits a Component, an API or a Resource as the source.bssh:ResourceUsageShapeadmits a Component or a Resource as the source, sincespec.dependsOnis declared on Resource too. The target stays Resource: aspec.dependsOnnaming a Component is a different relationship, andbs:has no term for it.bs:partOfSystem/bs:SystemMembershipandbs:usesResource/bs:ResourceUsagecarry the matchingarch:domainIncludes.
Both rules still bite. A bs:Group as the source of either relationship, and a Component as the target
of a bs:ResourceUsage, are each reported.
Neither asset is versioned up for this, since nothing consumes them yet: the correction is a fix to 0.1.0 / 0.2.0 rather than a new revision of either.
backstage-shapes remains off the default set until the corrected shape is the published
document. Assets are fetched at runtime, and what is served today still carries the Component-only
source constraint, so a default including the shapes would fail against the live file. Read the served
document rather than its version to tell them apart:
curl -sH 'Accept: text/turtle' https://meta.linked.archi/backstage/shapes \
| grep 'SystemMembership source'
Against the corrected shapes the playground catalog conforms — --shapes backstage-shapes --shapes
core-shapes, 14/16 target classes covered — and against the published ones it reports the two
spec.system sources and nothing else, the converter defect being gone.
Fixed — BPMN no longer reports a missing label on every gateway, event and sequence flow¶
Validating BPMN output against the core shapes reported a missing-label violation for almost every
object in the file. On the two playground diagrams that was 41 violations, 39 of them false —
18 sequence flows, 3 exclusive gateways, 3 end events, 2 start events and a bpmn:Documentation
node, none of which BPMN requires to be named.
The cause was upstream, not in the converter. core-shapes#ElementLabelShape required a
skos:prefLabel on every arch:Element, and the BPMN alignment declared
bpmn:BaseElement rdfs:subClassOf arch:Element — the root of the entire BPMN metamodel, which under
rdfs:subClassOf reasoning reached all 130 subclassed BPMN classes. A rule written for
architecture elements therefore landed on every serialization artefact in the document.
Both have been corrected on meta.linked.archi and this release adopts them:
- The core label shapes are withdrawn. Requiring a label on every element is not sound across notations, because routing and connector constructs — BPMN gateways and events, ArchiMate junctions — are genuine elements their notations permit to be unnamed.
- The alignment is declared on the classes it describes in
bpmn/onto, on 11 roots reaching 49 classes rather than all 130. Connectors arearch:QualifiedRelationshipand notarch:Element;bpmn:Documentation,ItemDefinition,Expressionand the rest of the metamodel bookkeeping are neither. Element counts for a BPMN model drop accordingly, and a query forarch:Elementno longer returns wrappers. - Naming is stated by
bpmnsh:RequiredNameShapeinbpmn-shapes, over the 12 classes BPMN expects to be named — activities and their task subtypes, processes, collaborations, conversations, pools, lanes, data objects and stores, messages.
The same two diagrams now conform, and an unnamed serviceTask still reports exactly one violation
naming that task, so the rule is enforcing rather than passing vacuously.
Changed — BPMN emits both bpmn:name and skos:prefLabel, and does so by default¶
This changes the emitted RDF. --emit-skos-labels used to move the name: skos:prefLabel
replaced bpmn:name rather than accompanying it, and it was off by default — the only converter
in the suite where it was, against a shared core default of on.
Both halves were wrong against the published converter contract, which asks for the two together:
<…/element/Task_CheckInventory>
a bpmn:ServiceTask ;
bpmn:name "Check Inventory" ; # source attribute, verbatim
skos:prefLabel "Check Inventory"@en . # canonical label, language-tagged
They are not redundant. skos:prefLabel is what every Linked.Archi consumer reads — documentation
generators, palette builders, deliverable templates, SPARQL reporting, SHACL validation — and it must
carry a language tag. bpmn:name is the round-trip record of the XML attribute, a plain
xsd:string, and it is the only name-bearing property on the BPMN classes that are not
architecture elements, so dropping it loses those names entirely.
Emitting both makes the option purely additive, which is what makes defaulting it to on safe: no
triple that used to be emitted stops being emitted. Output grows by one label per named element.
--no-emit-skos-labels withholds the label, which leaves the output non-conforming; it is kept for
round-trip work rather than as a supported publishing mode.
bpmn:id / skos:notation is unchanged — no shape requires a notation, so there is no contract to
meet there.
Added — every converter validates against the shapes for the notation it emits¶
Withdrawing the core label shapes left four converters with no naming rule at all:
plantuml, structurizr and backstage defaulted to core-shapes alone, and archimate
defaulted to relationships alone, so nothing checked that an element had a label. Each notation
publishes its own rule; none of them were wired in.
They are now, and the default set for each converter is the shapes for the notation it actually
emits plus core-shapes for the contracts that are not notation-specific:
| Converter | Notation emitted | Default shapes | Naming rule now enforced |
|---|---|---|---|
archimate2linkedarchi |
ArchiMate 3.2 | relationships, elements, core-shapes |
amelsh:ArchiMateElementShape |
plantuml2linkedarchi |
UML 2.5.1 | uml-shapes, core-shapes |
umlsh:NamedElementShape |
structurizr2linkedarchi |
C4 | c4-shapes, structurizr-shapes, core-shapes |
c4sh:C4ElementLabelShape, structsh:StructurizrElementLabelShape |
backstage2linkedarchi |
Backstage | backstage-shapes, core-shapes |
bssh:BackstageElementLabelShape |
The notation each converter emits is worth stating plainly, because two of them are not obvious.
PlantUML is a syntax, not a metamodel: a stereotype maps to uml:Class, uml:Interface,
uml:Association and the rest via UmlTypeDefaults, so UML is what its output is typed in and
uml-shapes is what applies. Structurizr emits C4: people, software systems, containers and
components map to c4:Person, c4:SoftwareSystem, c4:Container, c4:Component, every
relationship to c4:Using.
Newly registered, so all are selectable by name and appear in --list-assets: uml,
uml-shapes, c4, c4-shapes, structurizr-shapes, backstage, backstage-shapes.
structurizr-shapes resolves to c4/structurizr-shapes, not structurizr/shapes — the
deployment shapes live in the C4 namespace family because Structurizr is a C4 tool, and only the
short name says Structurizr. Requesting structurizr/shapes 404s; that is a naming mismatch, not
a missing document.
Each converter's all alias expands to its own full set, and element-labels / labels now name
that notation's label shape, so a script written against the withdrawn core aliases keeps working
against a shape that exists.
Verified against the playground output of each converter: plantuml conforms (coverage 5/26),
structurizr conforms (7/11). ArchiMate reported 2 violations, both from archimate3:AssignmentShape
in the relationships set that was already the default — the newly added elements and
core-shapes contributed none. Those two were the playground fixture's own modelling error, corrected
in the entry below; ArchiMate now conforms too.
backstage-shapes was held out of the default set while two problems stood: a shape narrower than the
descriptor format, and a converter defect that left a relationship pointing at a node nothing
described. Both are fixed — see the entries at the top of this release — and it is now a default
alongside core-shapes.
CoreShapesValidateCommand is removed. It existed for converters that had no notation shapes
published yet; every converter now declares its own set, so nothing extended it.
Changed — --without-shape aliases are per notation, and BPMN disables nothing by default¶
bpmn2linkedarchi used to pass --without-shape labels by default, to suppress the core label rule
described above. With that rule withdrawn the default did nothing at all, and worse than nothing:
sh:deactivated on a shape absent from the graph succeeds silently, so the run reported disabling a
rule while changing nothing.
- The BPMN default is removed. Nothing needs suppressing now that the naming rule matches what BPMN actually requires.
- The shared aliases
element-labels/view-labels/labelsare gone fromcore, because the shapes they named no longer exist. Naming is notation-owned, so aliases are declared per converter. bpmn2linkedarchiacceptsrequired-names, pluslabelsandelement-labelsas synonyms for scripts that carry them. All three now disablebpmn/onto-shapes#RequiredNameShape— a shape that exists.- A converter with no aliases says so and asks for the full IRI, rather than listing aliases it does not have:
Validation error: Unknown --without-shape value 'labels'. This converter defines no aliases, so pass
the full IRI of the shape to disable — the SHACL report names it as sh:sourceShape.
Every converter defines labels and element-labels for its own notation's rule — see the entry
above, which wires those rules into each default shape set — so a script that passed either name
keeps working. What changed is which shape it disables: the notation's, not the withdrawn core one.
Changed — BPMN validates against core-shapes by default, and stops fetching bpmn-suite¶
core-shapes joins bpmn-shapes and bpmn-infra-shapes in the BPMN default shape set. The BPMN
shapes deliberately do not restate what core covers: bpmn/onto declares the connector classes to be
arch:QualifiedRelationship, and core-shapes#QualifiedRelationshipShape is what then checks each
one has exactly one arch:source and arch:target. Without it that contract went unchecked.
bpmn-suite is no longer loaded. It is now an owl:imports-only aggregator with no axioms of its
own, and SHACL processors do not dereference owl:imports, so fetching it added nothing — the
alignment it used to carry moved into bpmn/onto. The name still resolves for
--ontology bpmn-suite; only the default set changed. A BPMN validate run makes one fewer network
request.
Fixed — an index value YAML read as a mapping is refused instead of silently dropped¶
A links: or data: value written as a mapping without target: was skipped without a word. The way
an author reaches that is not by writing a mapping on purpose: YAML reads anything starting with {
as one, so an unquoted literal like {n/a} arrives as a map, and the statement disappears.
[ERROR] Malformed value for 'x:note' on 'order-service' in index.yaml (entry 'catalog'): a mapping
with no 'target:' — got [n/a]. Write 'x:note: { target: …, direction: … }', or quote the value if it
was meant as a literal: YAML reads anything starting with '{' as a mapping.
The { target: …, direction: … } form is unaffected. Colons and
quoting now documents this alongside the other
YAML reading worth knowing about in an index full of CURIEs — notably that an unquoted 12:30 is
sexagesimal in YAML 1.1 and reaches the graph as "750", while a colon inside a word (a CURIE, a
Backstage entity ref) needs no quoting at all.
Fixed — a PlantUML view's label was the punctuation of the list it was stored in¶
Every converted PlantUML view carried a wrong skos:prefLabel, in one of two ways. A titled diagram
published "[Order Flow]", brackets included. An untitled one published "NULL".
Both came from one line. A PlantUML title is a Display, which is a list of lines, and the
converter read it with toString() — so a title rendered as the list's [a, b], and the "no title"
sentinel rendered as the four characters NULL, which takeIf { isNotEmpty() } had no reason to
reject. The lines are now read as lines, and the sentinel is tested rather than its text, so a title
that genuinely reads NULL still works. A title wrapped over two lines becomes one label: the author
wrapped it for the picture's sake, and a newline inside a skos:prefLabel only moves the problem to
whoever renders it.
Element names were never affected — stripBrackets has always compensated for the same punctuation
there — and they are deliberately left alone, because an element's name is slugified into its id and
so into its published IRI. Unifying the two paths would change those IRIs, which is an identity
change rather than a cleanup.
Added — a view is labelled with a name someone wrote, not with its own id¶
Falling back to the id was the other half of the bug above: it is what the graph got whenever no title
existed, and order-flow-v2 is an address rather than something written to be read.
A title is now resolved from the index first, then the diagram, then the id — the rule identity
already follows, the index wins and the notation fills the gap. Using the index title also keeps
skos:prefLabel and dct:title from disagreeing, since the index title becomes the latter.
When neither declares one the run says so:
[WARN] flow.puml declares no title, so the view id 'flow' is published as its skos:prefLabel. Add
'title:' to the diagram index entry, or a 'title' line to the diagram — or pass --require-title to
make this an error.
--require-title makes it an error, for a publishing pipeline that should not ship a diagram
labelled with its own address. The sibling of --require-id: that one is about the address being
declared, this one about the name. Either route satisfies it.
Added — Backstage can be annotated from the diagram index¶
The index route was documented as reaching Backstage and never did: the module had no
--emit-extension-data, no --ns-global-id, and no call into the shared assertion emitter, so
elements: entries against a catalog were read and silently discarded.
diagrams:
- id: catalog
file: catalog.yaml
links:
arch:architectureState: arch:Baseline # about the model
elements:
order-service: # about one entity
links:
am:realizes: kg:CAP-OrderManagement
It is the same route every other converter has, with the same prefixes, direction tokens and
--emit-extension-data gate. The index is the only route, and that is deliberate rather than
pending: a catalog file is Backstage's schema, and inventing a key in it would be a claim on a format
these converters only read.
An entity resolves by its bare name, its entity ref in either casing
(Component:default/order-service, component:default/order-service) or its minted id
(default--order-service) — Backstage's docs write the ref lowercase while the catalog file spells
kind capitalised, so the two are one reference typed from two places. An ambiguous bare name is
refused and reported rather than attached to whichever entity was parsed last.
elements: now also works in the legacy flat schema, which it never did. That gap did not matter
while the schema's only users were diagram notations; a catalog is one entry per file, so the flat
schema is its natural form and the gap made the feature unusable there.
Two things were shared out rather than copied while doing this: AnnotationSubjects, whose real
content is the rule that an ambiguous name is refused rather than guessed at, and the --ns-global-id
terminator fix, which existed as an identical private helper in two converters and is now
UriUtil.ensureNamespaceTerminator.
The specifications now drive the Backstage jar as well, which is what would have caught this.
Added — the decision ontology and its SHACL shapes are resolvable by name¶
--ontology arch-decision and --shapes arch-decision-shapes now resolve, following their
publication at meta.linked.archi. Neither is in any converter's default set: a graph converted from
diagrams links to decisions but contains none, so shapes targeting ad:Decision would be downloaded
to validate nothing. Add them when validating a graph that holds decision records.
Added — architectureState: says whether a diagram is baseline, target or transitional¶
A repository holds the same architecture several times over: how it is today, how it is meant to end up, and the plateaus in between. Nothing said which, so the graph merged them — a target-state component and its baseline counterpart became one resource making contradictory claims, and a query for "what runs today" returned things nobody had built yet.
model:
id: order-domain
architectureState: baseline # the model is the baseline architecture
views:
- id: planned
file: planned.puml
architectureState: target # this diagram describes the target
Both emit arch:architectureState with arch:Baseline / arch:Target / arch:Transitional from the
core ontology. The skos:altLabels the ontology itself records are accepted — as-is, current,
to-be, future, transition, intermediate — as is any casing and -/_/space mix, so As Is
works.
A first-class field rather than a generic link, so a typo fails the run. The same statement is
writable through the annotation route added below, and there arch:architectureState arch:Targt
emits a reference to a term nothing defines and nothing complains, because a generic predicate has no
set of values to check against:
[ERROR] Conversion failed: Unknown architectureState 'Targt' for 'planned.puml' in
/repo/models/diagram-index.yaml. Accepted: baseline, target, transitional. … Note this is not the
publication status: 'draft' and 'publish' belong under 'status:'.
That last sentence is there because the two fields are the thing most likely to be confused. They
answer different questions — status: asks whether the diagram has been released, architectureState:
asks which reality it describes — and neither implies anything about the other. A target-state diagram
can be perfectly published: finished, reviewed, and describing a future that does not exist yet.
Three decisions worth knowing:
- Nothing is inherited. Unlike
status:, where a model-level value is only a default and the model is never the subject,arch:architectureStateis defined on a Model as readily as on a view and the core ontology's own example uses a Model. So each level states its own, and a consumer asking about a view that says nothing followsdct:isPartOfto its model. - The two routes must agree. If a
.pumland the index both declare a state and they differ, the run fails rather than one winning. Baseline and target are opposite claims about one diagram, so unlike a prefix declared twice there is no sense in which either declaration is more specific. - It is emitted into the semantic graph, not provenance, and is not gated by
--emit-extension-data. It says what the diagram is about rather than how the file came to be, and the generic route writes the same predicate into the semantic graph — one predicate landing in two named graphs depending on which route wrote it would silently halve any graph-scoped query.
Saying nothing asserts nothing: there is no default, because a repository that has never distinguished baseline from target is not thereby claiming everything is baseline.
Available on PlantUML (both routes), BPMN and Backstage (index only — neither format has anywhere in-file to declare it). See Architecture state.
Added — a statement can be about the diagram or the model, not only about an element¶
Extension data attached to an element by name: '!la-link OrderService … needs an OrderService in
the diagram. Some things are true of the diagram as a whole and of no element in it — which
architecture state it depicts, which viewpoint it conforms to, which decision option it articulates —
and had nowhere to go.
The subject can now be the view or the model:
| Subject | In the source file | In the diagram index |
|---|---|---|
| an element | '!la-link / '!la-data |
elements: on the view entry |
| the view | '!la-view-link / '!la-view-data |
links: / data: on the view entry |
| the model | — | links: / data: on the model: entry |
'!la-prefix arch: https://meta.linked.archi/core#
'!la-prefix x: https://example.org/vocab#
'!la-view-link arch:architectureState arch:Target
'!la-view-data x:reviewedBy Jane Doe
prefixes:
arch: https://meta.linked.archi/core#
model:
id: order-domain
links:
arch:architectureState: arch:Baseline # about the model
views:
- id: target-payments
file: target-payments.puml
links:
arch:architectureState: arch:Target # about this view
The view directives take no subject, because a .puml file is one view and there is nothing else
they could be about. Everything else is unchanged: the same prefixes, the same scalar-or-list-or-map
values, the same direction tokens, the same --emit-extension-data gate. A predicate that runs
towards the view is written Backward, as on the element route — arch:exposedInView goes from a
concept to the view, so '!la-view-link arch:exposedInView kg:ReadReplicas Backward emits
kg:ReadReplicas arch:exposedInView <this view>.
Two absences are deliberate. There is no in-file route for the model: a model spans several
files, so a model-wide claim written in one of them would have no evident owner, and two files
disagreeing would need a conflict rule nobody asked for. BPMN has no in-file route for the view:
bpmn:extensionElements nests inside the element it annotates, so there is nowhere in a .bpmn to
put a statement about the diagram. The index covers both.
Model-level statements stay at model level rather than being copied onto each view, for the reason
formerIds: is not merged across the two levels: they are different resources, and attaching a
model-wide claim to every view would say something about each diagram that the author said about the
whole model.
This is the primitive the architecture-state and decision vocabularies need, so both work through it
today with no vocabulary-specific code: arch:architectureState with arch:Baseline / arch:Target
/ arch:Transitional from core, and ad:Decision / ad:Option / ad:hasSelectedOption /
ad:hasAlternative from the arch-decision extension.
See Extension data and Statements about a view or a model.
Changed — every lifecycle state is converted by default, and --exclude-states withholds one¶
This changes what an existing pipeline produces. A build with status: draft or
status: in-review entries in its index used to leave them out. It now converts them. To keep the
old behaviour, add --exclude-states draft,in-review.
status: describes a diagram; it no longer decides whether the diagram exists. All five states are
converted and rendered, each emitting its own adms:status, and a consumer that wants only current
diagrams filters on that:
# Every state, each carrying its own adms:status
plantuml2linkedarchi convert models/*.puml --diagrams-index index.yaml -o out/model.trig
# A release build that must not publish unreleased work
plantuml2linkedarchi convert models/*.puml --diagrams-index index.yaml \
--exclude-states draft,in-review -o out/release.trig
The reason for the flip is which way each option fails. An inclusion list omits by default, so the mistake it invites is a state nobody named: diagrams silently missing, found later as dead links in documentation. An exclusion list publishes by default, so its mistake is something published early — visible at once, undone by a rerun. And withholding a draft hid it from exactly the SHACL validation and cross-model checks that would have caught its problems, so the first run to validate a diagram was the run that published it.
--exclude-states refuses to name all five states, because a run that converts nothing writes an
empty graph and exits zero, which reads as success.
--include-states and --include-drafts are accepted, ignored, and warned about. They only ever
added to the processed set, and that set is now everything, so anything they could have named is
already in. Reinterpreting --include-states as "only these" was rejected — it would quietly narrow
the output of every pipeline passing it:
[WARN] --include-states no longer changes anything: every state is converted by default, so draft
would have been included regardless. The flag is deprecated and ignored. To leave states out, use
--exclude-states.
Two rules that used to lean on selection now lean on a separate published property, so both still
hold: a published diagram still cannot go back to draft, and supersededBy: still cannot point at
an unreleased entry — no longer because a default build minted nothing for it, but because "read
this instead" should not send a reader to unfinished work.
The withheld-entry warning survives, narrowed to builds that pass --exclude-states, which are now
the only builds where a diagram can leave the output without anyone editing it:
[WARN] 'order-flow.puml' was published as 'publish' and is now 'in-review', which this run excludes,
so it is left out and its IRIs stop resolving. Drop 'in-review' from --exclude-states to keep
publishing it while it is out of use.
See ADR 0004, Amendment 2 and Lifecycle states.
Added — extension data for PlantUML, by two routes, and an index route for every notation¶
BPMN could carry statements the notation cannot express — which capability a process realizes, which
application serves a task — because bpmn:extensionElements exists to hold data a tool does not
understand. PlantUML has no such construct, so it had no way to say any of it.
It now has two ways, and they exist because the work is usually split between two people:
| Route | Written by | Lives in |
|---|---|---|
'!la-link / '!la-data comments |
whoever authors the diagram | the .puml file |
elements: entries |
whoever owns the pipeline | diagram-index.yaml |
'!la-prefix am: https://meta.linked.archi/archimate3/onto#
'!la-prefix kg: https://example.org/graph/
@startuml
class OrderService
'!la-link OrderService am:realizes kg:CAP-OrderManagement
'!la-data OrderService x:costCentre CC-4711
@enduml
prefixes:
am: https://meta.linked.archi/archimate3/onto#
kg: https://example.org/graph/
views:
- id: order-flow
file: orders.puml
elements:
OrderService:
links: { am:realizes: kg:CAP-OrderManagement }
data: { x:costCentre: CC-4711 }
Comments were chosen for the in-file route because every renderer ignores them and no tool rewrites
them, and because this repository already declares identity that way ('!la-model:) — one convention
rather than two. PlantUML's own $tags and [[url]] were considered and rejected: a tag is a single
token with no room for a predicate and an object, and [[url]] allows one per element and is already
claimed by render for graph hyperlinks.
The routes are additive. Both apply in one run and neither overwrites the other, so a platform team's index entries and an author's comments coexist. The only thing that has to resolve one way is a prefix declared in both places, and the source file wins — it is the more specific declaration, and the author editing it can see it.
The index route works for BPMN too, alongside its extensionElements route, and reaches
Backstage as well. It is the only route available to a notation with nowhere in-file to put anything,
which is why it lives in core rather than in a converter.
Everything the BPMN feature established is unchanged and now shared: the same CURIE precedence, the
same direction tokens, the same refusal to invent an IRI from a prefix nobody declared, the same
--emit-extension-data gate and --ns-global-id.
One asymmetry is worth knowing. BPMN's in-file route contains its statements — extensionElements
is nested inside the element — while every other route refers to one by name, and a reference can
be left behind by a rename. All of them report an annotation that matches no element rather than
dropping it:
[WARN] orders.puml:12: no element 'Nowhere' in this model, so <am:realizes> was not attached to
anything. Check the name against the diagram.
PlantUML makes referring harder than it should be, because an element there has up to three names:
class "Payment Gateway" as PayGw displays Payment Gateway, is written as PayGw, and is minted
as Payment_Gateway. Only the last reaches the graph and it is the least likely to be typed, so all
three resolve and an ambiguous name is refused rather than guessed at. Quote a name with spaces.
Full reference in Extension data.
Fixed — PlantUML render --base-iri now links shapes declared in bracket shorthand¶
Component and use case diagrams came out with almost nothing clickable. The renderer found an
element's declaration line only when it began with a keyword (component Foo, class "Foo" as F),
so every element written in PlantUML's bracket shorthand was skipped:
[WebUI] ' skipped — no link
[MobileApp] as Mobile ' skipped
() "PaymentApi" as PApi ' skipped
:Customer: ' skipped
(Browse Catalog) ' skipped
package "Frontend" { } ' linked
That is the shorthand most component and use case diagrams are written in, so the shapes a reader
would click were exactly the ones left unlinked. converter-plantuml's own
sample-component.puml produced 3 anchors for 9 elements — and all 3 were the packages drawn
behind the components. A use case diagram produced 1 anchor for 6 elements.
It was silent because the safety check asks PlantUML whether the diagram has any URL. One linked package answered yes, so a diagram whose every shape was unlinked was reported as a success.
Both are fixed: the shorthands [Component], (Use case), :Actor: and () Interface are now
recognised, with or without as Alias, and the elements that still got no link are listed at debug
level. Relationship lines are explicitly excluded, so [A] --> [B] cannot have a link appended to
its arrow. Nothing changes for diagrams that already worked — class and sequence diagrams use
keyword declarations and are unaffected — and the IRIs are unchanged, so no published link moves.
sample-component.puml now emits 9 anchors for 9 elements, matching the 9 element IRIs convert
mints for the same file exactly. Covered by PumlSvgLinkTest, which asserts on anchors in the
rendered SVG rather than on the annotated source, since PlantUML drops a link it cannot place.
Note that an SVG embedded via <img src> or a Markdown  image is never interactive in a
browser, whatever anchors it contains. Use <object>, <iframe> or inline <svg>.
Changed — --base-iri is required for BPMN and PlantUML, and no longer defaults to a urn:¶
Breaking for any command that omitted it. BPMN and PlantUML were the only two converters with a
default base IRI; ArchiMate, Backstage and Structurizr have always required one. The default was a
urn: placeholder, and an unconfigured run published every resource under it:
That IRI dereferences nowhere, and linkedarchi is not a registered URN namespace, so it is not a
well-formed URN either — it only looks like one. Worse, nothing said so: a pipeline that forgot the
flag published a whole model under a fake namespace and reported success.
The fix is to ask rather than to guess. Substituting an https default would have relocated every
IRI in an unconfigured run silently, which is the failure mode --require-id and
identity-lock.yaml exist to prevent elsewhere in this tool; a namespace is a publishing decision
and belongs to whoever is publishing. Omitting it now fails immediately:
Add --base-iri https://example.org/la/ to any BPMN or PlantUML convert that relied on the
default. Nothing in this repository did — every playground script, documentation example and
specification already passed one — so the change is invisible here and visible only to callers that
were producing urn: IRIs by accident. render --base-iri is unaffected and stays optional, since
omitting it there means "no hyperlinks" rather than "no namespace".
Added — BPMN extension data reaches the graph, with prefixes resolved from the document¶
A BPMN file could already express links to ArchiMate, LeanIX and a knowledge graph as foreign
children of bpmn:extensionElements, and they survived every tool round-trip, but none of them
reached the RDF: the payload was dropped and only an empty bpmn:ExtensionElements node came out.
BPMN was the only authored notation whose external references were unqueryable.
--emit-extension-data maps them. The extension element's own qualified name is the predicate,
and the triple is flattened into the one a query would otherwise have to reconstruct:
<bpmn:process id="Process_GeoMapping">
<bpmn:extensionElements>
<am:realizes rdf:resource="kg:CAP-MasterDataManagement"/>
</bpmn:extensionElements>
<…/bpmn/geo-mapping/element/Process_GeoMapping>
am:realizes <https://example.org/graph/CAP-MasterDataManagement> .
Nothing is configured to make that work, because XML already did it. Prefixes in element names
are resolved by any XML parser — it is the one place they are — so <am:realizes> arrives carrying
the namespace the document bound to am. There is no carrier vocabulary to adopt and no namespace
option for predicates.
What XML will not resolve is a prefix in an attribute value, which is why kg:CAP-MasterDataManagement
needs the converter's help and is the only place it does prefix resolution of its own. First match
wins:
| Form | Result |
|---|---|
| CURIE whose prefix the document declares | expanded |
Absolute http/https IRI |
passed through |
No prefix, with --ns-global-id set |
resolved against that base |
| Anything else | warning, kept as a literal |
A declared prefix wins over reading the value as an absolute IRI, because lx:APP-x and
doi:10.1000/x have the same shape and the author's own declaration is the better evidence. A
prefixed value with no declaration is reported rather than emitted as an IRI that resolves nowhere —
silently emitting one was the defect this replaces, so it is not replaced with a quieter version of
itself. A few schemes without :// (urn:, mailto:, tag:, doi:, data:) are accepted as
absolute, since they are unambiguously IRIs, but the documentation and fixtures use dereferenceable
https IRIs throughout.
Whether an element is a link or a literal follows from its object, not its name: rdf:resource (the
RDF/XML convention) says "reference" outright, text that names a resource anyway is read as one, and
any other text is a literal. Element children make a node, so two groups on one element stay
distinguishable.
Write the name from a published ontology. Terms already exist for most of what a BPMN file wants to say about the world outside it:
| To say | Use |
|---|---|
| this process realizes a capability | am:realizes |
| this application serves the task | am:serves with direction="Backward" |
| this task adds detail to a coarser concept | arch:refines |
| this task is part of a larger concept | arch:partOf |
| where its master data lives | arch:masterDataSource |
| it is the same thing as that | owl:sameAs |
| it corresponds to / relates to that | skos:exactMatch, skos:related |
Nothing constrains a consumer to those — point at your own ontology and every rule here is unchanged — but the examples, fixtures and specifications in this repository model the published vocabulary rather than inventing one to demonstrate with.
Identity therefore needs no separate mechanism: an element named owl:sameAs is an owl:sameAs
triple. That matters because most cross-model references are typed relations (am:realizes,
arch:refines) that identity alone cannot express, so --ns-global-id supplies a base for bare
rdf:resource values rather than designating an identity field as it does in the ArchiMate
converter.
direction is honoured — Forward (the default, and what None is read as), Backward, Both.
An unrecognised value is reported and read as Forward rather than dropping a well-formed link.
Backward is what lets an author reuse a term that runs the other way, am:serves being the
obvious one, instead of inventing an inverse.
Data with no published term is written the same way and needs no namespace option either, because
the predicate namespace comes from the element: with xmlns:x="https://example.org/vocab#",
<x:costCentre>CC-4711</x:costCentre> becomes vocab:costCentre "CC-4711". Two extension
vocabularies therefore cannot collide. An unprefixed attribute — <x:criticality level="high"> — has
no namespace in XML and is read against the element it qualifies; direction and rdf:resource are
never emitted as data. A namespace that does not end in # or / would run into the local name, so
it is reported rather than emitted. dc: and dcterms: children are left alone, since those are
already read into the provenance graph and mapping them twice would state one fact under two
vocabularies.
Off by default: a file authored in a modelling tool carries that tool's extension elements, and turning Camunda form data into triples unasked would change the graph of every existing model.
bpmn:relationship remains out of scope and is now documented as such — it is valid BPMN and
looks like the intended mechanism, so leaving it an undocumented no-op was its own trap. Full
reference in BPMN → Extension data.
An earlier draft of this feature wrapped the predicate in a carrier element and put it in a type
attribute (<x:link type="am:realizes" target="kg:CAP-1"/>). That shape is refused with an
instruction, not silently mapped, because taking the carrier's own name as the predicate would
emit literals called type and target — junk that looks like data. It was dropped rather than
kept alongside the direct form because it is strictly worse: it needed a carrier vocabulary nobody
had published, and it moved the predicate into an attribute where XML will not resolve its prefix,
so the converter had to re-implement resolution the XML parser already does for element names. No
--ns-vocab option exists for the same reason.
Changed — bpmn:ExtensionElements is no longer counted as an architecture element¶
It is a container the serialization needs, holding nothing of its own, so typing it arch:Element
and arch:ModelConcept overstated a model's elements and put a wrapper in the way of any query for
them. It keeps its BPMN type and its folder membership; what it carries is now mapped onto the
element it hangs off, which is where a consumer would look for it. bpmn:definitions and
bpmn:Import were already excluded on the same grounds.
Element counts in convert output drop by one per annotated element.
Fixed — a failed BPMN render says what is wrong with the file¶
The Camunda parser reports every parse failure as SAXException while parsing input stream and
puts the detail in the cause, but only the top-level message reached the [WARN] Skipped … line.
The commonest cause was invisible: bpmn:source and bpmn:target are xsd:QName, so an
undeclared prefix — tns: is the usual one — fails render while convert succeeds, because only
Camunda's parser resolves the QName. The message named neither the prefix nor the line.
[WARN] Skipped flow.bpmn: SAXException while parsing input stream — Error: URI=null Line=16:
UndeclaredPrefix: Cannot resolve 'tns:Process_1' as a QName: the prefix 'tns' is not declared.
Added — status: accepts the whole lifecycle, and reaches the graph as adms:status¶
A diagram index could say publish or draft. It can now say in-review, deprecated or
archived too, and the state is published rather than discarded once it has been used to filter.
The states split on whether an IRI for the diagram has ever been public, which is what decides their treatment:
status: |
Converted and rendered | adms:status |
|---|---|---|
draft |
no | adms/status/UnderDevelopment |
in-review |
no | adms/status/UnderDevelopment |
publish (default) |
yes | adms/status/Completed |
deprecated |
yes | adms/status/Deprecated |
archived |
yes | adms/status/Withdrawn |
deprecated and archived are deliberately still published. Their IRIs are already linked from
documentation and from other models, so removing them would produce dead links rather than
communicate a deprecation — the same harm formerIds: exists to avoid. They stay, carrying a
triple that says what they are:
<https://example.org/la/plantuml/order-domain/graph/provenance> {
<…/view/current> adms:status <http://purl.org/adms/status/Completed> .
<…/view/leaving> adms:status <http://purl.org/adms/status/Deprecated> .
}
adms:status and its concepts come from ADMS, so a
consumer that already reads DCAT-AP understands the lifecycle without knowing anything about
Linked.Archi. No term is minted in the arch: namespace. status was previously the only index
field emitProvenance did not emit, so nothing in a published graph distinguished a current
diagram from one kept only for the record.
--include-states names states to add to the published set, and adds rather than replaces, so a
preview build gets the drafts and the live diagrams:
plantuml2linkedarchi convert models/*.puml --diagrams-index index.yaml \
--include-states draft,in-review -o out/preview.trig
--include-drafts still works and means --include-states draft. It is deprecated: it cannot
name the other states.
Full rules in ADR 0004.
Added — supersededBy: names what to read instead of a retired diagram¶
A status says a diagram is on its way out. This says where to go, which is the more useful half of a deprecation:
<…/view/order-flow-v1> dct:isReplacedBy <…/view/order-flow-v2> .
<…/view/order-flow-v2> dct:replaces <…/view/order-flow-v1> .
No owl:sameAs, which is what separates this from formerIds:. A former id is the same resource
under a previously published name, so asserting identity is right. A successor is a different
diagram, and claiming the two IRIs denote one thing would merge them — along with their elements,
under any reasoner. Use formerIds: to rename a diagram and supersededBy: to replace one.
Read at the level the entry sits on, like formerIds:. Refused if the successor is not in the index,
is in a state the build withholds, is the entry itself, or forms a cycle — all so that following
dct:isReplacedBy arrives somewhere real. Also refused on a publish entry, since "current, and
also replaced" usually means the two entries were swapped.
Added — the states form a lifecycle, and draft is no longer a synonym for in-review¶
Both are withheld, so the difference is whose turn it is: draft is work in progress with nobody
waiting on it, in-review is submitted and waiting for validation. That makes in-review the gate
into publication and the only state a published diagram can return to without being retired.
| From | May become |
|---|---|
draft |
in-review |
in-review |
draft, publish |
publish |
in-review, deprecated, archived |
deprecated |
publish, archived |
archived |
publish, deprecated |
Two moves are refused: anything published going back to draft, because its IRIs are public and
"unfinished" is not something a public resource can be; and draft straight to publish, which
skips the validation in-review exists to require. Nothing is a dead end, so a retired diagram can
always be reinstated.
identity-lock.yaml now records each entry's state alongside its ids, since a transition can only be
checked against what came before — the same argument ADR 0003 makes for rename detection. A state the
record has never seen is not a move, so adopting this needs no migration: the first run records, the
next enforces. --allow-state-change authorises a refused move.
[ERROR] Conversion failed: Illegal state change for 'order-flow.puml': 'publish' cannot become
'draft'. 'publish' has been published, and a published diagram cannot become work in progress
again — its IRIs are already public. Send it to 'in-review' instead.
Sending a live diagram back for validation is legal and takes it out of the build, so it warns once, on the run where the link breaks:
[WARN] 'order-flow.puml' was published as 'publish' and is now 'in-review', so it is left out of
this run and its IRIs stop resolving. Add --include-states in-review to keep publishing it while
it is out of use.
Withheld entries are reconciled too, which matters more than it sounds: the transitions that take a
diagram out of the output are exactly the ones whose entry is then never converted, so checking
only converted entries would have let publish → draft silently drop a published diagram.
Added — --no-render-archived¶
Stops render drawing entries with status: archived, for an output directory that should hold only
what is still in use. convert is unaffected, so the diagram keeps its IRIs and nothing in the graph
breaks; only the image stops being written, and any URL already published for it stops resolving.
Opt-in for that reason.
Fixed — an unrecognised status: fails the run instead of publishing¶
Breaking for an index carrying a status outside the accepted set. status: was a free string
in which only the exact token draft had any effect, so everything else fell through to
"publish": status: Draft published a draft, status: archived published an archived diagram as
current, and status: publsih published — making a typo in draft indistinguishable from
publish. The failure was silent and in the wrong direction.
Unrecognised values are now refused, naming the value, the file and the accepted set:
[ERROR] Conversion failed: Unknown status 'publsih' for 'order-flow.puml' in
/repo/models/diagram-index.yaml. Accepted: draft, in-review, publish, deprecated, archived.
Case and the choice of -, _ or a space no longer change the meaning, so In Review and
in-review are one state. published is accepted as a synonym for publish and obsolete for
deprecated; both published silently before, so recognising them is safer than refusing them.
Validation covers the whole index before entries are filtered out, so a typo on a draft fails on the run that introduced it rather than on the later run that first tried to publish it.
An entry that declares no status: is still published, and now emits no adms:status at all:
silence is not an author calling a diagram finished.
Added — render writes the SVG under a superseded view id, and an alias lasts as long as it is declared¶
A view id names the rendered file as well as the view IRI, so renaming a view relocated a published
image URL with nothing left at the old one. render now also writes the picture under each id the
entry declares under formerIds::
Closing that required moving the aliases from the lock record to the declaration, which fixes a
defect of its own. Alias triples were derived from identity-lock.yaml: the run that renamed an id
found the previous one in the record and emitted dct:replaces / owl:sameAs /
dct:isReplacedBy — and rewrote the record, so every later run published the model without the
alias and the previous IRI stopped resolving on the next re-publish. render could not have used
the record at all, since convert normally updates it first. Both commands now read formerIds:,
so the alias triples and the superseded SVG are emitted on every run for as long as the declaration
stands, and retiring an alias means removing the declaration.
--allow-identity-change consequently emits no aliases. It authorises a change without describing
one, and is documented for identities that were never published outside the repository.
formerIds: is now read at the level it sits on, where the two levels were previously merged into
one list:
model:
id: order-domain
formerIds: [orders] # aliases the model IRI
views:
- id: fulfilment
formerIds: [order-flow] # aliases the view IRI, and names a superseded SVG
file: flow.puml
Merging them meant a former view id authorised a renamed model, and a former model id would have
been minted as a former view IRI. Each level is now checked and aliased against its own
declarations. Former ids are also validated as slugs, like current ones: they are published as IRIs
and, for a view, written as filenames, so an unchecked ../x wrote outside -o.
Changed — a BPMN model has the same three folders as every other notation¶
BPMN grouped elements into six type-category folders — Processes, Tasks, Events,
Gateways, Data, Other — while PlantUML, Backstage and Structurizr emit Elements,
Relationships and Views under the model folder. BPMN now emits those same three.
The categories restated in folder membership what rdf:type already carries, and the mapping
from BPMN class to category lived in a when block that had to be extended for every new
subclass. A consumer selects tasks with ?e a bpmn:UserTask, which needs no folder:
# before
<…/bpmn/order-domain/folder/Tasks> a arch:Folder ; schema:name "Tasks" .
<…/bpmn/order-domain/element/Task_1> dct:isPartOf <…/bpmn/order-domain/folder/Tasks> .
# after
<…/bpmn/order-domain/folder/Elements> a arch:Folder ; schema:name "Elements" .
<…/bpmn/order-domain/element/Task_1> dct:isPartOf <…/bpmn/order-domain/folder/Elements> .
The six category folder IRIs and the dct:isPartOf and schema:itemListElement triples that
referenced them are no longer published; a consumer that navigated them reads rdf:type
instead. FolderStructure loses classifyElement and its MetaModel dependency. A scenario
per notation family pins the folder set.
Fixed — a BPMN model folder lists its children once, not once per view¶
The child-folder listing describes the model, so it belongs to the model. It was emitted per
input file, and its schema:ListItem entries are blank nodes, so nothing collapsed them: two
views of one model left the model folder holding two children at schema:position 1, three
views three. The listing is now emitted with the first view of a model. Member entries were
already continued across views by the position counters.
Added — diagram-index entries are validated: id shape, and id/file: uniqueness¶
DiagramsIndex.parse() took id verbatim, and two places then used it unescaped: as an
IRI path segment (IriMinting, {base}{notation}/{id}/…) and as the render output
filename. Anything that is not slug-shaped broke one or both of those, silently:
id |
What used to happen |
|---|---|
My Model |
a space in the published SVG name, and an IRI with a literal space — not valid RDF |
order#v2 |
everything after # became a fragment, so the IRI stopped identifying the model |
a/b, ../x |
extra IRI path segments, and an SVG written outside -o |
Order vs order |
two distinct IRIs that collide on a case-insensitive filesystem |
the same id twice |
one model's SVG overwrote the other's, and both shared one set of IRIs |
Ids must now match ^[a-z0-9][a-z0-9._-]*$. A violation fails the command with the
offending entries named, rather than being normalised: a silently rewritten id would put a
different identity in the graph than in the filename, which is exactly the mismatch the
shared id exists to prevent.
id and file: must also be one-to-one, and both directions now fail loudly:
- A repeated
idused to mean two models sharing one set of IRIs, with one SVG overwriting the other. - A repeated
file:used to mean the opposite kind of loss. The index is keyed by file, so only the last entry for a file took effect and the earlier id rendered nothing and minted nothing — no warning, no output. That looked like a way to publish one diagram under two names (an old id kept alive beside a new one during a migration) and never was. Equivalent spellings of one path —nested/x.bpmnand./nested/./x.bpmn— count as the same file. To publish an alias, copy or link the rendered SVG instead.
Validation runs in the shared parser, so convert and render agree, and it covers
status: draft entries too — a broken or colliding entry should surface while it is still
a draft, not on the day it is published.
This can fail an index that previously converted. Every index in the repo already complies; a non-compliant one was already producing one of the defects above.
Validation applies to the index, which is one of three places a model ID can come from.
The other two remain unvalidated: --model-id, and the input filename when neither an
index nor --model-id supplies anything. docs/architecture/diagram-index.md now states
the resolution chain per command instead of leaving it to be inferred, along with which
converters read an index at all — ArchiMate and Structurizr do not, because for them one
file is one whole model rather than a diagram to select.
Fixed — --model-id with several inputs is refused instead of misapplied¶
A model ID identifies one model, so one explicit value cannot serve a batch. Each converter handled that on its own and all of them differed:
- BPMN and PlantUML silently ignored
--model-idas soon as a second input appeared. An explicitly chosen identity was dropped without a word, and the filenames decided instead. - Backstage and Structurizr applied it to every input, so several sources collapsed
into one model IRI namespace and overwrote each other — the same defect the diagram index
rejects as a duplicate
id. Two catalogues converted with one--model-idproduced 263 triples under a singlebackstage/{id}/graph where the indexed run produces 291 under two.
ModelIdOption in core now decides for all of them: more than one active input is an
error naming the count and the remedy, nothing is written, exit 2. Passing --model-id
alongside an index is still allowed but warns that the index supplies the id, rather than
having the option quietly do nothing. ArchiMate is unaffected — -i takes one file, so the
option was never ambiguous there.
Changed — bad input reports a message, not a stack trace¶
IllegalArgumentException — what require throws, and what every input check in this
release raises — reached the user as a Picocli stack trace with exit 1 from the converters
that did not catch it themselves, and as a message plus a stack trace from the ones that
did. A rejected --model-id or an invalid index is something to read, not to debug.
CliRunner in core is now the single main entry point for all five CLIs. It prints
[ERROR] <message> and exits 2 for that class of failure, and keeps the full trace for
anything unexpected. The commands that catch their own failures use CliRunner.isUserError
to make the same distinction when logging.
Fixed — --include-di disabled diagram interchange instead of enabling it¶
Picocli treats the positive form of a negatable option as a toggle unless a fallback value
is declared. Every option here that defaults to on was declared without one, so naming it did the
opposite of what it says:
| Invocation | Before | Now |
|---|---|---|
| (nothing) | DI emitted | DI emitted |
--include-di |
DI dropped, no views graph, silently | DI emitted |
--no-include-di |
DI emitted | DI dropped |
playground/run-bpmn-full-trig.sh passes --include-di explicitly, so the playground had been
publishing BPMN output with no graph/views at all — 345 triples where there should have been
568. --dual-typing (BPMN), --include-views and the --emit-skos-* pair (Structurizr,
Backstage, and the shared BaseConvertCommand) had the same defect. PlantUML was already correct
and showed the fix: fallbackValue = "true".
A @flags feature now pins all three invocations, so the next negatable option cannot quietly
invert itself.
Changed — the playground demonstrates models holding views¶
The sample indexes taught the flat schema, and the two PlantUML diagrams shared no elements, so nothing showed why grouping matters.
plantuml/diagram-index.yamldeclares modelshopwith three views, and a newstock-replenishment.pumldeliberately reusesWarehouse,StockItemandSupplierfrominventory-domain.puml. Each is one element with anarchvis:ArchNodein two views.run-plantuml-views.shconverts all three in one run and prints the resulting structure.bpmn/diagram-index.yamldeclares modelorder-domain, and a newreturns-handling.bpmnis its second view.convertputs both in one namespace;renderwrites one SVG per view.- Backstage indexes stay flat, which is correct there: a catalog file is not a diagram and has no views.
Added — executable specifications for the diagram-index contract¶
A new :spec module holds Gherkin features run by Cucumber on the JUnit Platform. They describe
what the tool promises rather than how it is built: what an index means, which IRIs a run mints,
how a bad index fails, and which exit code comes out. If a scenario and
docs/architecture/diagram-index.md disagree, the scenario is the one that has been run.
models-hold-views.feature— grouping, an element shared across views,arch:Model/arch:Viewtyping, metadata levels, SVG naming, and the legacy flat schema.index-validation.feature— id shape per level, duplicate view ids and files, element collisions,--model-idrules, the no-match warning. The id matrix is oneScenario Outlinewith twoExamplestables.
Scenarios drive the fat JARs as subprocesses: the commands end in exitProcess(), which
would take the test JVM with them, and exit codes and the [ERROR] / [WARN] lines only exist
in a real process. Assertions read the emitted RDF with RDF4J, so "exactly one element" means one
subject in the graph rather than a string match. The module ships no production code and is not
part of the distribution.
Unit tests cover what is cheaper as a table than as a scenario: core gained DiagramsIndexTest
for schema detection, metadata inheritance, draft filtering, per-level id validation and
input-to-entry lookup. Writing it removed a validation branch that could never fire — an index is
read as grouped or as flat, never both, so no entry set can mix the two. Mixing the schemas in one
file is still refused, with a message that now says why.
docs/development/testing.md records where each level belongs and how to add a scenario.
Fixed — BPMN folder IRIs are inside the model namespace, like every other converter's¶
FolderStructure minted {base}folder/bpmn/{modelId}/{name}, putting a model's folders outside the
model that contains them, while PlantUML, Backstage, Structurizr and ArchiMate all use
{base}{notation}/{modelId}/folder/{name}. A query for a model's folders therefore had to be
written twice, and BPMN's version could not be derived from the model IRI.
BPMN now mints folders through the shared IriMinting.folderIri:
# before
<https://example.org/la/folder/bpmn/order-domain/Elements> a arch:Folder .
# after
<https://example.org/la/bpmn/order-domain/folder/Elements> a arch:Folder .
Its own percent-encoding went with it: folder segments are built from the notation slug, the model id and fixed folder names, and a declared id is validated rather than escaped.
This changes published folder IRIs for BPMN models. Folders are structural — nothing outside the model references them — but a consumer that queried the previous shape needs updating. Two scenarios now pin the shape, one per notation family.
Added — a model identity must be declared, and a published one cannot change silently¶
Also in this release: a declared id is validated wherever it is declared. The diagram index
validated its own entries, but --model-id and an in-file declaration reached IriMinting
unchecked, so --model-id "My Model" --require-id succeeded and minted …/plantuml/My%20Model.
Both are now held to ^[a-z0-9][a-z0-9._-]*$. The filename fallback stays permissive, since it is
a derivation rather than a declaration, but it no longer passes in silence:
[WARN] 'Order-Flow.puml' has no declared model id, and its filename is not a valid id
(^[a-z0-9][a-z0-9._-]*$), so 'Order-Flow' becomes an IRI path segment as it stands.
Declare an id, or pass --require-id to make this an error.
Implements ADR 0001 (ownership of model identity) and ADR 0003 (changing a published identity),
both under docs/architecture/decisions/.
--require-id removes the filename fallback. A model ID is the IRI namespace of everything in
the model; deriving it from a filename produces an identity that is unvalidated and changes when
the file is renamed. Available on every convert command and on render:
[ERROR] Conversion failed: --require-id is set and no model id was declared for 'Order Flow.puml'.
Without one the filename would become the identity, which is neither validated nor stable.
Declare one in the diagram index, or pass --model-id.
PlantUML sources can declare their own identity in header comments, read by convert and
render alike:
Only the lines before @startuml are read, so an identity cannot be declared among the shapes it
identifies. A declaration in the source and one in the index must agree; a disagreement fails the
run naming both values, rather than one side silently winning. BPMN identity stays index-declared
until a modeller round-trip is shown to preserve a custom <extensionElements> entry.
Changing a published identity is now an error. An indexed run records what it published in
identity-lock.yaml, written beside the index and committed like a dependency lock. A later run
whose IDs differ fails:
[ERROR] Conversion failed: Identity change for 'flow.puml': view id was published as
'order-flow', now 'fulfilment'. Declare the previous id under formerIds:, or pass
--allow-identity-change.
Declaring the change authorises it and keeps the previous IRI resolvable:
<…/plantuml/order-domain/view/fulfilment>
dct:replaces <…/plantuml/order-domain/view/order-flow> .
<…/plantuml/order-domain/view/order-flow>
owl:sameAs <…/plantuml/order-domain/view/fulfilment> ;
dct:isReplacedBy <…/plantuml/order-domain/view/fulfilment> .
--allow-identity-change accepts an undeclared change and updates the record. A run that fails for
any reason leaves the record untouched, and a run converting a subset of the index keeps the
entries for the files it did not process.
Cross-repository uniqueness (ADR 0001, decision 2) is not implemented here: a converter cannot see the other source repositories, so the check belongs to the aggregation repository.
Fixed — a diagram is a view of a model, and the index can finally say so¶
A diagram is an arch:View inside an arch:Model. The index could not express that: it had one
flat list, so for BPMN and PlantUML every file became its own model holding a single view.
Two diagrams showing the same component minted two unrelated elements, each in its own set of
named graphs, with nothing asserting they were the same thing.
The index now names the model and lists its views:
model:
id: order-domain
title: "Order Domain"
views:
- id: order-flow
file: order-flow.puml
status: publish
- id: fulfilment-flow
file: fulfilment-flow.puml
models: takes a list of such models. Both diagrams convert into one namespace, so the shared
element is one resource with a node in each view:
<https://example.org/la/plantuml/order-domain> a arch:Model .
<https://example.org/la/plantuml/order-domain/element/OrderService> a arch:Element .
<https://example.org/la/plantuml/order-domain/view/order-flow> a arch:View , arch:Diagram .
<https://example.org/la/plantuml/order-domain/view/fulfilment-flow> a arch:View , arch:Diagram .
Three named graphs where there were six, and one OrderService where there were two.
What came with it:
arch:Modelandarch:Vieware emitted. Previously only ArchiMate typed the model, and BPMN, PlantUML and Structurizr typed viewsarch:Diagramwhile ArchiMate usedarch:View, so a query for one class missed the other. BPMN and PlantUML now emit both classes on a view plusdct:isPartOfto the model.- Provenance splits by level.
dct:source,title,created,modifiedanddescriptionattach to the view;author,versionand the model title to the model. rendernames SVGs after the view. Element hyperlinks still resolve to{model}/element/{id}, since elements belong to the model.- Element collisions fail the run. Sharing an element namespace is the point of grouping, but
BPMN ids are only unique per file, so two views claiming one id with different names or types
is a defect.
ElementCollisionsreports the IRI and both source files, and writes nothing. Two views showing the same element consistently merge silently, as intended. - View ids are validated like model ids — slug shape, unique within their model.
- An indexed BPMN view id overrides the process-as-view remap, so a named view keeps its own identity instead of being folded onto the process it depicts.
The legacy flat diagrams: list is untouched, including its provenance subjects and SVG paths:
models: counts as the grouped form only when its entries carry views:, so no existing index
changes meaning or needs a version marker. Adopting the grouped form does move IRIs, since the
model id changes — pair it with the alias mechanism in
ADR 0001 (docs/architecture/decisions/0001-model-identity-ownership.md).
Added — architecture/identifiers.md and a decision record for model identity¶
"ID" meant three distinct things across the docs — the model ID that names an IRI namespace,
the view ID of a diagram surface inside it, and the element ID of a single node — with no page
stating which is which or where each comes from. docs/architecture/identifiers.md now sets
them out per converter, including the derivations that surprise people: PlantUML element IDs
are slugified display names, so renaming a shape moves its IRI; BPMN remaps a BPMNDiagram
onto the IRI of the process it depicts rather than minting a view of its own; Backstage
relationship IDs are an emission counter; ArchiMate is the only converter that percent-encodes
local IDs, while Structurizr view keys reach the IRI with their spaces intact.
Two decision records accompany it, both Proposed — nothing in this release implements either:
- 0001, where model identity lives — central index, source files or a hybrid, evaluated against authoring ownership, conflict resolution and versioning. Proposes the hybrid with the location change sequenced last, behind strict-ID mode and an alias mechanism.
- 0002, models hold views — the index calls an entry's
ida model ID, but for BPMN and PlantUML one file becomes a model and its only view, so the same element drawn in two diagrams mints two IRIs instead of one element with two views. Proposes an opt-inmodels:/views:grouping, with the element-merge semantics stated and a collision check, and re-scopes 0001's cross-repo gap to the archi-graph repository where models are merged.
Fixed — one diagram index, read the same way by every command¶
Both convert commands looked up index entries differently, and neither handled the two
file: spellings an index can legitimately use:
index file: |
BPMN convert (before) |
PlantUML convert (before) |
render (before) |
|---|---|---|---|
processes/order.bpmn |
no match | match | match |
order.bpmn, file nested |
match | no match | match |
BPMN matched bare file names only, PlantUML relative paths only, render both. So the same
index selected different diagrams depending on which command read it, and a non-match was
indistinguishable from having nothing to do: BPMN reported Converted OK — 1 file(s) while
writing an empty graph, because the summary counted the inputs it was given rather than the
ones it processed.
Selection now lives in one DiagramSelection in core, built through
DiagramsIndex.select(index, root, includeDrafts) and used by all four index-aware
commands — BPMN, PlantUML and Backstage convert, plus BaseRenderCommand:
- Relative path first, then the bare file name, so both spellings work everywhere. The
bare-name fallback applies only when that name identifies one entry, so
processes/x.bpmnandevents/x.bpmncannot be confused for each other. - An index that matched no input logs
[WARN] No input file matched the diagram index. Check --diagrams-root and the file: paths.Previously only Backstage said anything. - BPMN's summary counts processed files, matching the other converters.
parseAsFileMap and three near-identical private resolveRelativePath helpers are gone.
ModelsIndex.kt is now DiagramsIndex.kt, matching --diagrams-index,
diagram-index.yaml and the DiagramsIndex object — the file name was the last place
still calling it a model index. The legacy models: root key is still accepted.
Fixed — PlantUML render --base-iri now actually produces hyperlinks¶
PumlSvgRenderer.render() accepted an elementBaseIri and never read it. The command
computed the right IRI and passed it in, where it was dropped: the rendered SVG had no
links at all, while both the renderer's own KDoc and docs/architecture/svg-rendering.md
claimed shapes were wrapped in <a href>.
Links are now produced by annotating a copy of the source with PlantUML's own [[iri]]
syntax on each element declaration, letting PlantUML emit the anchors:
- The IRI matches what
convertmints ({base}plantuml/{modelId}/element/{id}), so a click lands on a node that exists in the graph. - The
.pumlfile on disk is never modified, and links the author already wrote are left alone. - Elements declared only inside a relationship (
[A] --> [B]) have no declaration line to annotate and stay unlinked. - If the annotated source stops parsing, or PlantUML reports no links in the result, the diagram renders without links and logs a warning instead of emitting something broken.
Added — PlantUML render --layout, and Graphviz is no longer required¶
Class, component, use case, state and deployment diagrams need Graphviz's dot algorithm.
Previously a machine without the dot binary got a valid SVG whose entire content was
"Cannot find Graphviz", reported as Rendered: …. That silent failure is gone, and the
dependency itself is now optional: PlantUML bundles Smetana, a pure-Java port of dot,
which the renderer enables by inserting !pragma layout smetana into a copy of the
source (a layout pragma the author wrote always wins).
--layout |
Behaviour |
|---|---|
AUTO (default) |
dot when installed, otherwise Smetana with a warning |
DOT |
Require dot; skip the file with an explanatory message when missing |
SMETANA |
Always the bundled engine — no external dependency, same layout everywhere |
AUTO keeps existing setups on the reference layout while making a bare container work
out of the box. SMETANA is the one to pin in CI, since under AUTO the same source
renders differently depending on whether the agent has Graphviz installed. Sequence
diagrams never needed either engine and are unaffected.
Two smaller fixes came with it: a diagram that produces no image at all now fails instead of reporting success without writing a file, and the SVG is written straight to the requested path instead of being renamed after PlantUML picks its own name.
Changed — both render commands now share one implementation¶
BaseRenderCommand held only the option declarations, and PlantUML did not even extend
it, so the two commands had drifted apart in ways nobody would choose deliberately:
| Behaviour | BPMN (before) | PlantUML (before) |
|---|---|---|
| Index lookup key | relative path, falling back to file name | relative path only |
| Output sub-directory | from the index file: path |
from the path relative to --diagrams-root |
| Setup failure (bad index/config) | exit 2 with a message |
uncaught exception + stack trace |
| Failed file | left an Error: … SVG behind |
left nothing |
| Missing input file | aborted the whole batch | aborted the whole batch |
The whole render loop now lives in BaseRenderCommand — index selection, model IDs,
element IRI minting, output paths, logging and exit codes — and each converter supplies
only a notation slug and a renderOne(). Notation-specific options stay on the
subclasses (--shape-config for BPMN, --layout for PlantUML). The CLI surface is
unchanged.
Behaviour is now identical on both sides, and two things changed for the better in the process:
- A failed file leaves no output. BPMN used to write a red "Error: …" SVG and then report the file as skipped, so a consumer scanning the output directory saw a file that looked rendered.
- A missing input file is a skipped file (exit
1), not an aborted batch (exit2). One bad path inrender *.bpmnno longer discards the status of everything that rendered.
Exit codes, both converters: 0 all rendered, 1 at least one file skipped, 2 setup
failed before rendering started.
Fixed — svg-rendering.md claims the code did not support¶
- The design rationale said
renderreuses the parser thatconvertuses. It does not: BPMN parses twice with two different libraries (StAX/CMOF for RDF, Camunda for BPMNDI geometry), and PlantUML uses two entry points into plantuml-mit. The actual reasons for co-locatingrender— shared index, shared IRI scheme, one image — are stated instead. - The
idfield was described as naming both the.trigand the.svgoutput. It names the SVG;convertwrites wherever-opoints, and thereidfeeds IRI minting. - Added what was missing:
--shape-configfor BPMN, the Graphviz requirement for PlantUML, and a section on how each renderer produces hyperlinks.
1.2.0 — 2026-07-27¶
First tagged release of the consolidated converters. Everything published before this
point was an overwriting 0.1.0-SNAPSHOT build of the default branch, so this is the
first version that can be pinned. The number starts at 1.2.0 rather than 1.0.0 so it
stays above the 1.1.0 label that circulated on manually tagged container images
before releases were automated.
At a glance:
- Requires a Java 25 runtime (was JRE 21) — see the breaking change below before upgrading.
bpmn2linkedarchi renderproduces SVG again; in the consolidated project it had been a stub that printed a message and wrote nothing.- RDF 1.2 reification is now emitted in standard syntax instead of RDF4J-proprietary IRIs.
--versionon every converter now reports the version of the artifact it was built from, read from the JAR manifest. The hardcoded strings had drifted (the BPMN CLI claimed0.2.0-SNAPSHOTwhile the JAR was0.1.0-SNAPSHOT).- Tagged builds now publish a version-tagged container image (
:1.2.0and:v1.2.0) alongside:latestand:<short-sha>, and the distribution tarball name tracks the release instead of being frozen at0.1.0-SNAPSHOT.
Breaking changes — runtime requires Java 25¶
RDF4J was updated from 5.3.1 to 6.0.0, whose artefacts are compiled for Java 25 (class file version 69). That requirement applies to the whole 6.0.0 line, including the milestones, so it cannot be avoided by pinning an earlier 6.x build.
- The JARs now require a Java 25 runtime. Running them on Java 21 fails with
UnsupportedClassVersionError. Previously JRE 21 was enough. - The Gradle toolchain and Kotlin
jvmTargetmoved from 21 to 25 across all modules. - The Gradle wrapper moved from 8.14.4 to 9.6.1, which is needed to target a Java 25
toolchain. One build script change came with it:
fileModewas removed in Gradle 9, so the distribution'sbin/permissions now usefilePermissions { unix("755") }. - Container images moved to
gradle:9-jdk25(build) andeclipse-temurin:25-jre-alpine(runtime).
Fixed — BPMN render actually renders again¶
bpmn2linkedarchi render was a scaffold: it accepted the options and printed
Render command scaffolded — port SVG rendering here. without writing anything. The SVG
renderer itself had already been moved into converter-bpmn/…/svg/, only the CLI wiring
was missing, so the command silently produced no output.
The command is now implemented on top of that renderer and is feature-equivalent to the
pre-consolidation bpmn tool:
- SVG from the BPMNDI geometry in the XML — pools/lanes, sub-processes, event circles, gateway diamonds, task boxes with type icons, sequence-flow arrows.
--base-iriwraps each shape in<a href>pointing at{base}bpmn/{modelId}/element/{id}, the same IRIconvertmints.--shape-config(restored) for per-typefill/stroke/strokeWidth/iconoverrides. A reference file ships asconfig/shape-config-bpmn.yml.--diagrams-index/--diagrams-root/--include-draftsshare the index withconvert: drafts are skipped,idnames the SVG, and index metadata (title, author, version, dates, description) is drawn into an info box, merged over the metadata embedded in the BPMN file.- Per-file failures warn and exit
1; setup failures exit2.
Two behaviour changes against the old bpmn tool: output is named <id>.svg from the
index (was always <filename>.svg) and mirrors sub-directories of the index file: path,
matching the PlantUML renderer and the documented index contract. Without an index the old
flat <filename>.svg naming applies.
Fixed — RDF 1.2 reification now serializes correctly¶
The ArchiMate converter already emitted rdf:reifies for --emit-direct-rel-triples, but
on RDF4J 5.x the triple term was written as an opaque, RDF4J-proprietary IRI that no other
tool can interpret:
RDF4J 6 introduces a real TripleTerm type, and the same code now produces standard
RDF 1.2 syntax:
The API rename ValueFactory.createTriple → createTripleTerm was the only source change
the upgrade required.
Fixed — type-mapping lookups now use the ontology type name¶
--emit-direct-rel-triples and --emit-qualified-rel-triples looked up the ArchiMate
predicates: / qualifiedPredicates: tables with the raw Exchange xsi:type
(ServingRelationship), while the ontology — and every shipped type-mapping file — names
the type Serving. The keys never matched, so both flags silently emitted nothing with the
bundled type-mapping-archimate.yml, which also meant the rdf:reifies bridge never fired.
All mapping lookups now go through one normalisation that resolves the ontology name first
and accepts the raw Exchange name as a fallback, so mappings keyed either way work. This
applies to the class mappings too, where an explicit relationships: override keyed with
the ontology name was previously ignored.
Build — Java 25 toolchain resolution¶
The toolchain requirement is now declared in the repository rather than left to each
developer's local setup. gradle.properties enables toolchain auto-detection and leaves
auto-download off, so a machine without JDK 25 fails with a clear message
("Cannot find a Java installation ... matching: {languageVersion=25}") instead of silently
downloading a JDK mid-build.
Gradle only scans conventional install locations, so a JDK installed elsewhere (Homebrew,
for example) may need an explicit path in your user ~/.gradle/gradle.properties. The
README documents that, and documents applying the foojay resolver as an opt-in alternative
for anyone who does want Gradle to provision the JDK.
Also removed the turtlestar / trigstar entries from the bundled RDF4J service files:
RDF4J 6 folds RDF-star into the base Turtle and TriG parsers and no longer ships those
modules, so listing them raised ServiceConfigurationError: Provider ... not found on
every run.
Breaking changes — emitted RDF¶
The output contract changed in four ways. Queries and shapes written against earlier output need updating.
- Relationship endpoints renamed:
arch:relSource/arch:relTarget→arch:source/arch:target, for BPMN, PlantUML, Structurizr and Backstage. The published core ontology definesarch:source/arch:targetand records therelSource/relTargetpair as a superseded earlier design. The ArchiMate converter already emitted the canonical form. - Labels are language-tagged:
skos:prefLabelis now emitted as anrdf:langString("Name"@en) rather than a plain literal, because the core SHACL shapes require it. A query matching a plain string (?e skos:prefLabel "Name") silently returns nothing; match a variable instead. The tag defaults toenand is configurable with the new--label-language. arch:ModelConceptemitted explicitly: elements and relationships now carry the common supertype alongsidearch:Element/arch:QualifiedRelationship. The core ontology declares it viardfs:subClassOf, but RDF4J applies subclass reasoning tosh:targetClassand not tosh:classconstraint values, so relying on inference failed validation.- BPMN serialization containers are no longer architecture elements:
bpmn:definitionsandbpmn:importskeep theirinfra:types but are no longer dual-typedarch:Element/arch:ModelConcept. They describe the source document, not the architecture, and typing them as elements obliged them to carry a label that BPMN does not require.
Breaking changes — ontologies and shapes are fetched at runtime¶
Metamodel artefacts are no longer bundled in the JARs. All ontologies and SHACL shapes are
fetched from meta.linked.archi, which becomes the single source of truth: a corrected
shape takes effect on the next run, with no converter rebuild or release.
- Removed the 13 vendored TTL files from
converter-bpmn. BPMN shapes are now published at/bpmn/onto-shapes,/bpmn/di-shapes,/bpmn/di-core-shapes,/bpmn/dc-shapesand/bpmn/infra-shapes. --cache-diris renamed--asset-dirand its default moved from.cache/linked-archito<java.io.tmpdir>/linked-archi-assets, so fetched files no longer land in the project tree. A per-converter subdirectory is always appended, so converters cannot collide or read each other's documents.--cache-diris kept as an alias.--no-downloadis renamed--offline, kept as an alias. It now fails naming the exact path it expected, instead of a generic message.--shapesvalues are resolved as a published asset name, a URL, or a local file path. Names come from the new registry: runvalidate --list-assets.--shapes corechanged meaning.coreis now the core ontology; the shapes arecore-shapes. Pipelines passing--shapes corewill load an ontology as their shape graph and validate nothing — a warning is logged for exactly this case.bpmn2linkedarchi ontology --dirnow downloads the published documents instead of exporting bundled copies. It gained--only ALL|ONTOLOGY|SHACL,--overwrite,--asset-dirand--offline.
Breaking changes — validate CLI¶
archimate2linkedarchi validatenow exits1when violations are found. It previously printed the report and exited0, so CI jobs relying on the exit status silently passed. All converters now use the same codes:0conforms,1violations,2execution error.- BPMN's default shape set is
bpmn-shapes+bpmn-infra-shapes, which match the normalized output. The diagram-interchange shapes constrain raw DI structures, so select them with--shapes dionly for output produced with--emit-raw-di-geometry.
Added¶
validatesubcommand for PlantUML, Structurizr and Backstage, which previously had none. The BPMNvalidatewas a stub that printed a placeholder; it now performs real validation.- Shared SHACL engine in
core(ShaclValidator), driving every converter'svalidatethrough one code path: shapes into RDF4J's SHACL shapes graph, ontologies forrdfs:subClassOfreasoning, bulk validation, and a Turtle report. - Shape coverage reporting. A SHACL run reports "no violations" both when data is correct and when no shape matched anything. Every run now logs how many
sh:targetClassdeclarations applied, and warns explicitly when none did, so a vacuous pass is no longer indistinguishable from a clean one. A shape graph with no class targets at all is also flagged, which catches an ontology being passed where shapes were expected. --without-shapeswitches off individual shapes for a run using SHACL's ownsh:deactivated, leaving every other constraint enforced. Accepts a full shape IRI or the aliaseselement-labels,view-labels,labels. BPMN disableslabelsby default, becausenameis optional on virtually every BPMN element includingbpmn:definitionsandbpmndi:BPMNDiagram.--list-assets,--no-store,-o/--ontology,-r/--report,--data-format,--no-rdfs-reasoningand--label-language.- Logging configuration: logs go to stderr so stdout carries only the SHACL report and can be redirected (
validate -i out.ttl > report.ttl). Previously RDF4J DEBUG output was written to stdout and corrupted any captured report.
Fixed¶
validatefailed withUnsupportedRDFormatException: Did not recognise RDF format object Turtle. RDF4J 5.x declares its Rio factories through JPMSmodule-info, and onlyrdf4j-rio-jsonldstill ships a legacy service file. Shading discards themodule-infodeclarations, soServiceLoaderfound only JSON-LD; Turtle, TriG, RDF/XML and N-Triples were present in the JAR but never registered.convertwas unaffected because it instantiates writers directly, whilevalidateparses through the registry. Fixed by authoring explicitMETA-INF/servicesfiles incore(which every converter bundles) and merging them withmergeServiceFiles(). Note thatmergeServiceFiles()andappend()alone cannot fix this: they only act on files that exist, and these modules contribute none.- Fetched-asset filename collision. Names were derived from the last URL segment, so
/core(ontology) and/core-shapesboth resolved tocore.ttl; whichever downloaded first won, and--shapes core-shapescould silently validate against the ontology, reporting zero coverage. Names are now derived from the full URL path. - Corrected the ArchiMate test fixture, which modelled
SystemSoftware --Assignment--> ApplicationComponent. ArchiMate does not permit an Assignment targeting an Application Component; deployment is modelled asSystemSoftware --assignment--> Artifact --realization--> ApplicationComponent. - Removed the programmatic
RDFWriterRegistryregistration inRdfIoand the ineffectiveappend(...)calls in the BPMN build, both workarounds for the registry problem above.
Documentation¶
- New Validation page: the engine, where shapes come from, selecting your own, where fetched documents are stored, offline runs, shape coverage, exit codes and reading a report.
- Documented the emitted-RDF contract in
docs/config/ontology-alignment.md, including the language-tagged labels and the endpoint rename above. - Corrected stale examples across the docs that still showed
arch:relSource/arch:relTargetand untaggedskos:prefLabelliterals.