Skip to content

ADR 0005 — What names a PlantUML relationship

Status: Accepted, implemented

Scope: how the local ID of an arch:QualifiedRelationship is synthesised for the PlantUML converter, which decides when two arrows drawn in two views are one resource and when they are two.

Depends on: ADR 0002 for why views of one model share a namespace at all; ADR 0003 for the principle that a published identity does not change silently — a principle this ADR has to invoke without being able to use its mechanism.

Context

Four of the five converters read a relationship's identifier out of the source: BPMN and ArchiMate take a tool-assigned id, Structurizr takes the workspace JSON id, Backstage counts. PlantUML has nothing to read. A relationship there is a line in a text file with no identity of its own, so the converter has to make one up, and whatever it makes up becomes a published IRI.

It used to be the pair of endpoints, {source}__{target}, with _2, _3 … appended when the same pair was connected more than once. That is wrong in a way that is easy to miss: the endpoints name a pair, not an edge between them. Two participants in a sequence diagram exchange many messages, and two classes can be joined by more than one kind of relationship.

The suffix that was supposed to separate them was counted per file, while the IRI namespace is per model. So two views of one model that connected the same pair each restarted at the bare {source}__{target} and collided. A synchronous call in one view and its reply in another became a single uml:Message carrying two skos:prefLabel values and two uml:messageSort values:

<…/plantuml/order-domain/relationship/ClientApp__OrderService> a uml:Message ;
    skos:prefLabel  "create order" , "order created" ;
    uml:messageSort uml:SynchCall ,    uml:Reply .

Forty-nine messages in one corpus were in that state. umlsh:MessageShape caught the duplicated messageSort as a cardinality violation, days later, on a merged graph with no route back to the file that caused it — and caught nothing at all when the two colliding arrows were both calls, in which case only the labels merged.

Options

A. Add the message sort to the ID

ClientApp__OrderService__SynchCall and ClientApp__OrderService__Reply. Fixes exactly the 49 reported cases, because every one of them was a call colliding with a reply. Leaves two calls in two views still merging their labels, so it treats the symptom the shapes happened to notice.

B. Scope the ID to the view

call__ClientApp__OrderService and reply__ClientApp__OrderService, on the reasoning that a UML Message belongs to an Interaction and one .puml is one Interaction. Makes the collision structurally impossible.

But it splits what should merge: the same arrow drawn in two diagrams becomes two resources, so a graph cannot answer "does ClientApp call OrderService" without unioning per-view edges. It also introduces a third identity rule into one converter — elements by name, messages by view, structural relationships by endpoints — and leaves participants model-scoped while messages become view-scoped.

C. Derive the ID from the label

ClientApp__OrderService__create_order. The label is what an author writes to tell one arrow from another, so it is the thing that actually distinguishes them.

The asymmetry with elements is deliberate and worth stating. An element has a name the source refers to it by — its code — so its ID is taken from that rather than from the text it displays (ADR 0010). An arrow has no such name: there is nothing in A -> B : create order that PlantUML or an author uses to refer to the arrow itself, so the only authored thing that distinguishes it from a second arrow between the same pair is its label. Option C takes the most identifier-like thing available, which for a relationship is the label and for an element is not.

Decision

C, with the kind as a fallback and order of appearance as a last resort.

notation = {source}__{target}__{discriminator}

discriminator = the label, slugified and lower-cased      when the arrow is labelled
              = the kind — call / reply / async, or the   when it is not
                relationship type for a structural edge
+ _2, _3 …      only for arrows identical in all of the above

local ID = notation.take(180) + "__" + sha256(notation).take(20)

The second line is the amendment described under Bounding the ID: the composed notation carries the identity, but it has no length limit, so the ID is that notation truncated with a digest of the whole of it appended. The untruncated notation is published as skos:notation.

Each part earns its place. The label does the work in the normal case. The kind carries unlabelled arrows, where PlantUML offers nothing else — A -> B and A --> B differ only in being a call and a reply. The counter survives only for arrows a reader could not tell apart either, and it is the only remaining order-dependence.

Three consequences of the decision, stated because each was a choice:

The rule applies to structural relationships too, not only sequence messages. A --> B: owns and A --> B: manages were A__B and A__B_2, positional; they are now A__B__owns and A__B__manages. One rule per converter is simpler than one rule per diagram kind.

The label is lower-cased into the ID, and left alone in the graph. create order and Create Order are the same message written twice and must not be two IRIs — an edited capital is not a new resource. The skos:prefLabel keeps the author's text, spacing and all, because rewriting someone's prose is not the converter's business.

The per-file counter stays per file. It was the bug's proximate cause only because the base ID did not discriminate. Now that it does, resetting per file is what makes the same arrow in two views mint one IRI and merge — which is the point of grouping views into a model.

Consequences

  • The same message drawn in two views is one arch:QualifiedRelationship with an archvis:Link in each view. Querying "does ClientApp call OrderService" needs no per-view union.
  • A call and a reply between one pair are two relationships, in one file or across several.
  • Relationship IRIs are stable under reordering. Inserting an arrow no longer renumbers the ones after it, except among arrows that are indistinguishable anyway.
  • Relationship IRIs move. Every published PlantUML relationship IRI changes with this release. See below.
  • Editing a label moves that relationship's IRI, exactly as renaming a shape moves an element's. Case and outer whitespace are absorbed; wording is not.
  • Two views spelling one message differently are reconciled with a warning, not refused. The ID folds case and punctuation, so both views always meant one message; refusing them would reject exactly the differences the ID derivation deliberately discards. ConceptCollisions keeps the first spelling, drops the other — two skos:prefLabel values in one language is not legal SKOS — and names both files. A genuine disagreement, about the type or about different words, still fails the run.

This corrects an earlier decision. The first implementation made any spelling difference fatal, on the reasoning that only the sources can settle which spelling is right. That is true and still worth saying, but a fatal error was the wrong instrument: on a 153-view corpus it aborted 13 runs and every one was cosmetic — eight differing in case, four in optional parentheses, one in where the bold markers sat — with no genuine conflict among them. An identifier that folds case cannot then demand case agreement. - uml:messageKind uml:Lost distinguishes A ->x B from A -> B, which were previously identical in the graph. It does not participate in identity, so a lost and an ordinary call sharing a label still separate by counter.

Bounding the ID (amendment)

Deriving the ID from content fixed identity and broke length. Composing two endpoint IDs and a label concatenates three pieces of unbounded author-supplied text, and while an IRI has no length limit, the things built from one do: a path segment becomes a filename wherever the graph is materialised as documents, and 255 bytes is the ceiling on every mainstream filesystem.

It is reached by ordinary practice. An author names a participant after the URL of its API contract, which is what you do when the participant is an API with a published specification:

https_git.example.org_group-one_spec-registry_alpha-api-contract_-_blob_main_openapi.yaml_alpha-api
  __ https_git.example.org_group-two_platform_beta-proxy_-_blob_develop_beta-2.0.yml_beta-endpoint
  __ wait_for_confirmation

Measured on the same corpus before and after the label-derived scheme, relationship segments over the limit went from 16 to 116 and the longest from 314 to 644 bytes. So the scheme made a pre-existing problem seven times more common, and the document generator aborts on every one of them.

Decision: truncate the notation to 180 characters and append a 20-hex-character SHA-256 digest of the whole of it. The notation is published as skos:notation.

Every identity property above survives untouched, because the digest is taken over exactly the string that used to be the ID: the same arrow in two views still digests the same and still merges, a call still separates from a reply, and the counter still separates indistinguishable arrows. What changes is that the address stops growing with its inputs.

Three choices inside that, each of which had a plausible alternative:

Every ID is hashed, not only the long ones. Leaving short IDs alone would keep most IRIs exactly as they are and change only the ones that break. It was rejected because it makes the scheme depend on the length, and that boundary is invisible: an author who lengthens a label from 178 to 182 characters would silently move a published IRI from one scheme to the other. One unconditional rule costs a single declared change now; a conditional rule costs an unbounded number of undeclared ones later.

A digest, not a counter. A counter is shorter but assigned by order of appearance, so inserting an arrow would renumber the ones after it. A digest is a function of content alone, which is what makes two views agree.

A readable prefix is kept. A bare digest would be shorter and completely opaque. Keeping the first 180 characters costs nothing against the limit and means an operator reading a graph can still see which participants and which label an ID came from without resolving it.

The residual: element IDs are unbounded. An element ID is a single slugified code, so it does not concatenate, and no element segment exceeded the limit on the corpus — but nothing prevents one, since an entity written without an as clause takes its display text as its code. Bounding it would re-address every element in every model, for a problem not yet observed, so it is left as a known risk rather than fixed pre-emptively. Giving a long-named participant a short alias avoids it at the source and shortens the relationship segments composed from it.

On the identity change, and why there is no lock for it

This breaks published IRIs, and unlike a model or view rename it cannot be declared. IdentityLock records modelId, viewId and lifecycle state only, so nothing detects a relationship ID change and --allow-identity-change has nothing to authorise. identifiers.md marks element and relationship IDs as unvalidated, which is the same gap seen from the other side.

That is accepted rather than solved. Extending the lock to every relationship in every model would make it a full inventory of the graph, which is a different artefact with different costs, and the alternative to changing the scheme is keeping one that merges distinct messages. The migration is therefore a republish: consumers holding a …/relationship/ClientApp__OrderService IRI must re-resolve. No owl:sameAs bridge is emitted, because the old IRI frequently denoted two messages at once and there is nothing single to point it at.

Recorded as an open question below.

Example

' call.puml — view "order-request"
ClientApp -> OrderService: create order

' reply.puml — view "order-reply"
ClientApp --> OrderService: order created
<…/plantuml/order-domain/relationship/ClientApp__OrderService__create_order__60f79e38d012a705ddc3>
    a uml:Message ; uml:messageSort uml:SynchCall ;
    skos:notation  "ClientApp__OrderService__create_order" ;
    skos:prefLabel "create order"@en .

<…/plantuml/order-domain/relationship/ClientApp__OrderService__order_created__086887b6418401eac9dd>
    a uml:Message ; uml:messageSort uml:Reply ;
    skos:notation  "ClientApp__OrderService__order_created" ;
    skos:prefLabel "order created"@en .

The trailing 20 characters are the digest; the part before it is the notation, repeated in full in skos:notation. Where the notation is longer than 180 characters the IRI keeps only its first 180 and skos:notation is the only complete copy.

Unlabelled arrows between the same pair, and the lost-message marker — notations shown, each with a digest appended in the IRI:

A -> B          → notation A__B__call
A --> B         → notation A__B__reply
A ->> B         → notation A__B__async
A ->x B: ping   → notation A__B__ping    + uml:messageKind uml:Lost

Open questions

  • Should the lock cover relationship IDs? Today it cannot detect this class of change. A cheaper middle ground than a full inventory would be recording a scheme version per model, so a change of rule is detectable even when individual IDs are not.
  • Unlabelled arrows remain weakly identified. Several unlabelled calls between one pair separate only by order of appearance. Nothing in PlantUML fixes this; a diagram that needs stable IRIs for those arrows needs labels.
  • Messages through gates are still dropped. [-> B and A ->] parse as MessageExo rather than Message, and the parser reads only messages, so they never reach the graph. That is also why uml:messageKind uml:Found is never emitted — the notation that would produce one is not read.
  • Element IDs have no length bound. See the residual noted under Bounding the ID: an element ID does not concatenate, so it has not been observed to exceed the filename limit, but nothing enforces it.
  • Should a spelling variance be optionally fatal? A corpus that wants strict agreement between views has no way to ask for it now. A --fail-on-label-variance switch would let it opt in, rather than every corpus paying for the strict reading.