IRI path shape, for graph consumers¶
For anything that turns this graph into an artefact with a shape of its own — a documentation site,
one file per resource, a static export, a rendered diagram. rdf2docs is the case this page was
written for, but every consumer that maps an IRI onto a filename, a route or a directory hits the same
questions.
If you only read one thing: a local ID may contain a /, and it is not a filename.
The pattern¶
Every converter mints through IriMinting, so the layout is the same in all six notations:
| Resource | IRI |
|---|---|
| Model | {base}{notation}/{modelId} |
| Named graphs | {base}{notation}/{modelId}/graph/semantic · /graph/model · /graph/views · /graph/provenance |
| Per-source semantic graph | {base}{notation}/{modelId}/graph/semantic/{repo}/{path} |
| Element | {base}{notation}/{modelId}/element/{elementId} |
| Relationship | {base}{notation}/{modelId}/relationship/{relationshipId} |
| View | {base}{notation}/{modelId}/view/{viewId} |
| View node | {base}{notation}/{modelId}/view/{viewId}/node/{nodeId} |
| View link | {base}{notation}/{modelId}/view/{viewId}/link/{linkId} |
| Folder | {base}{notation}/{modelId}/folder/{name} |
Where those IDs come from, per notation, is Identifiers. This page is only about the shape of the result.
A local ID may contain slashes¶
{localId} is one or more path segments, not one. Three cases produce more than one today:
| Converter | Example | Why |
|---|---|---|
| Backstage | element/component/default/payments-service |
The Backstage entity reference is kind:namespace/name, and all three parts are the identity — a name is unique only per kind within a namespace. Written as a path it splits back unambiguously, which a single segment cannot do because a Backstage name may itself contain -, _ and . |
| BPMN | element/StartEvent_1/messageEventDefinition |
Definitions the XML leaves without an id (<bpmn:messageEventDefinition />) are anonymous composites of their parent. Nesting under the parent's IRI gives them a stable address without inventing one |
| all | view/order-flow/node/OrderService |
A view node belongs to one view, and its ID is only unique inside it |
A named graph does the same. Where a model is built from several inputs, graph/semantic gains the
input's slug — graph/semantic/group-orders/catalog-info-yaml — so the depth under graph/ is not
fixed either. Match on the graph/semantic prefix rather than on the whole segment, and reach a graph
through ?g prov:wasDerivedFrom ?src rather than by composing its name.
So do not assume localId is the last segment, and do not assume depth. Take everything after
/{modelId}/{segment}/ as the ID, or better, treat the whole IRI as the key and read what you need
from triples.
For a file-per-resource consumer
A nested ID maps naturally onto nested directories —
element/component/default/payments-service.html — and that is the intended reading. What breaks
is treating the ID as a single filename: component/default/payments-service.html written into one
directory either fails or silently creates paths you did not plan for. Create the intermediate
directories, or flatten with a separator of your own choosing; just do not assume the converter
picked one for you.
Length limits are per segment¶
Filesystem limits apply to a single name, not to the whole path: mainstream filesystems cap one
component at 255 bytes while allowing a much longer path. IriSegment bounds each generated
segment to a 180-character readable head plus a 20-character SHA-256 digest, leaving roughly 50 bytes
of headroom for an extension or prefix a consumer adds.
Two consequences:
- A nested ID is safer than a concatenated one. Three 60-character parts as three segments are fine; as one segment they would be truncated and digested, costing legibility.
- The digest suffix (
…__a1b2c3d4e5f6a7b8c9d0) is content-addressed, not a counter. Equal input gives an equal segment in every run and every process, which is what lets two views of one model mint the same IRI for the same arrow.
Escaping: assume nothing is encoded¶
IriMinting interpolates the local ID as-is. It does not percent-encode, and no shared validation
applies to element, relationship or view IDs.
| Converter | Local IDs percent-encoded |
|---|---|
| ArchiMate | Yes — every segment goes through URLEncoder.encode(…, UTF_8) with + rewritten to %20 |
| BPMN | No. BPMN id values are XML NCNames, so they cannot contain a space or a / anyway |
| PlantUML | Not needed — slugify() restricts IDs to [A-Za-z0-9_\-.], and it is applied to the code as well as to a display name, since an entity written without an as clause takes the display text as its code |
| Structurizr | No, and this is the practical gap: a view key of System Landscape reaches the IRI with a literal space |
| Backstage | No, and none is needed — Backstage constrains a name to [a-z0-9A-Z] separated by [-_.], so no part of a reference can contain a character that needs encoding |
| LeanIX | No, and none is needed — a fact sheet ID is a UUID |
A consumer that builds a URL or a filename from an IRI should encode defensively. Structurizr is the one notation where the graph can hand you a character you must handle.
Identity travels in triples, not in the path¶
The path is designed to be legible, not to be parsed. Everything a consumer needs is stated explicitly:
| You want | Read this, not the IRI |
|---|---|
| what kind of thing it is | rdf:type — arch:Element, arch:QualifiedRelationship, arch:View, plus the notation class |
| its name for display | skos:prefLabel, with skos:definition for a description |
| the source system's own identifier | skos:notation, and for Backstage the four retained fields bs:kind, bs:namespace, bs:name, bs:entityRef |
| which model it belongs to | arch:inModel, asserted on every element, relationship and view — and dct:isPartOf for its folder |
| which view depicts it | archvis:archElement from the view node back to the element |
| where it came from | the provenance graph — prov:wasDerivedFrom, prov:wasGeneratedBy, prov:qualifiedDerivation |
provenance/ is reserved¶
One top-level segment beside the notation slugs is reserved: {base}provenance/{kind}/… holds the
source-file, commit, person, run, agent and derivation nodes. A notation may not be called provenance.
These nodes sit outside {base}{notation}/{modelId}/ because they are identified by their content — a
file at a commit, a commit, a run — so one of them is one node across every model and every run. That is
what lets an aggregated graph ask "everything derived from this file" against a single subject. Only the
statements about them are per-model. See
ADR 0008 §2.
The IRIs are legible but still opaque in the sense above: every component is also asserted as a triple, so nothing needs to parse one.
The one exception in this repo is ConceptCollisions, which decides element-vs-relationship by testing
the path for /element/ or /relationship/. It does that deliberately, because --ns-core can
repoint the core namespace and make an rdf:type test miss — and note that it uses a substring test,
so it is indifferent to how deep the ID goes. That is the pattern to copy if you must read the path at
all: match a segment, never assume a position or a count.
What is stable and what is not¶
Stable. Re-converting an unchanged source file gives byte-identical IRIs. No IRI is derived from a filename, a timestamp or an ordering that varies between runs.
Not stable.
- Backstage element IRIs. They are the entity reference, and a Backstage entity has nothing but its name, so renaming one or moving it between namespaces or kinds relocates every address derived from it. BPMN, ArchiMate, Structurizr and LeanIX carry tool-assigned identifiers that survive a rename; PlantUML mints from the code — the name the source refers to an entity by — so retitling a shape moves its label and not its IRI (ADR 0010).
- PlantUML relationship IRIs. An arrow has no identifier of its own, so its ID is composed from its endpoints and its label — rewording the label moves it (ADR 0005).
- Backstage relationship IRIs.
rel-7is whichever relationship was emitted seventh, so inserting one earlier in a catalog file renumbers those after it. - Model IDs across repositories. Enforced within one index, unchecked across the repos an aggregation graph pulls from.
A consumer that caches by IRI, or publishes an IRI as a permalink, should expect the first two to move
and plan for a redirect. Changing a published identity deliberately is
ADR 0003; IdentityLock is what makes an accidental change fail
the run instead of silently republishing.
Related¶
- Identifiers — where each notation's IDs come from, and the three levels of "ID"
- ADR 0001: Where model identity lives
- ADR 0003: Changing a published identity
- Aggregation graph — how models are merged and overwritten