Skip to content

IRI path shape, for graph consumers

For anything that turns this graph into an artefact with a shape of its own — a documentation site, one file per resource, a static export, a rendered diagram. rdf2docs is the case this page was written for, but every consumer that maps an IRI onto a filename, a route or a directory hits the same questions.

If you only read one thing: a local ID may contain a /, and it is not a filename.

The pattern

Every converter mints through IriMinting, so the layout is the same in all six notations:

{base}{notation}/{modelId}/{segment}/{localId}
Resource IRI
Model {base}{notation}/{modelId}
Named graphs {base}{notation}/{modelId}/graph/semantic · /graph/model · /graph/views · /graph/provenance
Per-source semantic graph {base}{notation}/{modelId}/graph/semantic/{repo}/{path}
Element {base}{notation}/{modelId}/element/{elementId}
Relationship {base}{notation}/{modelId}/relationship/{relationshipId}
View {base}{notation}/{modelId}/view/{viewId}
View node {base}{notation}/{modelId}/view/{viewId}/node/{nodeId}
View link {base}{notation}/{modelId}/view/{viewId}/link/{linkId}
Folder {base}{notation}/{modelId}/folder/{name}

Where those IDs come from, per notation, is Identifiers. This page is only about the shape of the result.

A local ID may contain slashes

{localId} is one or more path segments, not one. Three cases produce more than one today:

Converter Example Why
Backstage element/component/default/payments-service The Backstage entity reference is kind:namespace/name, and all three parts are the identity — a name is unique only per kind within a namespace. Written as a path it splits back unambiguously, which a single segment cannot do because a Backstage name may itself contain -, _ and .
BPMN element/StartEvent_1/messageEventDefinition Definitions the XML leaves without an id (<bpmn:messageEventDefinition />) are anonymous composites of their parent. Nesting under the parent's IRI gives them a stable address without inventing one
all view/order-flow/node/OrderService A view node belongs to one view, and its ID is only unique inside it

A named graph does the same. Where a model is built from several inputs, graph/semantic gains the input's slug — graph/semantic/group-orders/catalog-info-yaml — so the depth under graph/ is not fixed either. Match on the graph/semantic prefix rather than on the whole segment, and reach a graph through ?g prov:wasDerivedFrom ?src rather than by composing its name.

So do not assume localId is the last segment, and do not assume depth. Take everything after /{modelId}/{segment}/ as the ID, or better, treat the whole IRI as the key and read what you need from triples.

For a file-per-resource consumer

A nested ID maps naturally onto nested directories — element/component/default/payments-service.html — and that is the intended reading. What breaks is treating the ID as a single filename: component/default/payments-service.html written into one directory either fails or silently creates paths you did not plan for. Create the intermediate directories, or flatten with a separator of your own choosing; just do not assume the converter picked one for you.

Length limits are per segment

Filesystem limits apply to a single name, not to the whole path: mainstream filesystems cap one component at 255 bytes while allowing a much longer path. IriSegment bounds each generated segment to a 180-character readable head plus a 20-character SHA-256 digest, leaving roughly 50 bytes of headroom for an extension or prefix a consumer adds.

Two consequences:

  • A nested ID is safer than a concatenated one. Three 60-character parts as three segments are fine; as one segment they would be truncated and digested, costing legibility.
  • The digest suffix (…__a1b2c3d4e5f6a7b8c9d0) is content-addressed, not a counter. Equal input gives an equal segment in every run and every process, which is what lets two views of one model mint the same IRI for the same arrow.

Escaping: assume nothing is encoded

IriMinting interpolates the local ID as-is. It does not percent-encode, and no shared validation applies to element, relationship or view IDs.

Converter Local IDs percent-encoded
ArchiMate Yes — every segment goes through URLEncoder.encode(…, UTF_8) with + rewritten to %20
BPMN No. BPMN id values are XML NCNames, so they cannot contain a space or a / anyway
PlantUML Not needed — slugify() restricts IDs to [A-Za-z0-9_\-.], and it is applied to the code as well as to a display name, since an entity written without an as clause takes the display text as its code
Structurizr No, and this is the practical gap: a view key of System Landscape reaches the IRI with a literal space
Backstage No, and none is needed — Backstage constrains a name to [a-z0-9A-Z] separated by [-_.], so no part of a reference can contain a character that needs encoding
LeanIX No, and none is needed — a fact sheet ID is a UUID

A consumer that builds a URL or a filename from an IRI should encode defensively. Structurizr is the one notation where the graph can hand you a character you must handle.

Identity travels in triples, not in the path

The path is designed to be legible, not to be parsed. Everything a consumer needs is stated explicitly:

You want Read this, not the IRI
what kind of thing it is rdf:type — arch:Element, arch:QualifiedRelationship, arch:View, plus the notation class
its name for display skos:prefLabel, with skos:definition for a description
the source system's own identifier skos:notation, and for Backstage the four retained fields bs:kind, bs:namespace, bs:name, bs:entityRef
which model it belongs to arch:inModel, asserted on every element, relationship and view — and dct:isPartOf for its folder
which view depicts it archvis:archElement from the view node back to the element
where it came from the provenance graph — prov:wasDerivedFrom, prov:wasGeneratedBy, prov:qualifiedDerivation

provenance/ is reserved

One top-level segment beside the notation slugs is reserved: {base}provenance/{kind}/… holds the source-file, commit, person, run, agent and derivation nodes. A notation may not be called provenance.

These nodes sit outside {base}{notation}/{modelId}/ because they are identified by their content — a file at a commit, a commit, a run — so one of them is one node across every model and every run. That is what lets an aggregated graph ask "everything derived from this file" against a single subject. Only the statements about them are per-model. See ADR 0008 §2.

The IRIs are legible but still opaque in the sense above: every component is also asserted as a triple, so nothing needs to parse one.

The one exception in this repo is ConceptCollisions, which decides element-vs-relationship by testing the path for /element/ or /relationship/. It does that deliberately, because --ns-core can repoint the core namespace and make an rdf:type test miss — and note that it uses a substring test, so it is indifferent to how deep the ID goes. That is the pattern to copy if you must read the path at all: match a segment, never assume a position or a count.

What is stable and what is not

Stable. Re-converting an unchanged source file gives byte-identical IRIs. No IRI is derived from a filename, a timestamp or an ordering that varies between runs.

Not stable.

  • Backstage element IRIs. They are the entity reference, and a Backstage entity has nothing but its name, so renaming one or moving it between namespaces or kinds relocates every address derived from it. BPMN, ArchiMate, Structurizr and LeanIX carry tool-assigned identifiers that survive a rename; PlantUML mints from the code — the name the source refers to an entity by — so retitling a shape moves its label and not its IRI (ADR 0010).
  • PlantUML relationship IRIs. An arrow has no identifier of its own, so its ID is composed from its endpoints and its label — rewording the label moves it (ADR 0005).
  • Backstage relationship IRIs. rel-7 is whichever relationship was emitted seventh, so inserting one earlier in a catalog file renumbers those after it.
  • Model IDs across repositories. Enforced within one index, unchecked across the repos an aggregation graph pulls from.

A consumer that caches by IRI, or publishes an IRI as a permalink, should expect the first two to move and plan for a redirect. Changing a published identity deliberately is ADR 0003; IdentityLock is what makes an accidental change fail the run instead of silently republishing.