Skip to content

ADR 0008 — What the provenance graph says about how a graph was made

Status: Accepted, implemented

Implementation summary:

  • The conversion event is described with PROV-O: a prov:Activity for the run, a prov:SoftwareAgent for the container image that performed it, and a prov:Entity for the source file at the commit it was read from.
  • The generated view (or model) carries prov:wasGeneratedBy the run and prov:wasDerivedFrom the commit-scoped source entity.
  • Dublin Core keeps describing the diagram: title, author, dates, description. PROV describes the event. dct:created therefore stops meaning two things at once.
  • A node identified by its content — a source file at a commit, a commit, a run, an agent — lives under {base}provenance/{kind}/…, one node across every model and every run. Statements about it go into each model's own provenance graph.
  • A source file is identified by four facts: repository, path within it, commit and content digest.
  • Each named graph a conversion writes is described as a prov:Bundle: what generated it, from what, and when. Where a model is built from several inputs its semantic graph is split one per input, so "which input asserted this fact" is answerable. The curated model has its own graph, derived from the diagram index. Every concept carries arch:inModel, so the pieces reassemble from triples alone. See decision 8.
  • IRIs are opaque. What identifies a thing travels in triples, not in the IRI.
  • The activity and its agent are emitted by every conversion. They replace what dct:created and dct:creator used to say about the run, so making them optional would drop that rather than relocate it.
  • The git description inside them is what is opted into: --git-provenance auto reads GitLab's CI variables, overridable by explicit flags, off by default.
  • The three provenance emitters are not consolidated. See Why the emitters stay separate.

Scope: what the provenance named graph says about how a graph came to be — the run, the software that performed it, the sources it read, and the graphs it wrote. Descriptive metadata about the diagram is out of scope and unchanged.

Depends on: ADR 0001, which assigns IRIs to the resources this ADR describes.

Related: ADR 0004 set the precedent this ADR follows for vocabulary: reuse a published one (there ADMS, here PROV-O) rather than mint a term in the arch: namespace.

Context

Every converter writes a provenance named graph at {base}{notation}/{modelId}/graph/provenance. What lands in it depends entirely on which converter ran.

dct:created dct:source dct:creator PROV
PlantUML, Structurizr, Backstage (BaseLinkedArchiEmitter) wall clock filename literal "{notation}2linkedarchi" none
BPMN (own emitter) wall clock filename literal "bpmn2linkedarchi" none
ArchiMate (own emitter) — — — prov:generatedAtTime

Three things are wrong with this, and they are not stylistic.

dct:created means two different things on the same subject

BaseLinkedArchiEmitter.emitProvenance writes the conversion time:

val now = OffsetDateTime.now().format(DateTimeFormatter.ISO_OFFSET_DATE_TIME)
mb.namedGraph(provenanceGraph).subject(fileSubject)
    .add(vocab.dctCreated, vf.createLiteral(now))

and then, if the index declared created:, writes the author's date under the same predicate on the same subject in the same graph:

metadata.created?.let { … .add(vocab.dctCreated, vf.createLiteral(it)) }

RDF is a set of statements, so this is not an overwrite. Both survive. A consumer asking when the diagram was created gets two answers and no way to tell which is the authoring date and which is an artefact of the build that happened to run last night. The BPMN emitter reproduces this, with the wrinkle that its second date goes through typeMapping.metadataPredicateIri, so the two only collide when the type mapping does not override created.

Dublin Core is not at fault. dct:created is a perfectly good predicate for when the diagram was created. It was asked to also carry when this file was generated, which is a fact about an event, not about a diagram.

The agent is a string, and the thing that ran is not that string

dct:creator "bpmn2linkedarchi" names a wrapper script. There is no library, package or release by that name. The released artefact is a container image holding all five converters, tagged $CI_REGISTRY_IMAGE:$CI_COMMIT_SHORT_SHA and, on tags, :$CI_COMMIT_TAG.

So the string cannot answer the only question worth asking of it — exactly which build produced this? — and there is nowhere to put the answer, because a literal cannot carry a version, a digest and a source revision at once.

dct:source is a bare filename

dct:source "process.bpmn" does not say which repository, which revision, or which of several same-named files across models. Regenerating the graph, or auditing what a published IRI was derived from, needs all three.

Why PROV-O rather than more Dublin Core

Dublin Core has no vocabulary for an activity that used an input and produced an output, performed by a piece of software, at a time. PROV-O is the W3C recommendation for precisely that, RDF4J already ships the PROV vocabulary constant, and the ArchiMate converter already publishes prov:generatedAtTime — so the prefix is in released output and this is an extension of existing practice rather than a new dependency.

For the git half, the mapping follows the established ones rather than inventing a third: Git2PROV (De Nies et al., ISWC 2013) and GitLab2PROV. Both treat commits as activities, file versions as entities, and users as agents. Content was rephrased for compliance with licensing restrictions.

Decision

1. PROV-O describes the event; Dublin Core describes the diagram

dct:created reverts to one meaning: the date the author declared. The conversion time moves onto the activity as prov:startedAtTime / prov:endedAtTime.

The model keeps prov:generatedAtTime, which ArchiMate already publishes and which the other converters gain. It duplicates the activity's end time deliberately: it is the cheapest possible answer to "when was this made" for a consumer who should not have to traverse to the activity to get it.

2. A node identified by its content lives in one base-scoped namespace

Statements go into the model's own provenance graph. The nodes those statements are about do not belong to it: a node identified by its content — a file at a commit, a commit, a run — is one node across every model and every run, under a namespace derived from the base IRI.

Node Keyed by IRI
source file repository, commit, path {base}provenance/source/{repo}/{sha}/{path}
commit the sha {base}provenance/commit/{sha}
author their canonical email address, percent-encoded; a digest of the name where no address is known {base}provenance/person/{encoded-email}
run job URL, else start time and image {base}provenance/run/{token}
agent a digest of the image reference {base}provenance/agent/{digest}
derivation source and run {base}provenance/derivation/{repo}/{sha}/{path}/{run}

Why the base and not the graph. A provenance graph IRI is per model, so naming nodes as fragments of it makes a file read by two models into two nodes with nothing relating them, and makes one run into one activity per model. These graphs are published for an aggregating consumer, and the questions that consumer asks — everything derived from this file, which run produced this — are asked of a single subject or not at all. A prov:Bundle would not help: bundles separate provenance about provenance, not one run's description of a model from another's.

Why one reserved segment. {base} holds notation slugs — {base}bpmn/…, {base}backstage/… — so sources at {base}source/ would reserve a word among them, and this class of node needs six. Everything sits under {base}provenance/{kind}/…, which reserves one; a notation may not be called provenance. Spelled out rather than prov/, so it is not read as the prov: prefix.

Each model's provenance graph is self-contained. Every statement a model's description needs is in that model's graph, so a TriG consumer reading one graph sees a whole description, and a flat merge coalesces the repeated statements. Where one output file holds several models, the repetition across them is accepted. dct:identifier on the run carries the CI job URL.

Shas in an IRI are twelve characters. The sha is what separates two revisions of one file, and these nodes are shared across an aggregate, so the population it discriminates within is every commit that has touched that path. The full value is always dct:identifier, on both the source entity and the commit activity.

The IRIs are legible but not load-bearing, in the sense §3 means: every component is also asserted as a triple, so nothing has to parse or resolve one.

3. IRIs are opaque; identity travels in triples

{base}provenance/source/… is not required to parse or resolve. Everything its components carry is also a statement, so a consumer reads triples and never string-matches an IRI.

A source file is identified by four facts, each on its own predicate:

Fact Predicate Where it comes from
which repository dct:isPartOf, target typed prov:Collection the source map's repo:, or CI_PROJECT_URL
where in it schema:name the source map's path:, or the repo-relative path
which revision dct:identifier the commit sha
which bytes schema:sha256 the source map's contentSha256:

The repository is not optional detail. schema:name is the path inside a repository, and for a pulled input that repository is not the one doing the converting: every Backstage descriptor in every repository is called catalog-info.yaml, so the path alone discriminates nothing across a catalog collected from many sources. Repository and path together name the file; the commit picks the revision.

dct:isPartOf because that is the relation — a file is part of a repository, the same statement the term already makes about a view and its folder. The repository URL is the node IRI directly rather than a minted proxy: unlike a file at a commit, a repository has a stable published IRI that other graphs cite too, so a local proxy would fragment it for nothing. It is typed prov:Collection and not also prov:Entity, since PROV-O makes Collection a subclass and ?s a prov:Entity is how a consumer enumerates source files.

schema:sha256 and not a second dct:identifier, which carries the commit — two identifiers on one subject would reintroduce on this node the ambiguity §1 removes from the diagram. A digest and a commit answer different questions: which bytes, and which revision. Two commits sharing a digest are a file that moved commit without changing, which an incremental re-pull produces and a commit sha alone cannot show.

The source IRI is keyed on all three of repository, commit and path. Keying it on fewer lets two files share a node, and a node shared by files from two repositories would carry two dct:isPartOf values — saying a file is part of two repositories, which is false rather than merely imprecise.

4. The agent is the container image

<…/provenance/agent/7d1c40ba>
    a                      prov:SoftwareAgent, schema:SoftwareApplication ;
    schema:name            "bpmn2linkedarchi" ;
    schema:softwareVersion "1.4.0" ;
    dct:identifier         "registry.gitlab.com/linked-archi/tools/converters/linked-archi-converters@sha256:89ab…" ;
    schema:url             <https://gitlab.com/linked-archi/tools/converters/linked-archi-converters/-/commit/b28afe4…> .

schema:name is the entrypoint, because one image contains five converters and which one ran is part of the answer. dct:identifier is the image reference, digest-pinned when the caller pinned one. schema:url carries the converter source revision, as a URL rather than a bare sha so no term has to be minted for it — and not as a second dct:identifier, since two identifiers on one subject would repeat the ambiguity this ADR exists to remove.

schema:url and not rdfs:seeAlso, which is what this carried until 1.3.0. seeAlso says only "here is related information" and leaves the relationship to be guessed; schema:url says the object is the URL of the subject, which is the statement intended. The same objection retired seeAlso from the source entity in §5 — there prov:alternateOf already carried the link, so the term was dropped rather than replaced.

Resolution order, so the claim degrades honestly instead of guessing:

  1. --image-ref
  2. LA_IMAGE_REF, else CI_JOB_IMAGE
  3. a value baked in at image build time
  4. nothing — emit schema:softwareVersion and make no image claim

5. prov:wasDerivedFrom points at the source file at its commit

<…/bpmn/order-domain/view/current>
    prov:wasGeneratedBy <…/provenance/run/9c2f1e04> ;
    prov:wasDerivedFrom <…/provenance/source/linked-archi-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .

<…/provenance/run/9c2f1e04>
    a                      prov:Activity ;
    prov:startedAtTime     "2026-08-01T09:14:22Z"^^xsd:dateTime ;
    prov:endedAtTime       "2026-08-01T09:14:25Z"^^xsd:dateTime ;
    prov:wasAssociatedWith <…/provenance/agent/7d1c40ba> ;
    prov:used              <…/provenance/source/linked-archi-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> ;
    dct:identifier         "https://gitlab.com/linked-archi/models/-/jobs/9182734" .

<…/provenance/source/linked-archi-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn>
    a                   prov:Entity ;
    schema:name         "models/bpmn/order/process.bpmn" ;
    dct:isPartOf        <https://gitlab.com/linked-archi/models> ;
    dct:identifier      "8c44d1a4e9…" ;
    prov:wasGeneratedBy <…/provenance/commit/8c44d1a4e9f0> ;
    prov:alternateOf    <https://gitlab.com/linked-archi/models/-/blob/8c44d1a4e9…/models/bpmn/order/process.bpmn> .

<https://gitlab.com/linked-archi/models>
    a           prov:Collection ;
    schema:name "linked-archi/models" .

<…/provenance/commit/8c44d1a4e9f0>
    a                      prov:Activity ;
    prov:endedAtTime       "2026-07-30T12:01:02Z"^^xsd:dateTime ;
    prov:wasAssociatedWith <…/provenance/person/a.author%40example.org> ;
    dct:identifier         "8c44d1a4e9…" ;
    schema:url             <https://gitlab.com/linked-archi/models/-/commit/8c44d1a4e9…> .

<…/provenance/person/a.author%40example.org>
    a            prov:Person ;
    schema:name  "A. Author" ;
    schema:email "a.author@example.org" .

The chain diagram → generated by this run → of this file at this commit → committed by this person is then one traversal.

A derivation also says which run it went through

A derived resource carries both the plain and the qualified form, because the plain form cannot name the activity a derivation passed through and PROV does not license inferring it: prov:used and prov:wasGeneratedBy follow from a derivation, not the reverse. Naming it matters because a provenance graph holds many activities — one conversion run plus one commit activity per distinct commit — and an aggregate holds the runs of every conversion that contributed to it.

<…/element/component/default/order-service>
    prov:wasDerivedFrom      <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
    prov:qualifiedDerivation <…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04> .

<…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04>
    a                prov:Derivation ;
    prov:entity      <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
    prov:hadActivity <…/provenance/run/9c2f1e04> .

One derivation node per (source, run), shared by every resource derived from that file. A prov:Derivation states which entity was used and which activity did it, and names no derived entity, so one node per file per run is what the record says rather than a compression of it. The alternative is a node per resource, which costs a 500-entity catalog 500 nodes to state one fact 500 times.

An IRI, not a blank node. A blank node is relabelled on every parse, so two runs' descriptions of one derivation could not be recognised as one — and being able to recognise that is why the qualified form is here at all.

prov:hadUsage and prov:hadGeneration are not emitted. Supplying both makes this a precise-1 derivation under PROV-CONSTRAINTS, with ordering obligations between that usage and that generation. prov:hadGeneration is per-derived-entity by construction, so it is incompatible with the shared node, and its content would restate a prov:wasGeneratedBy the resource already carries; prov:hadUsage is shareable but states what {run} prov:used {source} states.

The property is prov:entity, lowercase — prov:Entity is the class. RDF4J exposes them as PROV.ENTITY_PROP and PROV.ENTITY, and a class IRI in the predicate position is valid RDF that means nothing, so the emitter's test asserts the predicate's IRI string rather than the constant.

Three predicates, three different jobs

dct:source, prov:alternateOf and prov:wasDerivedFrom all relate an output to its input somewhere in this block. They are not interchangeable and none of them is on the source entity by accident.

Predicate Subject Object Says
prov:wasDerivedFrom the output — model, view, each named graph, each concept the source entity this was produced from that file
dct:source the output — the view, or the model where there is no view the blob IRI, or the input's filename where no URL can be built the same thing in Dublin Core, and the only form of it a run with provenance off still emits
prov:alternateOf the source entity the blob IRI these two IRIs denote one document

The source entity carries no dct:source. It began as one, holding a bare filename, and a bare dct:source "catalog-info.yaml" is the same string for every descriptor in a pulled catalog: a consumer holding it cannot get back to the file. Repointing it at the blob URL fixed the value and broke the meaning — this node is the file at that commit, which is what prov:alternateOf says of it, what schema:sha256 digests and what the commit activity generated. Saying it was derived from that same blob would make the file derived from itself. The derivation holds of the output, and prov:wasDerivedFrom states it there, on the subject it is true of.

dct:source stays on the output, where "a related resource from which the described resource is derived" is exactly the relation. Its value is the commit-pinned blob IRI where the run can build one and the input's bare filename where it cannot — which is every run with provenance off, and the reason the predicate is kept rather than left to prov:wasDerivedFrom. Consumers must therefore accept either node kind; FILTER(isIRI(?o)) separates them. DCMI sanctions both: range rdfs:Resource, with a URI or "a string conforming to a formal identification system" recommended.

schema:name carries the identifying path. Recovering it from the blob URL would mean parsing a forge URL, which §3 says not to do. Same choice as the repository node above, which carries its group/subgroup/name path under schema:name. Not dct:title: a title is chosen for display, a path is not a title. The path is repo-relative, which requires the repository root: CI_PROJECT_DIR under CI, git rev-parse --show-toplevel locally. Where neither is available the path is emitted as given, which for a file with no repository is its bare name.

prov:alternateOf states the identity, and no rdfs:seeAlso sits beside it. Not owl:sameAs, which would licence inferences about the document this block does not intend, and not prov:atLocation, which says where to look rather than what a thing is. rdfs:seeAlso used to carry the same URL beside it and was dropped: it asserts no relationship between the two IRIs, so as a second link to the same blob it said nothing. Where a URL is the only pointer to a resource — the forge's page for a commit, the page for the converter revision — it is carried as schema:url, which states what seeAlso left implicit.

Enumerating source files is ?s a prov:Entity ; schema:name ?path. The blob is typed prov:Entity too, since prov:alternateOf has prov:Entity at both ends, and it carries nothing else — so the type alone over-matches and the name is what discriminates. The repository and the commit author are named but are not entities; a graph is typed prov:Bundle only. This replaces the older advice to filter on dct:source, which no longer appears on the entity.

A person is identified by their address, not by their name

<…/provenance/person/a.author%40example.org>
    a            prov:Person ;
    schema:name  "A. Author" ;          # what a human reads
    schema:email "a.author@example.org" .   # what identifies them

The node is keyed on the canonical email address: trimmed, lower-cased and percent-encoded as one IRI path segment. A display name is not an identity. It changes when someone marries, standardises on a different form, or configures one machine as J. Doe and another as Jane Doe — and keyed on the name, each variant is a different person, so no query can gather one author's commits. Git offers no other stable handle: a commit records a name and an address, and no username or numeric id. Canonicalising makes Jane@Example.org and jane@example.org one person. Percent encoding is reversible, so punctuation does not collapse distinct addresses: a+b@example.org and a-b@example.org remain different people. An exceptionally long encoded address keeps a readable prefix plus the full SHA-256 so the segment remains filesystem-safe.

The address is published, as schema:email. This reverses the position this ADR held until 1.3.0, that an address is "personal data in a published graph for no gain in provenance". The gain is stable identity, which is not available any other way, and the graph is internal — the converters cannot read a private catalog without being inside the network that holds one. It is also the judgement the rest of the suite had already made: the LeanIX converter keys arch:Stakeholder on a lower-cased address and publishes schema:email, and Backstage publishes bs:email from spec.profile. The provenance person was the outlier.

schema:email and not bs:email or lmm:subscriberEmail, each of which is one source's own term. LeanIX already chose schema:email so that people join across notations on more than a label; a git author and a LeanIX stakeholder are now joinable on the same predicate.

The address in the IRI, readable, rather than hashed. Following the same reasoning the LeanIX converter records: the graph carries the address as schema:email anyway, so hashing the IRI would hide nothing while making the resource unrecognisable. Worth knowing that an IRI travels further than a literal.

Where no address is known the node stays a digest of the name, exactly as before — a local git log on a repository with user.email unset, or a source map written before addresses were recorded. An address with no name is still a person: it identifies them, which is the harder half, and schema:name is simply omitted.

This does not yet join a git author to a Backstage User or a LeanIX stakeholder as one node. Those are model-scoped IRIs and reconciling them is a separate decision — see todo/UPSTREAM-REQUEST-provenance-source-identity.md. schema:email is what makes the join expressible.

A source that is not a file at a commit — a LeanIX or Backstage API snapshot — needs a different shape, which is not yet decided. See todo/converters-provenance-doc-recommendations.md, item B.

6. Git facts come from GitLab, off by default

--git-provenance auto reads CI_PROJECT_URL, CI_COMMIT_SHA, CI_COMMIT_TIMESTAMP, CI_COMMIT_AUTHOR, CI_JOB_IMAGE and CI_JOB_URL, falling back to git rev-parse HEAD and git log -1 for local runs. Explicit flags (--source-repo-url, --source-commit, --source-commit-time, --source-author) override.

GitLab-only is a deliberate limit, not an oversight: it is the only CI in use. The detection lives in one place in core with the environment injected, so it is testable without a repository or a pipeline, and a second forge is a second detector rather than a redesign.

Default off, because it changes published output.

7. --run-timestamp=commit|now|none

Deriving the activity time from the commit rather than the clock makes the provenance graph byte-stable for an unchanged model at an unchanged commit, which it has never been. now keeps today's behaviour; none omits the times for output that must not carry any.

It does not yet make the whole file byte-stable, and it is worth being exact about why. Folder membership is emitted with blank nodes, whose identifiers are freshly generated per run, so two conversions of an identical model at an identical commit still differ on those lines. Measured on a one-class diagram: the provenance graph is identical between runs and contains no blank nodes, while two elsewhere in the file change every time. Deterministic blank nodes — or schema:ListItem resources with minted IRIs — are what full reproducibility needs, and that is a separate decision about folder emission, not about provenance.

8. Each named graph is a prov:Bundle that says what it was lifted from

Derivation on a resource answers "where did this thing come from". It cannot answer "where did this fact come from", because a resource two inputs contribute to carries two prov:wasDerivedFrom values and nothing pairs a value with an input.

So the graph is a subject too. Every named graph a conversion writes is described in that model's provenance graph:

<…/backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml>
    a                    prov:Bundle ;
    prov:wasGeneratedBy  <…/provenance/run/9c2f1e04> ;
    prov:wasDerivedFrom  <…/provenance/source/group-orders/4f2c1ab8e0d1/catalog-info-yaml> ;
    prov:generatedAtTime "2026-08-29T09:03:11Z"^^xsd:dateTime .

prov:Bundle and not also prov:Entity. PROV-O makes Bundle a subclass of Entity, so the second type adds nothing a reasoner needs, and it adds something a non-reasoning consumer does not want: ?s a prov:Entity is how one reaches the inputs of a run, and a graph is not an input. The repository node is typed prov:Collection alone for the same reason.

The semantic graph is partitioned by input where a model has more than one. graph/semantic/{repo}/{path}, so "everything this input produced" is one GRAPH clause and a fact two inputs both assert is attributable to each of them. A model built from one input keeps the bare graph/semantic; partitioning one input would name a graph after the only file there is. The rule follows from the inputs and there is no flag for it. Only backstage2linkedarchi partitions today.

A graph is named after the file, not after the file at a commit. The prov:Entity for an input keeps its sha, because it identifies bytes at a revision and accumulating one node per revision is what gives an aggregate its history. A graph holds the current facts from a file and a later run replaces it, so keying it on the commit would make a re-pull write a second graph beside the first — and a union would then return that file's facts once per revision it had ever been read at. This is decision 2's rule about runs, applied to the source dimension.

The two IRIs are still relatable, the graph's segments being the entity's minus the revision, but nothing has to parse either: the link is stated as ?graph prov:wasDerivedFrom ?entity.

The curated model has its own graph, and its input is the diagram index. graph/model holds the arch:Model resource, its metamodel conformance, its folders and their ordering, and each concept's folder membership. Separating it is what lets graph/semantic be attributed to a source at all: while they shared a graph, 27 of 60 triples in one measured model were curation and attributable to nothing. Where a run was given no index, the graph is described as generated by the run and derived from nothing, which is the truth — the model's identity then comes from --model-id and its folders from the converter's own layout.

Membership is a statement. Every arch:ModelConcept carries arch:inModel to its model, in the same graph as the concept. That is what makes the partition reassemblable without depending on graph boundaries, and therefore what makes it survive Turtle, RDF/XML, N-Triples and a flattened merged graph.

The per-concept derivation is now recoverable rather than load-bearing. It is still emitted where it already was, since it is the form that survives flattening, but no converter needs to grow one:

CONSTRUCT { ?c prov:wasDerivedFrom ?src }
WHERE {
  ?g a prov:Bundle ; prov:wasDerivedFrom ?src .
  GRAPH ?g { ?c a arch:ModelConcept }
}

Why the emitters stay separate

The obvious cleanup — fold BPMN and ArchiMate into BaseLinkedArchiEmitter.emitProvenance — is rejected, because the three are not three copies of one behaviour:

  • BPMN resolves every metadata predicate through typeMapping.metadataPredicateIri(field, default), letting a type mapping redirect title, created, author and the rest. It is a core TypeMapping feature that only BPMN honours. It also merges dc:/dcterms: fields read out of the BPMN file with the index metadata. Folding it into the base emitter, which hardcodes its predicates, would silently drop both.
  • ArchiMate is an object built on LinkedHashModel and add(model, s, p, o, ctx), not ModelBuilder. Sharing the base emitter's code means porting it first.

So the shared unit is the PROV emission only, and it returns Statements rather than writing into a builder — which is what lets both ModelBuilder and LinkedHashModel consume it. The dct metadata paths stay where they are, and the dct:created collision is fixed in the two places it occurs.

This is recorded so the duplication reads as a decision rather than neglect. The thing that must not happen is a well-meaning consolidation that drops BPMN's predicate override.

Consequences

A repository that wants exact provenance must stop using :latest. Provenance is only as precise as the reference the caller pinned. image: …@sha256:… yields a digest; a moving tag yields the tag, which identifies a build far less firmly. The converter records what it was told and does not embellish. The converter's own source revision is recorded either way, so even an unpinned run identifies the code that produced it.

The digest cannot be baked into the image. It does not exist until after docker push. Build-time baking can therefore supply the tag and the source revision but never the digest, which is why the caller's CI_JOB_IMAGE outranks it. The Dockerfile also needs org.opencontainers.image.version and .revision, which it currently lacks.

Every download source has to publish the revision, or a repackager cannot pass it. Baking SOURCE_REVISION at image build time only works for a pipeline that has the converter commit — which the pipeline building from this repository does, and a pipeline downloading published JARs does not. Its own CI_COMMIT_SHA belongs to its repository, and a version is not a substitute: default-branch builds share one -SNAPSHOT version, so it does not distinguish two builds, and the revision is consumed as a bare sha to compose a commit URL from. Left unpublished, the only honest value a repackager can pass is none, which silently costs every graph it produces its source-commit claim. So the publishing jobs write SOURCE_REVISION and SOURCE_URL — one value per file, no parser needed — beside the artifacts in both the Pages downloads directory and each Package Registry version path. Pages is overwritten in place, so a fetch straddling a pipeline can still pair JARs with the next build's sha; a tagged registry version is immutable and is the source to use when the claim must be exact.

ArchiMate output changes most. It gains the dct provenance it has never emitted and the full PROV description, making its provenance graph comparable to the others for the first time.

prov: must be registered by all emitters. registerNamespaces in core and BPMN's own prefix block both need it. ArchiMate already has it.

Provenance is still unvalidated. The validate command's shapes do not cover the provenance graph, so nothing would catch a regression in it. Provenance shapes are worth adding and are not part of this decision.

Reproducible output needs one more change than this. --run-timestamp commit removes the timestamp churn, which is the part provenance is responsible for. Blank-node identifiers in folder membership remain, so a pipeline that wants git diff on generated graphs to be meaningful needs those addressed too — see decision 7.

BaseConvertCommand is not the place these options went. Nothing extends it: all five ConvertCommands implement Runnable directly, which is why --exclude-states is declared five times. The provenance options are a picocli @Mixin instead, which shares them without asking five commands to change their supertype. The dead base class is a pre-existing problem this decision deliberately does not take on.