ADR 0008 — What the provenance graph says about how a graph was made¶
Status: Accepted, implemented
Implementation summary:
- The conversion event is described with PROV-O: a
prov:Activityfor the run, aprov:SoftwareAgentfor the container image that performed it, and aprov:Entityfor the source file at the commit it was read from. - The generated view (or model) carries
prov:wasGeneratedBythe run andprov:wasDerivedFromthe commit-scoped source entity. - Dublin Core keeps describing the diagram: title, author, dates, description. PROV describes the
event.
dct:createdtherefore stops meaning two things at once. - A node identified by its content — a source file at a commit, a commit, a run, an agent — lives under
{base}provenance/{kind}/…, one node across every model and every run. Statements about it go into each model's own provenance graph. - A source file is identified by four facts: repository, path within it, commit and content digest.
- Each named graph a conversion writes is described as a
prov:Bundle: what generated it, from what, and when. Where a model is built from several inputs its semantic graph is split one per input, so "which input asserted this fact" is answerable. The curated model has its own graph, derived from the diagram index. Every concept carriesarch:inModel, so the pieces reassemble from triples alone. See decision 8. - IRIs are opaque. What identifies a thing travels in triples, not in the IRI.
- The activity and its agent are emitted by every conversion. They replace what
dct:createdanddct:creatorused to say about the run, so making them optional would drop that rather than relocate it. - The git description inside them is what is opted into:
--git-provenance autoreads GitLab's CI variables, overridable by explicit flags,offby default. - The three provenance emitters are not consolidated. See Why the emitters stay separate.
Scope: what the provenance named graph says about how a graph came to be — the run, the software that performed it, the sources it read, and the graphs it wrote. Descriptive metadata about the diagram is out of scope and unchanged.
Depends on: ADR 0001, which assigns IRIs to the resources this ADR describes.
Related: ADR 0004 set the precedent this ADR follows for vocabulary:
reuse a published one (there ADMS, here PROV-O) rather than mint a term in the arch: namespace.
Context¶
Every converter writes a provenance named graph at {base}{notation}/{modelId}/graph/provenance. What
lands in it depends entirely on which converter ran.
dct:created |
dct:source |
dct:creator |
PROV | |
|---|---|---|---|---|
PlantUML, Structurizr, Backstage (BaseLinkedArchiEmitter) |
wall clock | filename literal | "{notation}2linkedarchi" |
none |
| BPMN (own emitter) | wall clock | filename literal | "bpmn2linkedarchi" |
none |
| ArchiMate (own emitter) | — | — | — | prov:generatedAtTime |
Three things are wrong with this, and they are not stylistic.
dct:created means two different things on the same subject¶
BaseLinkedArchiEmitter.emitProvenance writes the conversion time:
val now = OffsetDateTime.now().format(DateTimeFormatter.ISO_OFFSET_DATE_TIME)
mb.namedGraph(provenanceGraph).subject(fileSubject)
.add(vocab.dctCreated, vf.createLiteral(now))
and then, if the index declared created:, writes the author's date under the same predicate on the
same subject in the same graph:
RDF is a set of statements, so this is not an overwrite. Both survive. A consumer asking when the
diagram was created gets two answers and no way to tell which is the authoring date and which is an
artefact of the build that happened to run last night. The BPMN emitter reproduces this, with the
wrinkle that its second date goes through typeMapping.metadataPredicateIri, so the two only collide
when the type mapping does not override created.
Dublin Core is not at fault. dct:created is a perfectly good predicate for when the diagram was
created. It was asked to also carry when this file was generated, which is a fact about an event,
not about a diagram.
The agent is a string, and the thing that ran is not that string¶
dct:creator "bpmn2linkedarchi" names a wrapper script. There is no library, package or release by
that name. The released artefact is a container image holding all five converters, tagged
$CI_REGISTRY_IMAGE:$CI_COMMIT_SHORT_SHA and, on tags, :$CI_COMMIT_TAG.
So the string cannot answer the only question worth asking of it — exactly which build produced this? — and there is nowhere to put the answer, because a literal cannot carry a version, a digest and a source revision at once.
dct:source is a bare filename¶
dct:source "process.bpmn" does not say which repository, which revision, or which of several
same-named files across models. Regenerating the graph, or auditing what a published IRI was derived
from, needs all three.
Why PROV-O rather than more Dublin Core¶
Dublin Core has no vocabulary for an activity that used an input and produced an output, performed by
a piece of software, at a time. PROV-O is the W3C recommendation for precisely that, RDF4J already
ships the PROV vocabulary constant, and the ArchiMate converter already publishes prov:generatedAtTime
— so the prefix is in released output and this is an extension of existing practice rather than a new
dependency.
For the git half, the mapping follows the established ones rather than inventing a third: Git2PROV (De Nies et al., ISWC 2013) and GitLab2PROV. Both treat commits as activities, file versions as entities, and users as agents. Content was rephrased for compliance with licensing restrictions.
Decision¶
1. PROV-O describes the event; Dublin Core describes the diagram¶
dct:created reverts to one meaning: the date the author declared. The conversion time moves onto the
activity as prov:startedAtTime / prov:endedAtTime.
The model keeps prov:generatedAtTime, which ArchiMate already publishes and which the other
converters gain. It duplicates the activity's end time deliberately: it is the cheapest possible
answer to "when was this made" for a consumer who should not have to traverse to the activity to get
it.
2. A node identified by its content lives in one base-scoped namespace¶
Statements go into the model's own provenance graph. The nodes those statements are about do not belong to it: a node identified by its content — a file at a commit, a commit, a run — is one node across every model and every run, under a namespace derived from the base IRI.
| Node | Keyed by | IRI |
|---|---|---|
| source file | repository, commit, path | {base}provenance/source/{repo}/{sha}/{path} |
| commit | the sha | {base}provenance/commit/{sha} |
| author | their canonical email address, percent-encoded; a digest of the name where no address is known | {base}provenance/person/{encoded-email} |
| run | job URL, else start time and image | {base}provenance/run/{token} |
| agent | a digest of the image reference | {base}provenance/agent/{digest} |
| derivation | source and run | {base}provenance/derivation/{repo}/{sha}/{path}/{run} |
Why the base and not the graph. A provenance graph IRI is per model, so naming nodes as fragments of
it makes a file read by two models into two nodes with nothing relating them, and makes one run into one
activity per model. These graphs are published for an aggregating consumer, and the questions that
consumer asks — everything derived from this file, which run produced this — are asked of a single
subject or not at all. A prov:Bundle would not help: bundles separate provenance about provenance,
not one run's description of a model from another's.
Why one reserved segment. {base} holds notation slugs — {base}bpmn/…, {base}backstage/… — so
sources at {base}source/ would reserve a word among them, and this class of node needs six. Everything
sits under {base}provenance/{kind}/…, which reserves one; a notation may not be called provenance.
Spelled out rather than prov/, so it is not read as the prov: prefix.
Each model's provenance graph is self-contained. Every statement a model's description needs is
in that model's graph, so a TriG consumer reading one graph sees a whole description, and a flat merge
coalesces the repeated statements. Where one output file holds several models, the repetition across them
is accepted. dct:identifier on the run carries the CI job URL.
Shas in an IRI are twelve characters. The sha is what separates two revisions of one file, and these
nodes are shared across an aggregate, so the population it discriminates within is every commit that has
touched that path. The full value is always dct:identifier, on both the source entity and the commit
activity.
The IRIs are legible but not load-bearing, in the sense §3 means: every component is also asserted as a triple, so nothing has to parse or resolve one.
3. IRIs are opaque; identity travels in triples¶
{base}provenance/source/… is not required to parse or resolve. Everything its components carry is also
a statement, so a consumer reads triples and never string-matches an IRI.
A source file is identified by four facts, each on its own predicate:
| Fact | Predicate | Where it comes from |
|---|---|---|
| which repository | dct:isPartOf, target typed prov:Collection |
the source map's repo:, or CI_PROJECT_URL |
| where in it | schema:name |
the source map's path:, or the repo-relative path |
| which revision | dct:identifier |
the commit sha |
| which bytes | schema:sha256 |
the source map's contentSha256: |
The repository is not optional detail. schema:name is the path inside a repository, and for a
pulled input that repository is not the one doing the converting: every Backstage descriptor in every
repository is called catalog-info.yaml, so the path alone discriminates nothing across a catalog
collected from many sources. Repository and path together name the file; the commit picks the revision.
dct:isPartOf because that is the relation — a file is part of a repository, the same statement the term
already makes about a view and its folder. The repository URL is the node IRI directly rather than a
minted proxy: unlike a file at a commit, a repository has a stable published IRI that other graphs cite
too, so a local proxy would fragment it for nothing. It is typed prov:Collection and not also
prov:Entity, since PROV-O makes Collection a subclass and ?s a prov:Entity is how a consumer
enumerates source files.
schema:sha256 and not a second dct:identifier, which carries the commit — two identifiers on one
subject would reintroduce on this node the ambiguity §1 removes from the diagram. A digest and a commit
answer different questions: which bytes, and which revision. Two commits sharing a digest are a file that
moved commit without changing, which an incremental re-pull produces and a commit sha alone cannot show.
The source IRI is keyed on all three of repository, commit and path. Keying it on fewer lets two files
share a node, and a node shared by files from two repositories would carry two dct:isPartOf values —
saying a file is part of two repositories, which is false rather than merely imprecise.
4. The agent is the container image¶
<…/provenance/agent/7d1c40ba>
a prov:SoftwareAgent, schema:SoftwareApplication ;
schema:name "bpmn2linkedarchi" ;
schema:softwareVersion "1.4.0" ;
dct:identifier "registry.gitlab.com/linked-archi/tools/converters/linked-archi-converters@sha256:89ab…" ;
schema:url <https://gitlab.com/linked-archi/tools/converters/linked-archi-converters/-/commit/b28afe4…> .
schema:name is the entrypoint, because one image contains five converters and which one ran is part
of the answer. dct:identifier is the image reference, digest-pinned when the caller pinned one.
schema:url carries the converter source revision, as a URL rather than a bare sha so no term has to
be minted for it — and not as a second dct:identifier, since two identifiers on one subject would
repeat the ambiguity this ADR exists to remove.
schema:url and not rdfs:seeAlso, which is what this carried until 1.3.0. seeAlso says only "here is
related information" and leaves the relationship to be guessed; schema:url says the object is the URL
of the subject, which is the statement intended. The same objection retired seeAlso from the source
entity in §5 — there prov:alternateOf already carried the link, so the term was dropped rather than
replaced.
Resolution order, so the claim degrades honestly instead of guessing:
--image-refLA_IMAGE_REF, elseCI_JOB_IMAGE- a value baked in at image build time
- nothing — emit
schema:softwareVersionand make no image claim
5. prov:wasDerivedFrom points at the source file at its commit¶
<…/bpmn/order-domain/view/current>
prov:wasGeneratedBy <…/provenance/run/9c2f1e04> ;
prov:wasDerivedFrom <…/provenance/source/linked-archi-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .
<…/provenance/run/9c2f1e04>
a prov:Activity ;
prov:startedAtTime "2026-08-01T09:14:22Z"^^xsd:dateTime ;
prov:endedAtTime "2026-08-01T09:14:25Z"^^xsd:dateTime ;
prov:wasAssociatedWith <…/provenance/agent/7d1c40ba> ;
prov:used <…/provenance/source/linked-archi-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> ;
dct:identifier "https://gitlab.com/linked-archi/models/-/jobs/9182734" .
<…/provenance/source/linked-archi-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn>
a prov:Entity ;
schema:name "models/bpmn/order/process.bpmn" ;
dct:isPartOf <https://gitlab.com/linked-archi/models> ;
dct:identifier "8c44d1a4e9…" ;
prov:wasGeneratedBy <…/provenance/commit/8c44d1a4e9f0> ;
prov:alternateOf <https://gitlab.com/linked-archi/models/-/blob/8c44d1a4e9…/models/bpmn/order/process.bpmn> .
<https://gitlab.com/linked-archi/models>
a prov:Collection ;
schema:name "linked-archi/models" .
<…/provenance/commit/8c44d1a4e9f0>
a prov:Activity ;
prov:endedAtTime "2026-07-30T12:01:02Z"^^xsd:dateTime ;
prov:wasAssociatedWith <…/provenance/person/a.author%40example.org> ;
dct:identifier "8c44d1a4e9…" ;
schema:url <https://gitlab.com/linked-archi/models/-/commit/8c44d1a4e9…> .
<…/provenance/person/a.author%40example.org>
a prov:Person ;
schema:name "A. Author" ;
schema:email "a.author@example.org" .
The chain diagram → generated by this run → of this file at this commit → committed by this person is then one traversal.
A derivation also says which run it went through¶
A derived resource carries both the plain and the qualified form, because the plain form cannot name the
activity a derivation passed through and PROV does not license inferring it: prov:used and
prov:wasGeneratedBy follow from a derivation, not the reverse. Naming it matters because a provenance
graph holds many activities — one conversion run plus one commit activity per distinct commit — and an
aggregate holds the runs of every conversion that contributed to it.
<…/element/component/default/order-service>
prov:wasDerivedFrom <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:qualifiedDerivation <…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04> .
<…/provenance/derivation/group-order-service/4f2c1ab8e0d1/catalog-info-yaml/9c2f1e04>
a prov:Derivation ;
prov:entity <…/provenance/source/group-order-service/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:hadActivity <…/provenance/run/9c2f1e04> .
One derivation node per (source, run), shared by every resource derived from that file. A
prov:Derivation states which entity was used and which activity did it, and names no derived entity, so
one node per file per run is what the record says rather than a compression of it. The alternative is a
node per resource, which costs a 500-entity catalog 500 nodes to state one fact 500 times.
An IRI, not a blank node. A blank node is relabelled on every parse, so two runs' descriptions of one derivation could not be recognised as one — and being able to recognise that is why the qualified form is here at all.
prov:hadUsage and prov:hadGeneration are not emitted. Supplying both makes this a precise-1
derivation under PROV-CONSTRAINTS, with ordering obligations between that usage and that generation.
prov:hadGeneration is per-derived-entity by construction, so it is incompatible with the shared node,
and its content would restate a prov:wasGeneratedBy the resource already carries; prov:hadUsage is
shareable but states what {run} prov:used {source} states.
The property is prov:entity, lowercase — prov:Entity is the class. RDF4J exposes them as
PROV.ENTITY_PROP and PROV.ENTITY, and a class IRI in the predicate position is valid RDF that means
nothing, so the emitter's test asserts the predicate's IRI string rather than the constant.
Three predicates, three different jobs¶
dct:source, prov:alternateOf and prov:wasDerivedFrom all relate an output to its input somewhere in
this block. They are not interchangeable and none of them is on the source entity by accident.
| Predicate | Subject | Object | Says |
|---|---|---|---|
prov:wasDerivedFrom |
the output — model, view, each named graph, each concept | the source entity | this was produced from that file |
dct:source |
the output — the view, or the model where there is no view | the blob IRI, or the input's filename where no URL can be built | the same thing in Dublin Core, and the only form of it a run with provenance off still emits |
prov:alternateOf |
the source entity | the blob IRI | these two IRIs denote one document |
The source entity carries no dct:source. It began as one, holding a bare filename, and a bare
dct:source "catalog-info.yaml" is the same string for every descriptor in a pulled catalog: a consumer
holding it cannot get back to the file. Repointing it at the blob URL fixed the value and broke the
meaning — this node is the file at that commit, which is what prov:alternateOf says of it, what
schema:sha256 digests and what the commit activity generated. Saying it was derived from that same
blob would make the file derived from itself. The derivation holds of the output, and
prov:wasDerivedFrom states it there, on the subject it is true of.
dct:source stays on the output, where "a related resource from which the described resource is
derived" is exactly the relation. Its value is the commit-pinned blob IRI where the run can build one and
the input's bare filename where it cannot — which is every run with provenance off, and the reason the
predicate is kept rather than left to prov:wasDerivedFrom. Consumers must therefore accept either node
kind; FILTER(isIRI(?o)) separates them. DCMI sanctions both: range rdfs:Resource, with a URI or "a
string conforming to a formal identification system" recommended.
schema:name carries the identifying path. Recovering it from the blob URL would mean parsing a forge
URL, which §3 says not to do. Same choice as the repository node above, which carries its
group/subgroup/name path under schema:name. Not dct:title: a title is chosen for display, a path is
not a title. The path is repo-relative, which requires the repository root: CI_PROJECT_DIR under CI,
git rev-parse --show-toplevel locally. Where neither is available the path is emitted as given, which
for a file with no repository is its bare name.
prov:alternateOf states the identity, and no rdfs:seeAlso sits beside it. Not owl:sameAs, which
would licence inferences about the document this block does not intend, and not prov:atLocation, which
says where to look rather than what a thing is. rdfs:seeAlso used to carry the same URL beside it and
was dropped: it asserts no relationship between the two IRIs, so as a second link to the same blob it said
nothing. Where a URL is the only pointer to a resource — the forge's page for a commit, the page for the
converter revision — it is carried as schema:url, which states what seeAlso left implicit.
Enumerating source files is ?s a prov:Entity ; schema:name ?path. The blob is typed prov:Entity
too, since prov:alternateOf has prov:Entity at both ends, and it carries nothing else — so the type
alone over-matches and the name is what discriminates. The repository and the commit author are named but
are not entities; a graph is typed prov:Bundle only. This replaces the older advice to filter on
dct:source, which no longer appears on the entity.
A person is identified by their address, not by their name¶
<…/provenance/person/a.author%40example.org>
a prov:Person ;
schema:name "A. Author" ; # what a human reads
schema:email "a.author@example.org" . # what identifies them
The node is keyed on the canonical email address: trimmed, lower-cased and percent-encoded as one IRI
path segment. A display name is not an identity. It changes when someone marries, standardises on a
different form, or configures one machine as J. Doe and another as Jane Doe — and keyed on the name,
each variant is a different person, so no query can gather one author's commits. Git offers no other
stable handle: a commit records a name and an address, and no username or numeric id. Canonicalising makes
Jane@Example.org and jane@example.org one person. Percent encoding is reversible, so punctuation does
not collapse distinct addresses: a+b@example.org and a-b@example.org remain different people. An
exceptionally long encoded address keeps a readable prefix plus the full SHA-256 so the segment remains
filesystem-safe.
The address is published, as schema:email. This reverses the position this ADR held until 1.3.0, that
an address is "personal data in a published graph for no gain in provenance". The gain is stable identity,
which is not available any other way, and the graph is internal — the converters cannot read a private
catalog without being inside the network that holds one. It is also the judgement the rest of the suite had
already made: the LeanIX converter keys arch:Stakeholder on a lower-cased address and publishes
schema:email, and Backstage publishes bs:email from spec.profile. The provenance person was the
outlier.
schema:email and not bs:email or lmm:subscriberEmail, each of which is one source's own term. LeanIX
already chose schema:email so that people join across notations on more than a label; a git author and a
LeanIX stakeholder are now joinable on the same predicate.
The address in the IRI, readable, rather than hashed. Following the same reasoning the LeanIX converter
records: the graph carries the address as schema:email anyway, so hashing the IRI would hide nothing
while making the resource unrecognisable. Worth knowing that an IRI travels further than a literal.
Where no address is known the node stays a digest of the name, exactly as before — a local git log on
a repository with user.email unset, or a source map written before addresses were recorded. An address
with no name is still a person: it identifies them, which is the harder half, and schema:name is simply
omitted.
This does not yet join a git author to a Backstage User or a LeanIX stakeholder as one node. Those are
model-scoped IRIs and reconciling them is a separate decision — see
todo/UPSTREAM-REQUEST-provenance-source-identity.md. schema:email is what makes the join expressible.
A source that is not a file at a commit — a LeanIX or Backstage API snapshot — needs a different shape,
which is not yet decided. See todo/converters-provenance-doc-recommendations.md, item B.
6. Git facts come from GitLab, off by default¶
--git-provenance auto reads CI_PROJECT_URL, CI_COMMIT_SHA, CI_COMMIT_TIMESTAMP,
CI_COMMIT_AUTHOR, CI_JOB_IMAGE and CI_JOB_URL, falling back to git rev-parse HEAD and
git log -1 for local runs. Explicit flags (--source-repo-url, --source-commit,
--source-commit-time, --source-author) override.
GitLab-only is a deliberate limit, not an oversight: it is the only CI in use. The detection lives in
one place in core with the environment injected, so it is testable without a repository or a
pipeline, and a second forge is a second detector rather than a redesign.
Default off, because it changes published output.
7. --run-timestamp=commit|now|none¶
Deriving the activity time from the commit rather than the clock makes the provenance graph
byte-stable for an unchanged model at an unchanged commit, which it has never been. now keeps
today's behaviour; none omits the times for output that must not carry any.
It does not yet make the whole file byte-stable, and it is worth being exact about why. Folder
membership is emitted with blank nodes, whose identifiers are freshly generated per run, so two
conversions of an identical model at an identical commit still differ on those lines. Measured on a
one-class diagram: the provenance graph is identical between runs and contains no blank nodes, while
two elsewhere in the file change every time. Deterministic blank nodes — or schema:ListItem
resources with minted IRIs — are what full reproducibility needs, and that is a separate decision
about folder emission, not about provenance.
8. Each named graph is a prov:Bundle that says what it was lifted from¶
Derivation on a resource answers "where did this thing come from". It cannot answer "where did this
fact come from", because a resource two inputs contribute to carries two prov:wasDerivedFrom values
and nothing pairs a value with an input.
So the graph is a subject too. Every named graph a conversion writes is described in that model's provenance graph:
<…/backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml>
a prov:Bundle ;
prov:wasGeneratedBy <…/provenance/run/9c2f1e04> ;
prov:wasDerivedFrom <…/provenance/source/group-orders/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:generatedAtTime "2026-08-29T09:03:11Z"^^xsd:dateTime .
prov:Bundle and not also prov:Entity. PROV-O makes Bundle a subclass of Entity, so the second
type adds nothing a reasoner needs, and it adds something a non-reasoning consumer does not want:
?s a prov:Entity is how one reaches the inputs of a run, and a graph is not an input. The repository
node is typed prov:Collection alone for the same reason.
The semantic graph is partitioned by input where a model has more than one.
graph/semantic/{repo}/{path}, so "everything this input produced" is one GRAPH clause and a fact two
inputs both assert is attributable to each of them. A model built from one input keeps the bare
graph/semantic; partitioning one input would name a graph after the only file there is. The rule
follows from the inputs and there is no flag for it. Only backstage2linkedarchi partitions today.
A graph is named after the file, not after the file at a commit. The prov:Entity for an input keeps
its sha, because it identifies bytes at a revision and accumulating one node per revision is what gives an
aggregate its history. A graph holds the current facts from a file and a later run replaces it, so keying
it on the commit would make a re-pull write a second graph beside the first — and a union would then
return that file's facts once per revision it had ever been read at. This is decision 2's rule about runs,
applied to the source dimension.
The two IRIs are still relatable, the graph's segments being the entity's minus the revision, but nothing
has to parse either: the link is stated as ?graph prov:wasDerivedFrom ?entity.
The curated model has its own graph, and its input is the diagram index. graph/model holds the
arch:Model resource, its metamodel conformance, its folders and their ordering, and each concept's
folder membership. Separating it is what lets graph/semantic be attributed to a source at all: while
they shared a graph, 27 of 60 triples in one measured model were curation and attributable to nothing.
Where a run was given no index, the graph is described as generated by the run and derived from nothing,
which is the truth — the model's identity then comes from --model-id and its folders from the
converter's own layout.
Membership is a statement. Every arch:ModelConcept carries arch:inModel to its model, in the
same graph as the concept. That is what makes the partition reassemblable without depending on graph
boundaries, and therefore what makes it survive Turtle, RDF/XML, N-Triples and a flattened merged graph.
The per-concept derivation is now recoverable rather than load-bearing. It is still emitted where it already was, since it is the form that survives flattening, but no converter needs to grow one:
CONSTRUCT { ?c prov:wasDerivedFrom ?src }
WHERE {
?g a prov:Bundle ; prov:wasDerivedFrom ?src .
GRAPH ?g { ?c a arch:ModelConcept }
}
Why the emitters stay separate¶
The obvious cleanup — fold BPMN and ArchiMate into BaseLinkedArchiEmitter.emitProvenance — is
rejected, because the three are not three copies of one behaviour:
- BPMN resolves every metadata predicate through
typeMapping.metadataPredicateIri(field, default), letting a type mapping redirecttitle,created,authorand the rest. It is acoreTypeMappingfeature that only BPMN honours. It also merges dc:/dcterms: fields read out of the BPMN file with the index metadata. Folding it into the base emitter, which hardcodes its predicates, would silently drop both. - ArchiMate is an
objectbuilt onLinkedHashModelandadd(model, s, p, o, ctx), notModelBuilder. Sharing the base emitter's code means porting it first.
So the shared unit is the PROV emission only, and it returns Statements rather than writing into a
builder — which is what lets both ModelBuilder and LinkedHashModel consume it. The dct metadata
paths stay where they are, and the dct:created collision is fixed in the two places it occurs.
This is recorded so the duplication reads as a decision rather than neglect. The thing that must not happen is a well-meaning consolidation that drops BPMN's predicate override.
Consequences¶
A repository that wants exact provenance must stop using :latest. Provenance is only as precise
as the reference the caller pinned. image: …@sha256:… yields a digest; a moving tag yields the tag,
which identifies a build far less firmly. The converter records what it was told and does not
embellish. The converter's own source revision is recorded either way, so even an unpinned run
identifies the code that produced it.
The digest cannot be baked into the image. It does not exist until after docker push. Build-time
baking can therefore supply the tag and the source revision but never the digest, which is why the
caller's CI_JOB_IMAGE outranks it. The Dockerfile also needs org.opencontainers.image.version
and .revision, which it currently lacks.
Every download source has to publish the revision, or a repackager cannot pass it. Baking
SOURCE_REVISION at image build time only works for a pipeline that has the converter commit —
which the pipeline building from this repository does, and a pipeline downloading published JARs does
not. Its own CI_COMMIT_SHA belongs to its repository, and a version is not a substitute: default-branch
builds share one -SNAPSHOT version, so it does not distinguish two builds, and the revision is
consumed as a bare sha to compose a commit URL from. Left unpublished, the only honest value a
repackager can pass is none, which silently costs every graph it produces its source-commit claim. So
the publishing jobs write SOURCE_REVISION and SOURCE_URL — one value per file, no parser needed —
beside the artifacts in both the Pages downloads directory and each Package Registry version path.
Pages is overwritten in place, so a fetch straddling a pipeline can still pair JARs with the next
build's sha; a tagged registry version is immutable and is the source to use when the claim must be
exact.
ArchiMate output changes most. It gains the dct provenance it has never emitted and the full PROV
description, making its provenance graph comparable to the others for the first time.
prov: must be registered by all emitters. registerNamespaces in core and BPMN's own prefix
block both need it. ArchiMate already has it.
Provenance is still unvalidated. The validate command's shapes do not cover the provenance graph,
so nothing would catch a regression in it. Provenance shapes are worth adding and are not part of this
decision.
Reproducible output needs one more change than this. --run-timestamp commit removes the
timestamp churn, which is the part provenance is responsible for. Blank-node identifiers in folder
membership remain, so a pipeline that wants git diff on generated graphs to be meaningful needs
those addressed too — see decision 7.
BaseConvertCommand is not the place these options went. Nothing extends it: all five
ConvertCommands implement Runnable directly, which is why --exclude-states is declared five
times. The provenance options are a picocli @Mixin instead, which shares them without asking five
commands to change their supertype. The dead base class is a pre-existing problem this decision
deliberately does not take on.