Output Serialization¶
Writing several artifacts from one conversion¶
--output is repeatable, and each value is path[:FORMAT[:PROFILE]]:
bpmn2linkedarchi convert models/*.bpmn --base-iri https://example.org/la/ \
-o out/model.trig \
-o out/model.ttl \
-o out/model-only.ttl:TURTLE:no-diagrams
Both suffixes are optional. Format resolution runs inline token → --format → file extension, so
-o out/model.trig needs no token at all, and --output x.ttl --format TURTLE keeps working as it
always has. A file extension is accepted where a format name is expected, so --format ttl means
Turtle.
Prefer this over running the converter once per file. The parse and type mapping happen once instead
of N times — about 2.4x faster for three artifacts on a large ArchiMate model — and, more importantly,
every artifact carries the same prov:generatedAtTime. Separate runs each stamp their own, so
three files from one publish would claim three different generation times and an aggregation store
would see three activities where there was one.
View profiles¶
A profile decides how much of the diagrams an artifact carries. The point is Turtle: a TriG consumer
can shed geometry after loading with DROP GRAPH <…/graph/views>, but Turtle has no graph boundaries,
so for those consumers the cut has to happen when the file is written.
| Profile | Keeps | Drops | Answers |
|---|---|---|---|
full (default) |
everything | — | — |
no-geometry |
arch:View, nodes, links, archvis:view/archElement |
bounds, bendpoints, styles | "what is drawn on which diagram" |
no-views |
arch:View with its name and viewpoint |
the whole views graph | "which diagrams exist" |
no-diagrams |
elements, relationships, folders, provenance | views graph and every arch:View |
"the architecture only" |
On a real ArchiMate model (Archisurance, 17 views) the sizes are 8,094 triples for full, 3,855 for
no-geometry and 2,038 for no-diagrams — geometry is about three quarters of the output.
Profiles filter the built model rather than changing what the emitters produce, which is only well
defined because the semantic/views boundary is now the same in every converter. no-diagrams also
removes the folder entries that listed the removed diagrams, so no folder is left asserting a
containment whose other half is gone.
The --format flag¶
--format is the default for any --output that names no format of its own. It is optional in every
converter; before, three of the six required it while the other three inferred from the extension.
All converters support the same values:
| Value | Extension | Named graphs | Description |
|---|---|---|---|
TRIG |
.trig |
✓ preserved | Recommended. W3C TriG — superset of Turtle with named graph blocks. |
TURTLE |
.ttl |
✗ merged | W3C Turtle — all named graphs flattened into a single default graph. |
JSONLD |
.jsonld |
✓ preserved | JSON-LD — JSON-based RDF. Named graphs via @graph. |
RDFXML |
.rdf |
✗ merged | RDF/XML — legacy XML serialization. No named graph support. |
NTRIPLES |
.nt |
✗ merged | N-Triples — one triple per line, no prefixes. |
NQUADS |
.nq |
✓ preserved | N-Quads — N-Triples plus a graph term. One quad per line, no prefixes: the streaming and line-processing format that still says which graph a fact is in. |
Why TriG is the default¶
A conversion produces four kinds of named graph:
{base}{notation}/{modelId}/graph/semantic ← facts lifted from an input, including the views themselves
{base}{notation}/{modelId}/graph/semantic/{slug} ← the same, one graph per input, where a model has several
{base}{notation}/{modelId}/graph/model ← the curated model: arch:Model, folders, ordering
{base}{notation}/{modelId}/graph/views ← diagram geometry only
{base}{notation}/{modelId}/graph/provenance ← PROV-O about all of the above
With TriG these are preserved as separate contexts. With Turtle they get merged and the graph boundaries are lost.
The semantic/views line falls around geometry rather than around diagrams: the arch:View
resource is a model concept and sits in semantic, while views holds only archvis: nodes, links,
points and their bounds. That is what makes views safe to drop on its own — for an ArchiMate model
it is around 80% of the output — without losing the diagrams or their names.
semantic and model¶
semantic holds what was lifted from an input. model holds what the conversion curated: the
arch:Model resource, its metamodel conformance, its folders and their ordering, and each concept's
folder membership. None of that comes from a diagram or a descriptor — its input is the diagram index —
and separating it is what lets semantic be attributed to a source at all. Measured on a
two-descriptor catalog, 27 of 60 semantic triples were curation.
One semantic graph per input¶
Where a model is built from several inputs, each one's facts go in graph/semantic/{slug}, named after
that input: its repository path where the run knows one, else its file name.
{base}backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-payments/catalog-info-yaml
That makes "everything this input produced" a single GRAPH clause, and it makes a fact two inputs
both assert attributable to each of them — a resource carrying two prov:wasDerivedFrom values cannot
say which value came from which file.
A model built from one input keeps the bare graph/semantic: partitioning one input would name a graph
after the only file there is. The rule follows from the inputs and there is no flag for it.
The slug carries no commit sha, deliberately. A graph holds the current facts from a file and a later
run replaces it, so it is named after the file; the prov:Entity describing that file keeps its sha,
because it identifies bytes at a revision. Keyed on the commit too, a re-pull that moved one file's
commit would write a second graph beside the first, and a union would return that file's facts once per
revision it had ever been read at.
Only backstage2linkedarchi partitions today. The other five write one graph/semantic per model.
Every graph says where it came from¶
Each graph is described in graph/provenance as a prov:Bundle:
<…/graph/semantic/group-orders/catalog-info-yaml> a prov:Bundle ;
prov:wasGeneratedBy <{base}provenance/run/8d5a6e38> ;
prov:wasDerivedFrom <{base}provenance/source/group-orders/4f2c1ab8e0d1/catalog-info-yaml> ;
prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime .
prov:Bundle and not also prov:Entity: PROV-O makes Bundle a subclass of Entity, so the second type
adds nothing a reasoner needs, and ?s a prov:Entity is how a consumer reaches the inputs of a run.
A graph is not an input.
Because the derivation is carried by the graph, the per-concept prov:wasDerivedFrom is recoverable
rather than load-bearing:
CONSTRUCT { ?c prov:wasDerivedFrom ?src }
WHERE {
?g a prov:Bundle ; prov:wasDerivedFrom ?src .
GRAPH ?g { ?c a arch:ModelConcept }
}
TriG output structure¶
@prefix arch: <https://meta.linked.archi/core#> .
@prefix bpmn: <https://meta.linked.archi/bpmn/onto#> .
@prefix archvis: <https://meta.linked.archi/core-vis#> .
@prefix dct: <http://purl.org/dc/terms/> .
# ─── Semantic graph: facts lifted from the input ───────────────────
<https://example.org/la/bpmn/demo/graph/semantic> {
<.../element/Task_ValidateOrder> a bpmn:UserTask, arch:Element, arch:ModelConcept ;
arch:inModel <.../bpmn/demo> ;
bpmn:name "Validate Order" ;
bpmn:id "Task_ValidateOrder" .
<.../relationship/Flow_1> a bpmn:SequenceFlow, arch:QualifiedRelationship, arch:ModelConcept ;
arch:inModel <.../bpmn/demo> ;
arch:source <.../element/Start_1> ;
arch:target <.../element/Task_ValidateOrder> .
}
# ─── Model graph: the curated model ────────────────────────────────
<https://example.org/la/bpmn/demo/graph/model> {
<.../bpmn/demo> a arch:Model ;
arch:modelConformsToMetamodel <https://meta.linked.archi/bpmn/metamodel#Bpmn202> .
<.../bpmn/demo/folder> a arch:Folder ;
schema:name "demo" ;
schema:itemListElement [ a schema:ListItem ;
schema:position 1 ;
schema:item <.../bpmn/demo/folder/Elements> ] .
<.../element/Task_ValidateOrder> dct:isPartOf <.../bpmn/demo/folder/Elements> .
}
# ─── Views graph: diagram geometry ─────────────────────────────────
<https://example.org/la/bpmn/demo/graph/views> {
<.../view/.../node/Shape_1> a archvis:ArchNode ;
archvis:view <.../view/Diagram_1> ;
archvis:archElement <.../element/Task_ValidateOrder> ;
archvis:bounds-x "270"^^xsd:double ;
archvis:bounds-y "190"^^xsd:double ;
archvis:bounds-w "100"^^xsd:double ;
archvis:bounds-h "56"^^xsd:double .
}
# ─── Provenance graph: how the graph came to be ────────────────────
<https://example.org/la/bpmn/demo/graph/provenance> {
<.../bpmn/demo>
# the input this model came from: the blob IRI where the run can name one, the filename otherwise
dct:source <https://git.example.org/group/models/-/blob/8c44d1a4e9f0/models/bpmn/order/process.bpmn> ;
prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
prov:wasGeneratedBy <https://example.org/la/provenance/run/8d5a6e38> ;
prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .
<https://example.org/la/provenance/run/8d5a6e38> a prov:Activity ;
prov:endedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
prov:wasAssociatedWith <https://example.org/la/provenance/agent/b01e544f> ;
prov:used <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> ;
dct:identifier "https://git.example.org/group/models/-/jobs/9182734" .
<https://example.org/la/provenance/agent/b01e544f>
a prov:SoftwareAgent, schema:SoftwareApplication ;
schema:name "bpmn2linkedarchi" ;
schema:softwareVersion "1.3.0" .
<https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn>
a prov:Entity ;
schema:name "models/bpmn/order/process.bpmn" ;
dct:isPartOf <https://git.example.org/group/models> ;
dct:identifier "8c44d1a4e9f0112233445566778899aabbccddee" ;
prov:wasGeneratedBy <https://example.org/la/provenance/commit/8c44d1a4e9f0> .
# each graph this run wrote, and what it came from
<.../bpmn/demo/graph/semantic> a prov:Bundle ;
prov:wasGeneratedBy <https://example.org/la/provenance/run/8d5a6e38> ;
prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .
<.../bpmn/demo/graph/views> a prov:Bundle ; … .
# the curated graph's input is the index, not the diagram
<.../bpmn/demo/graph/model> a prov:Bundle ;
prov:wasGeneratedBy <https://example.org/la/provenance/run/8d5a6e38> ;
prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-diagram-index-yaml> .
}
The provenance nodes are not under the model's IRI. {base}provenance/{kind}/… is a namespace
shared by every model and every run, so one file at one commit is one node however many models read it,
and one run is one activity however many models it converted. The statements about those nodes go into
each model's own provenance graph, which is why this graph is still a complete description on its own.
See ADR 0008 §2.
This matters for the format choice below. The derivation is asserted on the model resource as well as on each graph, so a Turtle consumer keeps it. The graph-level statements survive flattening too, as statements: the graph names stop naming anything, so "which graph is this triple in" is gone, but "what did this run write, and from what" is still answerable.
Turtle output (same data, flat)¶
@prefix arch: <https://meta.linked.archi/core#> .
@prefix bpmn: <https://meta.linked.archi/bpmn/onto#> .
<.../element/Task_ValidateOrder> a bpmn:UserTask, arch:Element, arch:ModelConcept ;
arch:inModel <.../bpmn/demo> ;
bpmn:name "Validate Order" .
<.../relationship/Flow_1> a bpmn:SequenceFlow, arch:QualifiedRelationship, arch:ModelConcept ;
arch:inModel <.../bpmn/demo> ;
arch:source <.../element/Start_1> ;
arch:target <.../element/Task_ValidateOrder> .
<.../view/.../node/Shape_1> a archvis:ArchNode ;
archvis:view <.../view/Diagram_1> ;
archvis:archElement <.../element/Task_ValidateOrder> .
<.../bpmn/demo> prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .
Graph boundaries lost, provenance and membership kept
In Turtle output there is no way to tell a semantic triple from a provenance triple after the fact, and "which input asserted this fact" is gone with the graph names. Use TriG when either matters.
What survives is everything asserted on a resource rather than carried by a
graph boundary: arch:inModel, so the model is still reassemblable; the
model's own prov:wasGeneratedBy and prov:wasDerivedFrom; and the bundle
descriptions, so the inventory of what the run wrote and from what is intact
even though the graph names no longer name graphs.
dct:created is not among them: it means the date the author declared, and
the conversion time is prov:generatedAtTime on the model and
prov:endedAtTime on the run
(ADR 0008 §1).
When to use which¶
| Scenario | Format |
|---|---|
| Loading into a triplestore (Jena/Fuseki, GraphDB, Oxigraph) | TRIG |
| SPARQL queries that filter by graph | TRIG |
| Merging output from multiple converters into one dataset | TRIG |
| Simple inspection / diffing / grepping | TURTLE |
| Tools that only accept Turtle (some validators) | TURTLE |
| Human reading in a text editor | TURTLE |
Streaming, grep/awk, or bulk loading, with graphs intact |
NQUADS |
| JSON-based workflows / APIs | JSONLD |
SPARQL examples (TriG-loaded data)¶
-- All elements in the semantic graph only
SELECT ?elem ?type ?label
FROM NAMED <https://example.org/la/bpmn/demo/graph/semantic>
WHERE {
GRAPH <https://example.org/la/bpmn/demo/graph/semantic> {
?elem a arch:Element, ?type .
OPTIONAL { ?elem bpmn:name ?label }
}
}
-- Provenance: what produced this model, from which file, and when?
--
-- Only four patterns are guaranteed: the two derivations, the source's `schema:name`, and the run's agent.
-- Everything else is OPTIONAL by design — `--run-timestamp none` omits the times, an unpinned run
-- makes no image claim, and a conversion outside a known forge names no repository. Binding any of
-- them non-optionally returns zero rows instead of a partial answer.
SELECT ?ended ?repo ?path ?commit ?converter ?image
WHERE {
GRAPH <https://example.org/la/bpmn/demo/graph/provenance> {
?model prov:wasGeneratedBy ?run ;
prov:wasDerivedFrom ?src .
?src schema:name ?path .
?run prov:wasAssociatedWith ?agent .
?agent schema:name ?converter .
OPTIONAL { ?run prov:endedAtTime ?ended }
OPTIONAL { ?agent dct:identifier ?image }
OPTIONAL { ?src dct:isPartOf ?repo }
OPTIONAL { ?src dct:identifier ?commit }
}
}
-- All view nodes with geometry
SELECT ?node ?x ?y ?w ?h
WHERE {
GRAPH <https://example.org/la/bpmn/demo/graph/views> {
?node a archvis:ArchNode ;
archvis:bounds-x ?x ; archvis:bounds-y ?y ;
archvis:bounds-w ?w ; archvis:bounds-h ?h .
}
}
-- Everything one input produced, named by its path rather than by its graph IRI
SELECT ?s ?p ?o
WHERE {
?g prov:wasDerivedFrom ?src .
?src schema:name "catalogs/orders/catalog-info.yaml" .
GRAPH ?g { ?s ?p ?o }
}
-- Which input asserted a given fact. Two answers where two inputs both assert it.
SELECT ?path
WHERE {
GRAPH ?g { <…/element/component/default/order-service> bs:lifecycleState bs:Production }
?g prov:wasDerivedFrom ?src .
?src schema:name ?path .
}
-- The model, reassembled across its partitions. Works in Turtle too.
SELECT ?concept
WHERE { ?concept arch:inModel <https://example.org/la/backstage/service-catalog> }
The first two need the store's default graph to be a union of the named graphs, or the ?g patterns
outside the GRAPH clause will not match. That is store configuration, not something the output can
carry — see the aggregation graph.
Usage examples¶
# Explicit format (recommended)
java -jar bpmn2linkedarchi.jar convert process.bpmn \
--format TRIG --base-iri https://example.org/la/ -o out.trig
# Turtle for human inspection
java -jar bpmn2linkedarchi.jar convert process.bpmn \
--format TURTLE --base-iri https://example.org/la/ -o out.ttl
# JSON-LD for API consumption
java -jar bpmn2linkedarchi.jar convert process.bpmn \
--format JSONLD --base-iri https://example.org/la/ -o out.jsonld