Skip to content

Output Serialization

Writing several artifacts from one conversion

--output is repeatable, and each value is path[:FORMAT[:PROFILE]]:

bpmn2linkedarchi convert models/*.bpmn --base-iri https://example.org/la/ \
  -o out/model.trig \
  -o out/model.ttl \
  -o out/model-only.ttl:TURTLE:no-diagrams

Both suffixes are optional. Format resolution runs inline token → --format → file extension, so -o out/model.trig needs no token at all, and --output x.ttl --format TURTLE keeps working as it always has. A file extension is accepted where a format name is expected, so --format ttl means Turtle.

Prefer this over running the converter once per file. The parse and type mapping happen once instead of N times — about 2.4x faster for three artifacts on a large ArchiMate model — and, more importantly, every artifact carries the same prov:generatedAtTime. Separate runs each stamp their own, so three files from one publish would claim three different generation times and an aggregation store would see three activities where there was one.

View profiles

A profile decides how much of the diagrams an artifact carries. The point is Turtle: a TriG consumer can shed geometry after loading with DROP GRAPH <…/graph/views>, but Turtle has no graph boundaries, so for those consumers the cut has to happen when the file is written.

Profile Keeps Drops Answers
full (default) everything — —
no-geometry arch:View, nodes, links, archvis:view/archElement bounds, bendpoints, styles "what is drawn on which diagram"
no-views arch:View with its name and viewpoint the whole views graph "which diagrams exist"
no-diagrams elements, relationships, folders, provenance views graph and every arch:View "the architecture only"

On a real ArchiMate model (Archisurance, 17 views) the sizes are 8,094 triples for full, 3,855 for no-geometry and 2,038 for no-diagrams — geometry is about three quarters of the output.

Profiles filter the built model rather than changing what the emitters produce, which is only well defined because the semantic/views boundary is now the same in every converter. no-diagrams also removes the folder entries that listed the removed diagrams, so no folder is left asserting a containment whose other half is gone.

The --format flag

--format is the default for any --output that names no format of its own. It is optional in every converter; before, three of the six required it while the other three inferred from the extension.

All converters support the same values:

Value Extension Named graphs Description
TRIG .trig ✓ preserved Recommended. W3C TriG — superset of Turtle with named graph blocks.
TURTLE .ttl ✗ merged W3C Turtle — all named graphs flattened into a single default graph.
JSONLD .jsonld ✓ preserved JSON-LD — JSON-based RDF. Named graphs via @graph.
RDFXML .rdf ✗ merged RDF/XML — legacy XML serialization. No named graph support.
NTRIPLES .nt ✗ merged N-Triples — one triple per line, no prefixes.
NQUADS .nq ✓ preserved N-Quads — N-Triples plus a graph term. One quad per line, no prefixes: the streaming and line-processing format that still says which graph a fact is in.

Why TriG is the default

A conversion produces four kinds of named graph:

{base}{notation}/{modelId}/graph/semantic          ← facts lifted from an input, including the views themselves
{base}{notation}/{modelId}/graph/semantic/{slug}   ← the same, one graph per input, where a model has several
{base}{notation}/{modelId}/graph/model             ← the curated model: arch:Model, folders, ordering
{base}{notation}/{modelId}/graph/views             ← diagram geometry only
{base}{notation}/{modelId}/graph/provenance        ← PROV-O about all of the above

With TriG these are preserved as separate contexts. With Turtle they get merged and the graph boundaries are lost.

The semantic/views line falls around geometry rather than around diagrams: the arch:View resource is a model concept and sits in semantic, while views holds only archvis: nodes, links, points and their bounds. That is what makes views safe to drop on its own — for an ArchiMate model it is around 80% of the output — without losing the diagrams or their names.

semantic and model

semantic holds what was lifted from an input. model holds what the conversion curated: the arch:Model resource, its metamodel conformance, its folders and their ordering, and each concept's folder membership. None of that comes from a diagram or a descriptor — its input is the diagram index — and separating it is what lets semantic be attributed to a source at all. Measured on a two-descriptor catalog, 27 of 60 semantic triples were curation.

One semantic graph per input

Where a model is built from several inputs, each one's facts go in graph/semantic/{slug}, named after that input: its repository path where the run knows one, else its file name.

{base}backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml
{base}backstage/service-catalog/graph/semantic/group-payments/catalog-info-yaml

That makes "everything this input produced" a single GRAPH clause, and it makes a fact two inputs both assert attributable to each of them — a resource carrying two prov:wasDerivedFrom values cannot say which value came from which file.

A model built from one input keeps the bare graph/semantic: partitioning one input would name a graph after the only file there is. The rule follows from the inputs and there is no flag for it.

The slug carries no commit sha, deliberately. A graph holds the current facts from a file and a later run replaces it, so it is named after the file; the prov:Entity describing that file keeps its sha, because it identifies bytes at a revision. Keyed on the commit too, a re-pull that moved one file's commit would write a second graph beside the first, and a union would return that file's facts once per revision it had ever been read at.

Only backstage2linkedarchi partitions today. The other five write one graph/semantic per model.

Every graph says where it came from

Each graph is described in graph/provenance as a prov:Bundle:

<…/graph/semantic/group-orders/catalog-info-yaml> a prov:Bundle ;
    prov:wasGeneratedBy  <{base}provenance/run/8d5a6e38> ;
    prov:wasDerivedFrom  <{base}provenance/source/group-orders/4f2c1ab8e0d1/catalog-info-yaml> ;
    prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime .

prov:Bundle and not also prov:Entity: PROV-O makes Bundle a subclass of Entity, so the second type adds nothing a reasoner needs, and ?s a prov:Entity is how a consumer reaches the inputs of a run. A graph is not an input.

Because the derivation is carried by the graph, the per-concept prov:wasDerivedFrom is recoverable rather than load-bearing:

CONSTRUCT { ?c prov:wasDerivedFrom ?src }
WHERE {
  ?g a prov:Bundle ; prov:wasDerivedFrom ?src .
  GRAPH ?g { ?c a arch:ModelConcept }
}

TriG output structure

@prefix arch: <https://meta.linked.archi/core#> .
@prefix bpmn: <https://meta.linked.archi/bpmn/onto#> .
@prefix archvis: <https://meta.linked.archi/core-vis#> .
@prefix dct: <http://purl.org/dc/terms/> .

# ─── Semantic graph: facts lifted from the input ───────────────────
<https://example.org/la/bpmn/demo/graph/semantic> {

    <.../element/Task_ValidateOrder> a bpmn:UserTask, arch:Element, arch:ModelConcept ;
        arch:inModel <.../bpmn/demo> ;
        bpmn:name "Validate Order" ;
        bpmn:id "Task_ValidateOrder" .

    <.../relationship/Flow_1> a bpmn:SequenceFlow, arch:QualifiedRelationship, arch:ModelConcept ;
        arch:inModel <.../bpmn/demo> ;
        arch:source <.../element/Start_1> ;
        arch:target <.../element/Task_ValidateOrder> .
}

# ─── Model graph: the curated model ────────────────────────────────
<https://example.org/la/bpmn/demo/graph/model> {

    <.../bpmn/demo> a arch:Model ;
        arch:modelConformsToMetamodel <https://meta.linked.archi/bpmn/metamodel#Bpmn202> .

    <.../bpmn/demo/folder> a arch:Folder ;
        schema:name "demo" ;
        schema:itemListElement [ a schema:ListItem ;
            schema:position 1 ;
            schema:item <.../bpmn/demo/folder/Elements> ] .

    <.../element/Task_ValidateOrder> dct:isPartOf <.../bpmn/demo/folder/Elements> .
}

# ─── Views graph: diagram geometry ─────────────────────────────────
<https://example.org/la/bpmn/demo/graph/views> {

    <.../view/.../node/Shape_1> a archvis:ArchNode ;
        archvis:view <.../view/Diagram_1> ;
        archvis:archElement <.../element/Task_ValidateOrder> ;
        archvis:bounds-x "270"^^xsd:double ;
        archvis:bounds-y "190"^^xsd:double ;
        archvis:bounds-w "100"^^xsd:double ;
        archvis:bounds-h "56"^^xsd:double .
}

# ─── Provenance graph: how the graph came to be ────────────────────
<https://example.org/la/bpmn/demo/graph/provenance> {

    <.../bpmn/demo>
        # the input this model came from: the blob IRI where the run can name one, the filename otherwise
        dct:source <https://git.example.org/group/models/-/blob/8c44d1a4e9f0/models/bpmn/order/process.bpmn> ;
        prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
        prov:wasGeneratedBy <https://example.org/la/provenance/run/8d5a6e38> ;
        prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .

    <https://example.org/la/provenance/run/8d5a6e38> a prov:Activity ;
        prov:endedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
        prov:wasAssociatedWith <https://example.org/la/provenance/agent/b01e544f> ;
        prov:used <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> ;
        dct:identifier "https://git.example.org/group/models/-/jobs/9182734" .

    <https://example.org/la/provenance/agent/b01e544f>
        a prov:SoftwareAgent, schema:SoftwareApplication ;
        schema:name "bpmn2linkedarchi" ;
        schema:softwareVersion "1.3.0" .

    <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn>
        a prov:Entity ;
        schema:name "models/bpmn/order/process.bpmn" ;
        dct:isPartOf <https://git.example.org/group/models> ;
        dct:identifier "8c44d1a4e9f0112233445566778899aabbccddee" ;
        prov:wasGeneratedBy <https://example.org/la/provenance/commit/8c44d1a4e9f0> .

    # each graph this run wrote, and what it came from
    <.../bpmn/demo/graph/semantic> a prov:Bundle ;
        prov:wasGeneratedBy <https://example.org/la/provenance/run/8d5a6e38> ;
        prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
        prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .

    <.../bpmn/demo/graph/views> a prov:Bundle ; … .

    # the curated graph's input is the index, not the diagram
    <.../bpmn/demo/graph/model> a prov:Bundle ;
        prov:wasGeneratedBy <https://example.org/la/provenance/run/8d5a6e38> ;
        prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-diagram-index-yaml> .
}

The provenance nodes are not under the model's IRI. {base}provenance/{kind}/… is a namespace shared by every model and every run, so one file at one commit is one node however many models read it, and one run is one activity however many models it converted. The statements about those nodes go into each model's own provenance graph, which is why this graph is still a complete description on its own. See ADR 0008 §2.

This matters for the format choice below. The derivation is asserted on the model resource as well as on each graph, so a Turtle consumer keeps it. The graph-level statements survive flattening too, as statements: the graph names stop naming anything, so "which graph is this triple in" is gone, but "what did this run write, and from what" is still answerable.

Turtle output (same data, flat)

@prefix arch: <https://meta.linked.archi/core#> .
@prefix bpmn: <https://meta.linked.archi/bpmn/onto#> .

<.../element/Task_ValidateOrder> a bpmn:UserTask, arch:Element, arch:ModelConcept ;
    arch:inModel <.../bpmn/demo> ;
    bpmn:name "Validate Order" .

<.../relationship/Flow_1> a bpmn:SequenceFlow, arch:QualifiedRelationship, arch:ModelConcept ;
    arch:inModel <.../bpmn/demo> ;
    arch:source <.../element/Start_1> ;
    arch:target <.../element/Task_ValidateOrder> .

<.../view/.../node/Shape_1> a archvis:ArchNode ;
    archvis:view <.../view/Diagram_1> ;
    archvis:archElement <.../element/Task_ValidateOrder> .

<.../bpmn/demo> prov:generatedAtTime "2026-06-06T07:33:44Z"^^xsd:dateTime ;
    prov:wasDerivedFrom <https://example.org/la/provenance/source/group-models/8c44d1a4e9f0/models-bpmn-order-process-bpmn> .

Graph boundaries lost, provenance and membership kept

In Turtle output there is no way to tell a semantic triple from a provenance triple after the fact, and "which input asserted this fact" is gone with the graph names. Use TriG when either matters.

What survives is everything asserted on a resource rather than carried by a graph boundary: arch:inModel, so the model is still reassemblable; the model's own prov:wasGeneratedBy and prov:wasDerivedFrom; and the bundle descriptions, so the inventory of what the run wrote and from what is intact even though the graph names no longer name graphs.

dct:created is not among them: it means the date the author declared, and the conversion time is prov:generatedAtTime on the model and prov:endedAtTime on the run (ADR 0008 §1).

When to use which

Scenario Format
Loading into a triplestore (Jena/Fuseki, GraphDB, Oxigraph) TRIG
SPARQL queries that filter by graph TRIG
Merging output from multiple converters into one dataset TRIG
Simple inspection / diffing / grepping TURTLE
Tools that only accept Turtle (some validators) TURTLE
Human reading in a text editor TURTLE
Streaming, grep/awk, or bulk loading, with graphs intact NQUADS
JSON-based workflows / APIs JSONLD

SPARQL examples (TriG-loaded data)

-- All elements in the semantic graph only
SELECT ?elem ?type ?label
FROM NAMED <https://example.org/la/bpmn/demo/graph/semantic>
WHERE {
    GRAPH <https://example.org/la/bpmn/demo/graph/semantic> {
        ?elem a arch:Element, ?type .
        OPTIONAL { ?elem bpmn:name ?label }
    }
}

-- Provenance: what produced this model, from which file, and when?
--
-- Only four patterns are guaranteed: the two derivations, the source's `schema:name`, and the run's agent.
-- Everything else is OPTIONAL by design — `--run-timestamp none` omits the times, an unpinned run
-- makes no image claim, and a conversion outside a known forge names no repository. Binding any of
-- them non-optionally returns zero rows instead of a partial answer.
SELECT ?ended ?repo ?path ?commit ?converter ?image
WHERE {
    GRAPH <https://example.org/la/bpmn/demo/graph/provenance> {
        ?model prov:wasGeneratedBy ?run ;
               prov:wasDerivedFrom ?src .
        ?src   schema:name ?path .
        ?run   prov:wasAssociatedWith ?agent .
        ?agent schema:name ?converter .
        OPTIONAL { ?run   prov:endedAtTime ?ended }
        OPTIONAL { ?agent dct:identifier ?image }
        OPTIONAL { ?src   dct:isPartOf ?repo }
        OPTIONAL { ?src   dct:identifier ?commit }
    }
}

-- All view nodes with geometry
SELECT ?node ?x ?y ?w ?h
WHERE {
    GRAPH <https://example.org/la/bpmn/demo/graph/views> {
        ?node a archvis:ArchNode ;
            archvis:bounds-x ?x ; archvis:bounds-y ?y ;
            archvis:bounds-w ?w ; archvis:bounds-h ?h .
    }
}

-- Everything one input produced, named by its path rather than by its graph IRI
SELECT ?s ?p ?o
WHERE {
    ?g prov:wasDerivedFrom ?src .
    ?src schema:name "catalogs/orders/catalog-info.yaml" .
    GRAPH ?g { ?s ?p ?o }
}

-- Which input asserted a given fact. Two answers where two inputs both assert it.
SELECT ?path
WHERE {
    GRAPH ?g { <…/element/component/default/order-service> bs:lifecycleState bs:Production }
    ?g prov:wasDerivedFrom ?src .
    ?src schema:name ?path .
}

-- The model, reassembled across its partitions. Works in Turtle too.
SELECT ?concept
WHERE { ?concept arch:inModel <https://example.org/la/backstage/service-catalog> }

The first two need the store's default graph to be a union of the named graphs, or the ?g patterns outside the GRAPH clause will not match. That is store configuration, not something the output can carry — see the aggregation graph.

Usage examples

# Explicit format (recommended)
java -jar bpmn2linkedarchi.jar convert process.bpmn \
  --format TRIG --base-iri https://example.org/la/ -o out.trig

# Turtle for human inspection
java -jar bpmn2linkedarchi.jar convert process.bpmn \
  --format TURTLE --base-iri https://example.org/la/ -o out.ttl

# JSON-LD for API consumption
java -jar bpmn2linkedarchi.jar convert process.bpmn \
  --format JSONLD --base-iri https://example.org/la/ -o out.jsonld