Skip to content

Linked.Archi Converter Development Guide

This guide explains how to build a new Linked.Archi converter — a tool that reads a source modelling notation (ArchiMate, BPMN, UML, C4, LeanIX, etc.) and produces an RDF dataset aligned to the Linked.Archi foundational ontology.

It is written for both human developers and AI coding agents (e.g. Kiro, Amazon Q).


1. What a converter does

A converter takes a source model file and produces a TriG/Turtle RDF dataset with:

  1. Semantic graph — elements, relationships, and views typed against the Linked.Archi foundational ontology and optionally a domain ontology
  2. Views graph — visual scaffolding (nodes, links, geometry) for diagram rendering
  3. Provenance graph — conversion metadata (timestamp, source file, tool name)
  4. Folder structure — arch:Folder hierarchy for navigation in rdf2docs

The output must be consumable by rdf2docs without any notation-specific code in the generator. This is the key design constraint: the generator is model-agnostic.


2. Foundational ontology contract

Every converter MUST emit these types and predicates. They are the interface between converters and consumers (rdf2docs, SPARQL queries, SHACL validation).

Required namespaces

Prefix IRI Purpose
arch: https://meta.linked.archi/core# Foundational types and predicates
archvis: https://meta.linked.archi/core-vis# View/diagram ontology
skos: http://www.w3.org/2004/02/skos/core# Labels and notations
dct: http://purl.org/dc/terms/ isPartOf, created, source
schema: https://schema.org/ Folder membership, properties
rdf: http://www.w3.org/1999/02/22-rdf-syntax-ns# Types, lists

Required triples per concept

Element:

<element/{id}>
    a arch:Element, <domain-type> ;
    skos:notation "{id}" ;
    skos:prefLabel "{name}"@en ;
    dct:isPartOf <{notation}/{model-id}/folder/Elements> .

Relationship:

<relationship/{id}>
    a arch:QualifiedRelationship, <domain-type> ;
    arch:source <element/{source-id}> ;
    arch:target <element/{target-id}> ;
    skos:notation "{id}" ;
    dct:isPartOf <folder/relationships> .

View (diagram):

<view/{id}>
    a arch:Diagram, <notation-view-type> ;
    skos:notation "{id}" ;
    skos:prefLabel "{name}"@en ;
    dct:isPartOf <folder/views> .

View node:

<view/{view-id}/node/{node-id}>
    a archvis:ArchNode ;
    archvis:view <view/{view-id}> ;
    archvis:archElement <element/{elem-id}> ;
    archvis:bounds-x {x} ; archvis:bounds-y {y} ;
    archvis:bounds-w {w} ; archvis:bounds-h {h} .

View link:

<view/{view-id}/link/{link-id}>
    a archvis:Link ;
    archvis:view <view/{view-id}> ;
    archvis:archRelationship <relationship/{rel-id}> ;
    archvis:source <view/{view-id}/node/{src-node-id}> ;
    archvis:target <view/{view-id}/node/{tgt-node-id}> ;
    archvis:points ( [ archvis:point-x {x} ; archvis:point-y {y} ] ... ) .

Folder:

<folder/{path}>
    a arch:Folder ;
    schema:name "{display-name}" ;
    dct:isPartOf <folder/{parent-path}> ;
    schema:itemListElement
        [ a schema:ListItem ; schema:position 1 ; schema:item <element/{id}> ] .


3. IRI minting strategy

All IRIs follow the pattern:

{base-iri}{notation}/{model-id}/{segment}/{local-id}
Segment Used for
element/ Semantic elements
relationship/ Semantic relationships
view/ Views / diagrams (or use element/ when the element IS the view)
view/{view-id}/node/ View nodes
view/{view-id}/link/ View links
graph/semantic Named graph for facts lifted from an input
graph/model Named graph for the curated model: arch:Model, conformance, folders, ordering, membership
graph/views Named graph for view triples
graph/provenance Named graph for provenance triples

The {notation} segment is fixed per converter: archimate, bpmn, uml, etc. The {base-iri} is the shared graph root passed via CLI (e.g. https://example.org/la/).

IRI stability rules

  • Use the source tool's stable identifier (e.g. ArchiMate identifier, BPMN id attribute)
  • Never use display names or labels in IRIs — they change
  • For anonymous elements (no id), derive from parent IRI: {parentIri}/{localName}
  • Add a counter suffix for duplicates: {parentIri}/{localName}_2
  • Last resort only: {base}#_gen_{uuid} — avoid this

4. Named graph structure

Emit a TriG dataset with four named graphs per model. Take the suffixes from IriMinting.GraphSuffix rather than writing the strings yourself:

{base}{notation}/{model-id}/graph/semantic    <- elements, relationships, views: what the input carried
{base}{notation}/{model-id}/graph/model       <- arch:Model, conformance, folders, ordering, membership
{base}{notation}/{model-id}/graph/views       <- archvis: nodes, links, geometry
{base}{notation}/{model-id}/graph/provenance  <- dct:source, PROV activity, adms:status, the bundles

When outputting Turtle (no named graph support), merge all graphs into one.

semantic versus model: did an input carry it, or did the conversion decide it? A concept's types and attributes came out of the file, so they are semantic. The arch:Model resource, the folder tree, the ordering and each concept's folder membership are the converter's own arrangement — the index is their only input — so they are model. Keeping them apart is what lets semantic be attributed to a source: a graph mixing the two has triples no input produced.

Two obligations follow:

  • Emit arch:inModel on every concept, into the same graph as the concept. It is what makes the model reassemblable without relying on a graph boundary, so it must not live in graph/model.
  • Write folder membership in graph/model, both directions. emitFolderMember in BaseLinkedArchiEmitter does both halves; do not add a dct:isPartOf beside the concept's own triples, which would put one relation in two graphs.

Describe the graphs you wrote. emitProvenance takes liftedGraphs and curatedGraphs, and turns each into a prov:Bundle saying what generated it and from what. A converter that partitions graph/semantic by input describes those graphs itself — see ConversionProvenance.emitBundles — because only it knows which input each one holds.

Writing the output

Build the model once, then hand it to RdfIo.writeAll with the specs from OutputSpec.fromOptions:

@Option(names = ["-o", "--output"], required = true, description = [OutputSpec.OUTPUT_HELP])
lateinit var outputs: List<String>

@Option(names = ["--format"], description = [OutputSpec.FORMAT_HELP])
var format: String? = null

// …

val specs = OutputSpec.fromOptions(outputs, format, warn = ::warn)   // before parsing, so a bad
                                                                     // spec fails the run at once
val merged = /* build */
RdfIo.writeAll(merged, specs, nsCore, nsCoreVis) { spec, triples -> report(spec, triples) }

--output is repeatable, so one conversion writes every artifact a repository publishes. Do not loop the whole conversion per artifact: besides repeating the parse, each run resolves its own ConversionProvenance.Run, so the files would claim different prov:generatedAtTime values and an aggregation store would see several publishes where there was one.

writeAll applies the view profile per spec and caches by profile, so a .trig and a .ttl of the same profile filter once.

Where the semantic/views boundary falls

The line is drawn around geometry, not around diagrams. It is easy to get wrong, and two converters did:

  • The arch:View resource goes in the semantic graph, with its skos:notation, label, arch:viewConformsToViewpoint, dct:isPartOf and properties. A diagram is a model concept — something the model contains — so a query for the model's diagrams scoped to the semantic graph must find it.
  • Only the presentation of that diagram goes in the views graph: archvis:ArchNode, archvis:Link, archvis:LabelNode, archvis:Point, their bounds, waypoints and styles.

Two rules follow, and both are worth checking before you open a merge request:

Never split one subject across the two graphs. If half a view's description is in semantic and half in views, a semantic-only query returns a typed diagram with no name, and a consumer that drops the views graph to shed geometry loses the diagram's label with it. Dropping geometry is the one operational reason the split exists — it is ~80% of an ArchiMate artifact — so it has to stay safe to do.

Never split one fact across the two graphs. Folder containment is stated twice, once from each end: the member's dct:isPartOf and the folder's ordered schema:itemListElement. Both belong in semantic, where the folder tree lives. Put one in views and neither graph can answer "what is in this folder" on its own.

The folder tree lists elements, relationships and views. Geometry is not a folder member; it is reached through archvis:view.

BpmnViewGraphBoundaryTest and the boundary tests in ArchiMateViewsTest pin all of this down. Worth copying for a new converter, since a misfiled triple is still present and still correct — nothing fails until something asserts on the graph.


5. Type-mapping YAML

Every converter should support a --type-mapping YAML file that maps source type names to domain ontology IRIs. The structure is shared across all converters:

# Source type name -> domain IRI for elements
elements:
  BusinessActor:   https://metadata.example.org/ont#BusinessActor

# Source type name -> domain IRI for relationships
relationships:
  Serving:         https://metadata.example.org/ont#Serving

# Relationship type -> direct predicate IRI (for --emit-direct-rel-triples)
predicates:
  Serving:         https://metadata.example.org/ont#serves

# Relationship type -> qualified predicate IRI, from the source element to the
# relationship resource. Always emitted; arch:hasQualifiedRelationship if unmapped.
qualifiedPredicates:
  Serving:         https://metadata.example.org/ont#qualifiedServes

# Namespace prefix declarations
namespaces:
  myont:           https://metadata.example.org/ont#

Fallback when a type is not listed: {notation-namespace}{TypeLocalName}.


6. View rendering contract

For rdf2docs to render an SVG diagram, the following must hold:

  1. The view IRI has a arch:Diagram (or arch:View)
  2. Each node has archvis:view <viewIri> and archvis:bounds-x/y/w/h
  3. Each link has archvis:view <viewIri>, archvis:source, archvis:target
  4. Bendpoints are stored as a proper RDF list on archvis:points:
    archvis:points ( [ archvis:point-x 100 ; archvis:point-y 200 ]
                     [ archvis:point-x 150 ; archvis:point-y 250 ] ) .
    
    This enables the SPARQL path archvis:points/rdf:rest*/rdf:first to work.
  5. SVG z-order: nodes must come before links in the SVG output so links render on top of container rectangles.

When the element IS the view

Some notations (BPMN Process, UML Package) are both semantic elements and diagram containers. In this case:

  • Use the element IRI as the view IRI (do not mint a separate view/ IRI)
  • Emit arch:Diagram on the element IRI
  • Point all archvis:view triples at the element IRI

This avoids a separate diagram page and makes the element the natural entry point in rdf2docs, with SVG, Elements, and Relationships tabs all on the same page.


7. Folder structure

The folder structure drives the rdf2docs sidebar. Every element and relationship must have dct:isPartOf pointing to a folder IRI.

Folder IRI pattern

{base}{notation}/{model-id}/folder/{name}

Minimum required folders

Folder Contents
{model-id}/Elements All elements
{model-id}/Relationships All relationships
{model-id}/Views All views / diagrams

Ordered membership

Use schema:itemListElement with schema:ListItem for ordered display:

<folder/elements>
    schema:itemListElement
        [ a schema:ListItem ; schema:position 1 ; schema:item <element/id-1> ] ,
        [ a schema:ListItem ; schema:position 2 ; schema:item <element/id-2> ] .

8. Implementation checklist

Use this checklist when building or reviewing a converter.

Parser

  • [ ] Reads source format (XML, JSON, YAML, etc.)
  • [ ] Extracts elements with stable IDs, types, names
  • [ ] Extracts relationships with source/target IDs
  • [ ] Extracts views with node/link geometry
  • [ ] Handles anonymous elements (no ID) with parent-derived IRIs

Emitter

  • [ ] Emits arch:Element on all elements
  • [ ] Emits arch:QualifiedRelationship on all relationships
  • [ ] Emits arch:ModelConcept on all elements and relationships
  • [ ] Emits arch:source / arch:target on all relationships
  • [ ] Emits arch:Diagram on all views (or on the element that IS the view)
  • [ ] Emits skos:prefLabel for names
  • [ ] Emits skos:notation for IDs
  • [ ] Emits dct:isPartOf on all elements, relationships, views pointing to a folder
  • [ ] Emits arch:inModel on every element, relationship and view, pointing at the model. Membership carried only by the named graph is lost the moment the output is Turtle, RDF/XML or N-Triples
  • [ ] Emits arch:Folder hierarchy with schema:name and schema:itemListElement
  • [ ] Emits archvis:ArchNode with archvis:view and archvis:bounds-* for view nodes
  • [ ] Emits archvis:Link with archvis:view, archvis:source, archvis:target for view links
  • [ ] Emits archvis:points as a proper RDF list for bendpoints
  • [ ] Emits the provenance graph via emitProvenance / ConversionProvenance, not by hand. Dublin Core there describes the diagram — dct:title, dct:creator, dct:created as the author declared them — and PROV describes the run: prov:wasGeneratedBy an activity, prov:wasDerivedFrom a source entity, prov:generatedAtTime on the model. A conversion timestamp under dct:created, or the converter's name under dct:creator, puts two answers on one predicate (ADR 0008)
  • [ ] Passes opts.baseIri to emitProvenance, so provenance nodes land in the shared {base}provenance/… namespace rather than under the model
  • [ ] Supports --type-mapping YAML for domain type overrides
  • [ ] Supports --base-iri for IRI minting
  • [ ] Supports --model-id for model scoping

CLI

  • [ ] --base-iri (shared graph root)
  • [ ] -o / --output (output file path)
  • [ ] --format (TRIG / TURTLE / JSONLD / RDFXML / NTRIPLES / NQUADS, inferred from extension)
  • [ ] --model-id (stable model identifier)
  • [ ] --ns-core (override arch: namespace)
  • [ ] --ns-core-vis (override archvis: namespace)
  • [ ] --type-mapping (domain type YAML)
  • [ ] --include-views / --include-di (enable view graph emission)
  • [ ] --emit-skos-labels (use skos:prefLabel for names)
  • [ ] --emit-skos-notation (use skos:notation for IDs)
  • [ ] --emit-direct-rel-triples (shortcut triples)

Verification

  • [ ] Zero _gen_ IRIs in output (all anonymous elements have parent-derived IRIs)
  • [ ] All archvis:view triples point to a valid arch:Diagram IRI
  • [ ] All archvis:source / archvis:target point to archvis:ArchNode IRIs
  • [ ] archvis:points is a proper RDF list (not a nested chain)
  • [ ] rdf2docs renders SVG correctly for at least one view
  • [ ] rdf2docs index shows all elements in the correct folders

9. Reference implementations

Notation Converter Key files
ArchiMate 3 Exchange XML converters/archimate/ ExchangeParser.kt, LinkedArchiEmitter.kt
BPMN 2.0 XML converters/bpmn/ BpmnXmlToRdfConverter.kt, LinkedArchiEmitter.kt

ArchiMate converter architecture

ArchiMate Exchange XML
  -> ExchangeParser (StAX)
  -> ExchangeModel (elements, relationships, views, folders)
  -> LinkedArchiEmitter (RDF4J ModelBuilder)
  -> TriG / Turtle output

The parser reads the Exchange XML in two passes (property definitions appear after elements in the file). The emitter builds the full RDF model in memory then writes it with RDF4J Rio.

BPMN converter architecture

BPMN 2.0 XML
  -> BpmnXmlToRdfConverter (CMOF-driven SAX parser)
  -> CMOF model (raw RDF with bpmn: / bpmndi: types)
  -> LinkedArchiEmitter (IRI remapping + archvis: emission)
  -> TriG / Turtle output

The BPMN converter is two-stage: the first stage produces a CMOF-typed RDF model using the OMG metamodel for type resolution; the second stage remaps IRIs and emits the Linked.Archi foundational types.


10. AI agent instructions (Kiro / Amazon Q)

When implementing a new converter as an AI coding task, follow these rules.

Source analysis

  1. Read the source format specification or sample files first
  2. Identify: what are the elements? what are the relationships? what are the views?
  3. Identify: what is the stable ID for each concept? (never use display names as IDs)
  4. Identify: is there a concept that is both an element AND a view? (e.g. BPMN Process)

Code structure

  1. Create a Parser class that reads the source format and produces an intermediate model (plain Kotlin data classes, not RDF)
  2. Create an Emitter class that takes the intermediate model and produces RDF using RDF4J ModelBuilder
  3. Create a ConversionOptions / Config data class for all CLI options
  4. Wire them together in a ConvertCommand (Picocli)

IRI minting

  1. Always use {base}{notation}/{model-id}/element/{id} for elements
  2. Always use {base}{notation}/{model-id}/relationship/{id} for relationships
  3. For anonymous elements: {parentIri}/{localName} with _N counter for duplicates
  4. Never use UUID.randomUUID() as a primary IRI strategy

Foundational types

  1. Always emit arch:Element on elements (when emitCoreTriples = true)
  2. Always emit arch:QualifiedRelationship on relationships
  3. Always emit arch:source / arch:target on relationships
  4. Always emit arch:Diagram on views (or on the element that IS the view)
  5. Always emit skos:prefLabel for names and skos:notation for IDs

View graph

  1. Emit archvis:ArchNode for each view node with archvis:view and bounds
  2. Emit archvis:Link for each view link with archvis:view, source, target
  3. Build bendpoint lists tail-first using rdf:first / rdf:rest / rdf:nil:
    var tail: Value = RDF.NIL
    for (point in points.reversed()) {
        val cell = vf.createBNode()
        mb.subject(cell).add(RDF.FIRST, point).add(RDF.REST, tail)
        tail = cell
    }
    mb.subject(link).add(archvisPoints, tail)
    
  4. Never use a nested chain (mb.add(list, point) then mb.subject(point)) — this produces incorrect structure for rdf:rest*/rdf:first SPARQL paths

Folder structure

  1. Emit arch:Folder for each category with schema:name and dct:isPartOf
  2. Emit schema:itemListElement with ordered schema:ListItem entries
  3. Emit dct:isPartOf on every element/relationship pointing to its folder

Verification queries

After generating output, verify with these SPARQL queries:

# All elements have arch:Element
SELECT (COUNT(?e) AS ?n) WHERE { ?e a arch:Element }

# No relationships missing endpoints
SELECT ?r WHERE {
    ?r a arch:QualifiedRelationship .
    FILTER NOT EXISTS { ?r arch:source ?s }
}

# No view nodes missing archvis:view
SELECT ?n WHERE {
    ?n a archvis:ArchNode .
    FILTER NOT EXISTS { ?n archvis:view ?v }
}

# No _gen_ IRIs remain
SELECT ?s WHERE {
    ?s ?p ?o .
    FILTER(CONTAINS(STR(?s), "_gen_"))
}

# All views have at least one node
SELECT ?v WHERE {
    ?v a arch:Diagram .
    FILTER NOT EXISTS { ?n archvis:view ?v }
}

11. Testing

Every converter should have:

  1. Unit tests for the parser — verify element/relationship/view extraction from sample files
  2. Unit tests for the emitter — verify correct RDF triples for known inputs
  3. Integration test — convert a sample file and verify the output TTL with SPARQL queries (see verification queries above)
  4. rdf2docs smoke test — run rdf2docs on the output and verify the HTML renders at least one view with an SVG diagram

Sample files should be minimal (< 10 elements) but cover: - Elements with and without names - Relationships between elements - At least one view with nodes and links - At least one anonymous element (no ID)


12. Key design decisions and rationale

Why arch:Diagram not arch:View?

arch:Diagram is a subtype of arch:View and is more specific — it implies a visual rendering with geometry. BPMN processes and ArchiMate diagrams are both diagrams in this sense. Using arch:Diagram allows rdf2docs to trigger SVG rendering for any resource typed as arch:Diagram regardless of notation.

Why is the BPMN Process the view, not BPMNDiagram?

The bpmn:Process is the semantic root — it has a name, contains all the elements, and is the concept users navigate to. The bpmndi:BPMNDiagram is a rendering artefact. Treating the Process as the view means: - One page per process (not two: one for the process, one for the diagram) - The process page shows the SVG, elements, and relationships in one place - No BPMN-specific code needed in rdf2docs

Why RDF lists for bendpoints?

The SPARQL property path archvis:points/rdf:rest*/rdf:first only works with proper RDF lists (rdf:first / rdf:rest / rdf:nil). A custom nested chain structure requires a different query pattern and breaks the generic SVG renderer.

Why dual typing (notation type + domain type + foundational type)?

  • Notation type (bpmn:UserTask, amate:BusinessActor): enables notation-specific queries and preserves the original modelling semantics
  • Domain type (tam:BusinessTask): enables cross-notation queries using your organisation's own ontology
  • Foundational type (arch:Element): enables generic queries across all notations without knowing the domain ontology

All three are needed for a complete knowledge graph.