Skip to content

Intermediate Model Strategy: POJO vs Direct RDF

Current design

The converters use two different strategies for transforming source files into RDF:

flowchart LR
    subgraph pojo["POJO Strategy (ArchiMate, PlantUML, Backstage, Structurizr)"]
        A1["Source file"] --> B1["Parser"]
        B1 --> C1["Intermediate model\n(Kotlin data classes)"]
        C1 --> D1["LinkedArchiEmitter"]
        D1 --> E1["RDF4J Model"]
    end

    subgraph direct["Direct RDF Strategy (BPMN)"]
        A2["Source file"] --> B2["CMOF-driven Parser"]
        B2 --> C2["RDF4J Model #1\n(CMOF-typed, fragment IRIs)"]
        C2 --> D2["LinkedArchiEmitter\n(IRI remapping)"]
        D2 --> E2["RDF4J Model #2\n(path IRIs, named graphs)"]
    end

POJO strategy (ArchiMate example)

XML → ExchangeParser → ExchangeModel (data classes) → LinkedArchiEmitter → RDF4J Model → serialize
      (3 XML passes)   (lists of ExchangeElement,      (iterates POJOs,
                        ExchangeRelationship, etc.)      writes triples)

The parser produces typed Kotlin data classes:

data class ExchangeElement(
    val id: String,
    val type: String?,
    val name: LangString?,
    val documentation: LangString?,
    val properties: List<ExchangePropertyValue>
)

The emitter iterates these and writes RDF:

for (e in exchangeModel.elements) {
    val eIri = elementIdToIri.getValue(e.id)
    add(model, eIri, RDF.TYPE, iri(nsCore, "Element"), ctxSemantic)
    add(model, eIri, SKOS.NOTATION, lit(e.id), ctxSemantic)
    e.name?.let { addLabel(model, eIri, it, ctxSemantic) }
}

Direct RDF strategy (BPMN example)

XML → BpmnXmlToRdfConverter → RDF4J Model (raw) → LinkedArchiEmitter → RDF4J Model (final) → serialize
      (CMOF-driven StAX,       (fragment IRIs,       (remaps IRIs,
       writes triples as        bpmn: types)           adds arch: types,
       elements are parsed)                            splits into graphs)

The parser writes directly to an RDF4J Model during XML traversal:

// Inside the StAX parsing loop:
model.add(iri, RDF.TYPE, metaModel.classIri(pkgKey, className))
model.add(parent.iri, pred, iri)
model.add(iri, pred, vf.createLiteral(attrValue))

Why two approaches exist

The BPMN spec has 144 metaclasses defined in OMG CMOF machine-readable files. Defining a Kotlin data class for each would be impractical and fragile — every spec revision would require manual updates. Instead, the CMOF-driven parser handles all classes generically by reading the metamodel at runtime.

The other notations (ArchiMate ~55 types, PlantUML ~15, C4 ~7, Backstage ~7) are small enough that typed intermediate models are practical and provide stronger guarantees.


Pros and cons

POJO intermediate model

Advantage Example
Testability val model = parser.parse(file); assertEquals("Merchant", model.elements[0].name?.value) — no RDF knowledge needed
Separation of concerns Parser understands XML quirks (xsi:type, nested bounds, attribute variants). Emitter understands RDF (named graphs, IRI patterns, SKOS). Neither knows about the other.
Debuggability Log or inspect the intermediate model to see exactly what was parsed, before any RDF is produced
Multiple consumers The same ExchangeModel could drive a JSON exporter, diff tool, or migration script — not locked to RDF
Refactoring safety Change RDF emission (e.g. switch from schema:additionalProperty to direct predicates) without touching the parser. Change XML format support without touching the emitter.
Compile-time safety Typos in field names are caught at build time. Missing fields cause compile errors.
Disadvantage Detail
Memory overhead The full model lives in memory as data classes, then again as RDF triples — double representation
Multiple XML passes ArchiMate currently reads the XML 3 times (property definitions, model, folders) because <propertyDefinitions> appears after <elements> in the Exchange format
Code duplication Each new element type needs a data class field + parser handler + emitter handler
Not scalable to huge specs Impractical for 100+ class metamodels (would need 100+ data class definitions)

Direct-to-RDF model

Advantage Detail
Spec-completeness Handles any class in the metamodel without explicit enumeration — new spec versions are handled automatically
Single memory representation Only one RDF model exists; no data class allocations
Fewer passes possible Can emit triples as elements are encountered in the XML stream
Metamodel-driven Type resolution, containment, and relationship detection happen from the CMOF spec files, not from handwritten code
Disadvantage Detail
Harder to test Must query an RDF model to verify parsing: assertTrue(model.contains(iri, RDF.TYPE, bpmnUserTask)) vs assertEquals("UserTask", element.type)
Coupled concerns Parsing and RDF emission happen together; can't test one without the other
Debugging requires RDF literacy Intermediate state is a triple store, not a simple object graph
Double model in BPMN The current implementation builds RDF model #1 (CMOF-typed, fragment IRIs) then remaps into RDF model #2 (path IRIs, named graphs) — negating the single-model advantage
Runtime type resolution All type checks are string-based at runtime; errors surface at execution time not compile time

Current rationale

Clarity of debugging and testing has more value than performance optimisation given the current state of the project and the diversity of input formats.

The project handles 5 notations (ArchiMate, BPMN, PlantUML, Structurizr, Backstage) with very different source formats (XML, JSON, YAML, custom DSL AST). Consistency in the codebase — being able to reason about each converter the same way — is more valuable than squeezing milliseconds. The POJO approach gives:

  • New contributors can understand the parser by looking at data classes
  • Tests are simple assertions on fields, not SPARQL-like model queries
  • The emitter is self-documenting: you can read it top-to-bottom and see exactly what triples are produced
  • Refactoring the output format (e.g. adding a new predicate) requires changes only in the emitter, not in the parser

The performance cost is negligible at the model sizes encountered in practice. The ArchiMate converter processes 3300 elements + 6022 relationships + properties in ~400ms total (including serialization). Typical CI pipelines spend 10-30x more time on Docker setup than on conversion.


Future optimisation path

If performance becomes a concern (e.g. processing 100K+ element models, or running converters in a hot loop), the following changes would help:

  1. Single-pass XML with deferred resolution — parse ArchiMate XML once, writing to RDF4J Model directly. When <propertyDefinitions> is reached at the end of the file, patch property name triples on already-emitted elements. This eliminates 2 of 3 XML passes and the POJO allocation.

  2. Single-model BPMN — mint path IRIs directly during XML parsing instead of fragment IRIs, eliminating the second RDF4J Model and the IRI remapping pass. Forward references (e.g. sourceRef="Task_1" before Task_1 is seen) handled via a deferred-resolution queue.

  3. Streaming serialization — instead of building the full RDF model in memory then serializing, use an RDFHandler to stream triples directly to the output file as they are produced. This reduces peak memory from O(triples) to O(1).

These are independent optimisations that can be applied per-converter without changing the output format or the public API. The intermediate model strategy is an internal implementation detail — the CLI interface, output contract, and named-graph structure remain identical.


Summary

Notation Strategy Reason
ArchiMate POJO ~55 types, complex properties, testability priority
PlantUML POJO ~15 types, clean separation, extends BaseLinkedArchiEmitter
Structurizr POJO ~7 types, simple JSON → data class mapping
Backstage POJO ~7 types, simple YAML → data class mapping
BPMN Direct RDF 144 types, CMOF-driven, spec-complete without enumeration

The BPMN approach is technically superior for large metamodels. The POJO approach is superior for developer experience and maintainability on smaller notations. Both produce identical output conforming to the Linked.Archi foundational ontology contract.