Skip to content

Converter Alternatives

This document evaluates alternative approaches to building Linked.Archi converters — tools that transform source architecture models into the Linked.Archi knowledge graph.

The evaluation is framed from the Linked.Archi perspective: what does the platform need from its converters, and how well does each approach satisfy those needs?


What Linked.Archi Requires from a Converter

A Linked.Archi converter must produce an RDF dataset that satisfies a strict contract:

  1. Foundational typing — every element carries arch:Element, every relationship carries arch:QualifiedRelationship, every diagram carries arch:Diagram
  2. Qualified Relationship Pattern — relationships are first-class nodes with arch:source, arch:target, optional direct triples, qualified triples, and RDF 1.2 reification
  3. Proper RDF lists — bendpoints on archvis:points must form valid rdf:first/rdf:rest/rdf:nil chains (required for SPARQL property paths in rdf2docs SVG rendering)
  4. Four named graphs — semantic, model, views, provenance — with dynamically minted graph IRIs per model, and semantic split one graph per input file where a model has several
  5. Stable IRI minting — based on source identifiers, never display names, with deterministic path segments
  6. SKOS labelling — skos:prefLabel, skos:notation, skos:definition for uniform label resolution across all notations
  7. Folder structure — arch:Folder hierarchy with ordered schema:itemListElement entries driving the rdf2docs sidebar
  8. Triple typing — notation type + domain type + foundational type on every resource
  9. Configurable mapping — type-mapping YAML for domain ontology alignment, with fallback, strict, and skip-unmapped modes
  10. View geometry — archvis:ArchNode with exact bounds predicates, archvis:Link with source/target references

These constraints exist because rdf2docs is notation-agnostic — it renders any model that hits the contract, with zero notation-specific code. Breaking the contract means broken rendering, broken SPARQL queries, or broken SHACL validation.


Approach 1: Custom Code Converters (Current)

Architecture: Parser → Intermediate Model → Emitter (Kotlin + RDF4J)

How It Works

Each converter is a JVM application with three layers: - A parser (StAX/SAX) that reads the source format into plain data classes - An emitter (RDF4J ModelBuilder) that transforms the intermediate model to RDF - A CLI (Picocli) that wires configuration, I/O, and format selection

A shared core module provides the foundational ontology constants, IRI minting utilities, folder structure builder, and the emitter base class. Each notation-specific converter only implements parsing and type resolution.

Strengths for Linked.Archi

  • Full control over the output shape — the contract is enforced at compile time
  • Domain-aware error reporting ("relationship X references unknown element Y")
  • Handles complex parsing scenarios: two-pass XML, element-is-view patterns, conditional emission based on CLI flags
  • The shared core module guarantees cross-notation consistency
  • Streaming (StAX) for large models without memory pressure
  • Built-in SHACL validation as a first-class command
  • CI/CD ready: fat JAR, Docker, GitLab CI templates

Limitations

  • Each new notation requires writing a parser from scratch
  • Mapping logic lives in code — changes require recompilation
  • Higher initial effort per converter compared to declarative approaches
  • Kotlin/JVM ecosystem dependency

When This Approach Is Essential

  • Source formats with complex structure (nested XML with forward references)
  • Notations where the element-is-view pattern applies (BPMN Process)
  • Sources requiring two-pass parsing (ArchiMate property definitions)
  • When the qualified relationship pattern, RDF lists, or named graph splitting is central to the output

Approach 2: RML (RDF Mapping Language)

Architecture: Declarative mapping rules → Generic RML engine → RDF output

How It Works

RML extends R2RML (the W3C standard for relational-to-RDF mapping) to support heterogeneous sources: XML, JSON, CSV, APIs. You write mapping documents that declare how source data maps to RDF triples using logical sources, subject maps, and predicate-object maps.

Engines: RMLMapper (reference), Morph-KGC (Python, streaming), SDM-RDFizer.

What It Can Handle

  • Flat-to-moderate source structures (YAML, JSON, tabular data)
  • Simple type mapping: source field → RDF type
  • IRI templates: {base}element/{id} patterns
  • Named graph assignment via rr:graphMap
  • Join conditions between logical sources (for relationships)
  • FnO (Function Ontology) extensions for string manipulation, conditionals

Where It Struggles with Linked.Archi Requirements

Requirement RML Capability
Qualified Relationship Pattern Requires multiple TriplesMap entries with joins per relationship; verbose and fragile
Proper RDF lists (bendpoints) Not natively supported — requires post-processing or custom FnO functions
Conditional emission (flags) Outside RML's scope — would need preprocessing or separate mapping files per mode
Two-pass parsing Not supported — sequential processing only
Type normalisation (strip suffixes, fallback chains) Requires FnO functions for each rule
Element-is-view pattern Control flow logic cannot be expressed declaratively
Cross-notation contract enforcement No shared module concept — each mapping file is independent

Where It Could Complement Linked.Archi

  • Simple flat sources: Backstage catalog-info.yaml, CSV inventories, JSON APIs where the source structure is regular and the mapping is straightforward
  • User-defined enrichment: A thin RML layer on top of converter output to pull in additional properties from external sources (CMDBs, spreadsheets) without modifying converter code
  • Rapid prototyping: Quickly produce approximate RDF from a new source to validate feasibility before committing to a full converter

Operational Considerations

  • RMLMapper is JVM-based (same ecosystem as current converters)
  • Morph-KGC is Python-based (different runtime, better streaming)
  • Mapping files are verbose for complex patterns — can exceed the complexity of equivalent code while being less expressive
  • No built-in validation — SHACL checking is a separate pipeline step
  • Error messages are generic (join failures, template errors) rather than domain-specific

Approach 3: SPARQL Anything

Architecture: Virtual RDF graph over source file → SPARQL CONSTRUCT → RDF output

How It Works

SPARQL Anything (based on Apache Jena) exposes any file (XML, JSON, CSV, YAML, binary formats) as a virtual RDF graph using the SPARQL SERVICE clause with the x-sparql-anything: protocol. You write CONSTRUCT queries that reshape the virtual triples into your target ontology.

What It Can Handle

  • Ad-hoc querying of any source format without writing a parser
  • CONSTRUCT queries for reshaping data
  • SPARQL's full expressiveness: FILTER, BIND, string functions, aggregation
  • Multiple sources in one query via multiple SERVICE clauses
  • Federated queries combining local files with remote SPARQL endpoints

Where It Struggles with Linked.Archi Requirements

Requirement SPARQL Anything Capability
Named graph datasets (TriG output) Designed for CONSTRUCT into default graph; named graph output is awkward
Proper RDF lists CONSTRUCT cannot produce well-formed rdf:first/rdf:rest chains from sequences
Qualified Relationship Pattern Expressible but extremely verbose — multiple CONSTRUCT blocks with shared variables
Configurable type-mapping YAML No native config mechanism — mapping logic lives in the query itself
Large models (10k+ elements) Entire source loaded into virtual RDF model in memory
Streaming Not supported — full materialisation before query execution
CLI flags / conditional emission Would need query generation or template preprocessing
Stable IRI minting with encoding SPARQL's ENCODE_FOR_URI exists but complex path assembly is cumbersome

Where It Could Complement Linked.Archi

  • Source exploration and discovery: Before building a converter, use SPARQL Anything to query an unfamiliar source format interactively — understand its structure, identify elements/relationships/views, find stable identifiers
  • Validation queries: Query converter output alongside source files to verify completeness (e.g., "which source elements have no corresponding RDF resource?")
  • One-off extractions: Pull specific data from a source without building a full converter (e.g., extract just the element names for a report)
  • Prototyping target shape: Quickly test whether a source has enough information to satisfy the Linked.Archi contract before investing in a full converter

Operational Considerations

  • Apache Jena dependency (separate from RDF4J used by current converters)
  • No incremental processing — full source must fit in memory
  • Queries become unmaintainable for complex transformations
  • No compile-time contract enforcement
  • Good for interactive exploration, poor for production pipelines

Approach 4: Hybrid (Code Converters + Declarative Layers)

Architecture: Code converters for complex sources + RML/SPARQL for enrichment

How It Works

Keep the current code converter architecture for sources that require it (ArchiMate, BPMN, PlantUML), but add a declarative layer for:

  1. Simple sources — new converters for flat formats (Backstage YAML, LeanIX JSON exports, ServiceNow CMDB dumps) could use RML mappings executed by the same Gradle pipeline, with a thin wrapper that ensures the Linked.Archi contract is met
  2. Post-conversion enrichment — RML mappings that merge additional data into existing graph elements (e.g., attach cost data from a CSV to ArchiMate ApplicationComponents via owl:sameAs / @id alignment)
  3. Exploration — SPARQL Anything as a developer tool for investigating new source formats before committing to a converter implementation

Contract Enforcement in the Hybrid Model

The risk of declarative mappings is contract drift. Mitigations:

  • Run SHACL validation (validate command) on all output regardless of how it was produced — the published shapes catch missing arch:Element types, broken relationship endpoints, malformed view structures
  • Provide a "Linked.Archi RML template" with pre-built subject maps, graph maps, and predicate-object maps for the foundational types — new mappings start from this template rather than from scratch
  • Integration tests in CI that verify rdf2docs can render the output (SVG smoke test)

When to Use Which

Source Complexity Recommended Approach
Complex XML with forward references, multi-pass needs Code converter
Notation where element-is-view pattern applies Code converter
Sources requiring RDF lists (diagram geometry) Code converter
Flat YAML/JSON with regular structure RML mapping (with template + SHACL gate)
CSV/tabular enrichment of existing graph RML mapping
Unknown source format (exploration phase) SPARQL Anything
One-off data extraction SPARQL Anything

Summary

Approach Contract Safety Complex Sources Simple Sources Exploration Production Ready
Code converters High (compile-time) Excellent Adequate Poor Yes
RML Low (runtime only) Poor Good Moderate Yes (with SHACL gate)
SPARQL Anything None (query-time) Poor Moderate Excellent No
Hybrid High (layered) Excellent Good Excellent Yes

The Linked.Archi platform's strength is its rigid contract — that rigidity is what makes rdf2docs, SHACL validation, and cross-notation SPARQL queries work without notation-specific code. Any alternative approach must be evaluated primarily on whether it can reliably produce output that satisfies that contract. For the core converters handling complex notations, custom code remains the right choice. For simpler sources and enrichment scenarios, declarative approaches can reduce effort while the SHACL validation layer provides the safety net.

The Complexity Threshold

The key insight is that declarative mapping only wins while the mapping stays simple. Once a source requires the full Linked.Archi contract — the qualified relationship pattern, proper RDF lists, named graph splitting, conditional emission, type normalisation — the RML mapping or SPARQL CONSTRUCT query becomes complex enough to be effectively equivalent to programming. But it is programming in a less expressive language, without type safety, without the shared core module enforcing the contract at compile time, and without domain-specific error messages.

At that point the declarative approach has all the cost of code with none of its advantages. This is why the decision hinges on source complexity: below the threshold, declarative mapping saves effort; above it, custom code is both easier to write and safer to maintain.