Skip to content

Validation

Every converter ships a validate subcommand that runs SHACL validation over RDF that a converter has already produced. Validation is a separate step from convert: you convert first, then validate the output file.

bpmn2linkedarchi convert process.bpmn -o out.ttl --format TURTLE
bpmn2linkedarchi validate -i out.ttl

What actually happens

All converters share one validation engine (archi.linked.converter.core.validation.ShaclValidator). A run performs these steps:

  1. Load the shape graph. The SHACL shapes are loaded into RDF4J's dedicated shapes graph (rdf4j:SHACLShapeGraph) of an in-memory ShaclSail.
  2. Load supporting ontologies. Class/property definitions are loaded into the default graph so that rdfs:subClassOf reasoning can resolve type hierarchies. Without this, a shape targeting a superclass would not apply to instances of its subclasses.
  3. Load the data and validate. Your input file is added inside a transaction using the Bulk validation approach; committing triggers validation of the whole dataset.
  4. Report. If any constraint is violated, the SHACL validation report is serialized as pretty-printed Turtle (to stdout, or to --report <file>). With --report, successful validation also writes a report with sh:conforms true, replacing any previous failure report at that path.

Validation is read-only: it never modifies your input file.

Where the shapes come from

Ontologies and SHACL shapes are fetched from meta.linked.archi at runtime. Nothing is bundled in the converter JARs: meta.linked.archi is the single source of truth, so a corrected shape takes effect without rebuilding or re-releasing a converter.

List everything that is available, and see which documents the converter you are running uses by default:

bpmn2linkedarchi validate --list-assets
Published assets on https://meta.linked.archi (fetched at runtime):
  core                   https://meta.linked.archi/core [default ontology]
                         Linked.Archi core ontology
  core-shapes            https://meta.linked.archi/core-shapes
                         Linked.Archi core SHACL shapes
  ...
Fetched files are stored under: /var/folders/.../T/linked-archi-assets/bpmn

Default shape set per converter

Each converter validates against the shapes for the notation it emits, plus core-shapes for the contracts that are not notation-specific. PlantUML is a syntax rather than a metamodel, so its output is typed in UML; Structurizr's is typed in C4.

Converter Notation emitted Default shapes Extra names accepted
archimate2linkedarchi ArchiMate 3.2 (am:) relationships, elements, core-shapes all
bpmn2linkedarchi BPMN 2.0.2 (bpmn:) bpmn-shapes, bpmn-infra-shapes, core-shapes di, all, the individual bpmn-*-shapes names
plantuml2linkedarchi UML 2.5.1 (uml:) uml-shapes, core-shapes all, any registered name
structurizr2linkedarchi C4 (c4:) c4-shapes, structurizr-shapes, core-shapes all, any registered name
backstage2linkedarchi Backstage (bs:) backstage-shapes, core-shapes all
leanix2linkedarchi LeanIX Meta Model v4 (lmm:) leanix-shapes, core-shapes all, attributes

LeanIX loads four ontologies, and two of them are easy to miss

The default ontology set is leanix, leanix-v3, arch-processes and core. The last two are load-bearing, and the published shapes document names the set it needs in its own header:

  • leanix-v3, because LeanIXFactSheetLabelShape targets lmm3:FactSheet as well as lmm:FactSheet, even though the converter emits v4.
  • arch-processes, because the lifecycle stage values a fact sheet carries are its individuals — the LeanIX ontology reuses ap:atLifecycleStage and the five stages of ap:ApplicationLifecycle rather than minting its own. Without it, sh:in ( ap:Plan … ) has nothing to compare against.

A shape whose target class or value set the graph cannot resolve validates nothing while reporting nothing — which looks exactly like conformance. --list-assets shows the full set.

Twenty-one LeanIX shapes apply, in three groups: the naming rule; four attribute shapes covering the status, subscription-type and lifecycle-stage value sets, the completion datatype and range, and the cardinalities the ontology deliberately keeps out of OWL (nothing is an owl:FunctionalProperty, because that would infer owl:sameAs between two conflicting values and silently merge a data error instead of reporting it); and one endpoint shape per relationship.

The endpoint shapes are what make a converter's own type mapping checkable — a --type-mapping entry pointing a relation at the wrong class is caught by name ("A lmm:Provision must run from an ITComponent."), which validated clean before they existed. Note they make a dangling endpoint report twice, once from core and once naming the type that failed to arrive; that is deliberate, and the second message is the one that tells you which type to add to the pull.

--without-shape attributes turns the four attribute shapes off together, for a graph converted before that vocabulary was published.

A multi-repo Backstage run reports its cross-repo references

A single catalog conforms. A run spanning repositories does not, and the shapes are right to say so. spec.owner: team-payments in one repo mints {model}/element/group/default/team-payments inside that model's namespace, while the Group entity declared in another repo is {other-model}/element/group/default/team-payments — two IRIs for one team. bssh:OwnershipShape reports a target that is not a bs:Group, and merging the graphs does not reconcile them, because neither IRI is the other.

This is a cross-model identity question — which namespace owns an entity referenced from a repo that does not declare it — and ADR 0001 places cross-repository uniqueness with the aggregating repository, so it is not decided in the converters.

It is not a new failure. Those references already failed core-shapes#QualifiedRelationshipShape, which was in the Backstage default set before the Backstage shapes were: 6 violations before, 12 after, all one root cause. The Backstage shapes name the problem precisely (Ownership target must be a Group) instead of generically.

Each converter also loads the ontologies its shapes depend on, so rdfs:subClassOf reasoning can resolve the class hierarchy. For BPMN that is the five BPMN ontologies plus core; for ArchiMate, archimate3 plus core.

BPMN includes core-shapes in its default set because the BPMN shapes deliberately do not restate what core already covers: bpmn/onto declares SequenceFlow, MessageFlow and the other connectors to be arch:QualifiedRelationship, and core-shapes#QualifiedRelationshipShape is what then checks each one has exactly one arch:source and arch:target.

bpmn-suite is registered but not loaded by default. It is an owl:imports-only aggregator with no axioms of its own, and SHACL processors do not dereference owl:imports, so loading it would cost a fetch to add nothing. The BPMN-to-core alignment it used to carry now lives in bpmn/onto, on the classes it describes.

BPMN diagram-interchange shapes

The bpmn-di-shapes, bpmn-di-core-shapes and bpmn-dc-shapes documents constrain raw DI structures (di:bounds pointing at a dc:Bounds node, di:waypoint). By default the converter flattens geometry into archvis:bounds-*, so those shapes are not part of the BPMN default set — select them with --shapes di only for output produced with --emit-raw-di-geometry. Applying them to normalized output reports a missing di:bounds per shape, which means the shape set does not match the data.

Selecting your own shapes

--shapes replaces the converter's default set and is repeatable. Each value is one of:

  • a published asset name from --list-assets (fetched at runtime)
  • a URL (fetched at runtime)
  • a path to a local file
# A subset of the published BPMN shapes
bpmn2linkedarchi validate -i out.trig --shapes bpmn-shapes

# Two published ArchiMate sets at once
archimate2linkedarchi validate -i out.trig --shapes relationships --shapes elements

# Your own rules, from a local file
plantuml2linkedarchi validate -i out.ttl --shapes ./shapes/my-rules.ttl

# Mix published and local
bpmn2linkedarchi validate -i out.trig --shapes bpmn-shapes --shapes ./shapes/house-rules.ttl

--ontology works the same way for the reasoning ontologies, and --extra-ontology adds to the defaults rather than replacing them.

Where fetched files are stored

Fetched documents are written to a per-converter subdirectory so that converters never read each other's files or collide on a name:

<asset-root>/<converter>/<slug-derived-from-url>.ttl
  • <asset-root> defaults to linked-archi-assets inside the JVM temp directory (java.io.tmpdir), so nothing lands in your project tree. On macOS that resolves to something like /var/folders/.../T/linked-archi-assets, and on Linux /tmp.
  • <converter> is archimate, bpmn, plantuml, structurizr, backstage or leanix.
  • The file name comes from the whole URL path, not just its last segment, so /core and /core-shapes (or /archimate3/onto and /archimate3/shapes) cannot overwrite each other.

The exact directory is logged on every run and shown by --list-assets:

Assets directory: /var/folders/.../T/linked-archi-assets/bpmn

Point it somewhere else with --asset-dir. The per-converter subdirectory is always appended, so even a deliberately shared root stays isolated:

plantuml2linkedarchi validate -i out.ttl --asset-dir ./.assets
# -> ./.assets/plantuml/core-shapes.ttl

A document is downloaded once and reused on later runs. To keep nothing at all, use --no-store, which deletes the fetched files when the run finishes:

bpmn2linkedarchi validate -i out.trig --no-store

Offline and air-gapped runs

--offline never touches the network: every required document must already be present in the asset directory. If one is missing, the run fails with the exact path it expected, so you know what to seed.

# First run populates the directory
plantuml2linkedarchi validate -i out.ttl --asset-dir ./.assets

# Later runs need no network
plantuml2linkedarchi validate -i out.ttl --asset-dir ./.assets --offline
Validation error: Offline mode: required asset is not present at
./.assets/plantuml/core-shapes.ttl. Run once without --offline to fetch it, or point
--asset-dir at a populated directory.

--no-download is accepted as an alias for --offline, and --cache-dir for --asset-dir, so existing scripts keep working.

Shape coverage: avoiding a vacuous pass

A SHACL run reports "no violations" both when your data is genuinely correct and when the shapes simply did not apply to any node in your data. The second case is a vacuous pass and is misleading, so every run reports how many of the shapes' sh:targetClass declarations actually matched something:

Shape coverage: 2/4 target class(es) matched the data.

If nothing matched, the run warns explicitly instead of quietly claiming success:

Shape coverage: 0 of 156 target class(es) matched the data — no constraint was
actually checked. The shapes and the data most likely use different namespaces.
Validation reported no violations, but the result is VACUOUS (see coverage above).

Treat a vacuous result as "not validated". The usual cause is a namespace or version mismatch between the shape set and the converter output, so check that the shapes target the same namespaces the converter emits.

Exit codes

Every converter uses the same exit codes, so a validate step fails a CI job on its own.

Code Meaning
0 Data conforms to the shapes
1 SHACL violations found (report written to stdout or --report)
2 Execution error (missing input, unparseable RDF, missing cached asset)

Log output goes to stderr, so stdout carries only the report and can be redirected safely:

plantuml2linkedarchi validate -i out.ttl > report.ttl

Options

The same options are available on every converter:

Option Default Description
-i, --input (required) Input RDF file to validate
-s, --shapes converter default Shape source: published asset name, URL, or local file. Repeatable
-o, --ontology converter default Reasoning ontology source, same forms as --shapes. Repeatable
-e, --extra-ontology none Additional ontology on top of the defaults. Repeatable
--data-format inferred from extension Override input format: TURTLE, TRIG, JSONLD, RDFXML, NTRIPLES, NQUADS
-r, --report stdout Write the SHACL report to this file
--asset-dir <java.io.tmpdir>/linked-archi-assets Parent directory for fetched documents. Alias: --cache-dir
--no-store false Delete fetched documents when the run finishes
--offline false Never fetch; require documents already present. Alias: --no-download
--no-rdfs-reasoning false Disable rdfs:subClassOf reasoning
--without-shape converter default Do not enforce a specific shape. Alias or full IRI. Repeatable
--list-assets — List the published documents and exit

Excluding a rule that does not apply

Some rules in a shared shape set do not make sense for every notation. Rather than weakening the shapes, --without-shape switches off individual shapes for a run using SHACL's own sh:deactivated, leaving every other constraint enforced.

# Turn off BPMN's naming rule only
bpmn2linkedarchi validate -i out.ttl --without-shape required-names

# Same thing by full IRI
bpmn2linkedarchi validate -i out.ttl \
  --without-shape https://meta.linked.archi/bpmn/onto-shapes#RequiredNameShape

Aliases are per converter, because naming rules are owned by each notation's own shapes graph. Every converter accepts labels and element-labels for its own naming rule, so a script written against the former core aliases keeps working:

Converter Alias Disables
bpmn2linkedarchi required-names, labels, element-labels bpmn/onto-shapes#RequiredNameShape
plantuml2linkedarchi labels, element-labels uml/shapes#NamedElementShape
structurizr2linkedarchi labels, element-labels c4/shapes#C4ElementLabelShape and c4/structurizr-shapes#StructurizrElementLabelShape
archimate2linkedarchi labels, element-labels archimate3/element-shapes#ArchiMateElementShape
backstage2linkedarchi labels, element-labels backstage/shapes#BackstageElementLabelShape

Passing a name a converter does not know lists the ones it does. The SHACL report always names the shape you would need as sh:sourceShape, so a full IRI works everywhere.

Disabled shapes are logged and excluded from the coverage count, so the reported coverage still reflects only rules that were actually enforced.

The core label shapes were withdrawn

core-shapes#ElementLabelShape and #ViewLabelShape used to require a skos:prefLabel on every arch:Element and arch:View, and the aliases element-labels / view-labels / labels disabled them. Both shapes are gone: requiring a label on every element is not a sound cross-notation rule, because routing and connector constructs — BPMN gateways and events, ArchiMate junctions — are genuine elements that their notations permit to be unnamed.

Naming is now stated by each metamodel, over the classes it actually requires to be named, and every converter's default shape set includes its own notation's rule. For BPMN that is bpmn/onto-shapes#RequiredNameShape, which does not target gateways, events or sequence flows — so those no longer produce violations and there is normally nothing to switch off. No converter disables a shape by default.

There is no core arch:View label rule either, and no notation currently supplies a replacement, so a diagram's own title is not checked. bpmndi:BPMNDiagram makes it optional, which is why the core rule could not stand.

The input format is inferred from the file extension (.ttl, .trig, .jsonld, .nt, .rdf), defaulting to Turtle. Use --data-format when the extension does not match the actual serialization.

Adding or changing a published document

Shapes and ontologies live in the meta.linked.archi repository, not here. To change a rule, publish the corrected document there — converters pick it up on their next run, with no code change or release.

To register a new document so it gets a short name, add an entry to PublishedAssets in the core module and reference it from the converter's defaultShapeAssets() / defaultOntologyAssets(). That is the only place URLs are declared.

Until a document is published, validate fails with a message naming the missing URL:

Validation error: Published asset not found: https://meta.linked.archi/bpmn/onto-shapes.
It may not be published yet on https://meta.linked.archi — see
publish/meta.linked.archi/README.md.

In the meantime you can validate against local copies with --shapes ./path/to/shapes.ttl.

Reading the report

The report is a standard SHACL validation report. Each violation identifies the offending node, the property path and which constraint failed:

[] a sh:ValidationReport;
  sh:conforms false;
  sh:result [ a sh:ValidationResult;
      sh:focusNode <https://example.org/la/element/order-service>;
      sh:resultPath <https://meta.linked.archi/core#source>;
      sh:sourceConstraintComponent sh:MinCountConstraintComponent;
      sh:resultSeverity sh:Violation
    ] .
  • sh:focusNode — the resource that failed
  • sh:resultPath — the property involved
  • sh:sourceConstraintComponent — the kind of constraint (MinCount, Datatype, Class, ...)

Structural checks vs SHACL

Separate from SHACL, the core library provides ConversionVerifier — lightweight in-memory structural checks (relationships have both endpoints, view nodes reference a view, no leftover generated IRIs) that operate on a model without a SHACL engine. These are used in converter unit tests rather than exposed as a CLI command; validate is the user-facing entry point.