Validation¶
Every converter ships a validate subcommand that runs SHACL validation over RDF
that a converter has already produced. Validation is a separate step from convert:
you convert first, then validate the output file.
bpmn2linkedarchi convert process.bpmn -o out.ttl --format TURTLE
bpmn2linkedarchi validate -i out.ttl
What actually happens¶
All converters share one validation engine (archi.linked.converter.core.validation.ShaclValidator).
A run performs these steps:
- Load the shape graph. The SHACL shapes are loaded into RDF4J's dedicated
shapes graph (
rdf4j:SHACLShapeGraph) of an in-memoryShaclSail. - Load supporting ontologies. Class/property definitions are loaded into the default
graph so that
rdfs:subClassOfreasoning can resolve type hierarchies. Without this, a shape targeting a superclass would not apply to instances of its subclasses. - Load the data and validate. Your input file is added inside a transaction using
the
Bulkvalidation approach; committing triggers validation of the whole dataset. - Report. If any constraint is violated, the SHACL validation report is serialized
as pretty-printed Turtle (to stdout, or to
--report <file>). With--report, successful validation also writes a report withsh:conforms true, replacing any previous failure report at that path.
Validation is read-only: it never modifies your input file.
Where the shapes come from¶
Ontologies and SHACL shapes are fetched from meta.linked.archi at runtime. Nothing
is bundled in the converter JARs: meta.linked.archi is the single source of truth, so a
corrected shape takes effect without rebuilding or re-releasing a converter.
List everything that is available, and see which documents the converter you are running uses by default:
Published assets on https://meta.linked.archi (fetched at runtime):
core https://meta.linked.archi/core [default ontology]
Linked.Archi core ontology
core-shapes https://meta.linked.archi/core-shapes
Linked.Archi core SHACL shapes
...
Fetched files are stored under: /var/folders/.../T/linked-archi-assets/bpmn
Default shape set per converter¶
Each converter validates against the shapes for the notation it emits, plus core-shapes
for the contracts that are not notation-specific. PlantUML is a syntax rather than a
metamodel, so its output is typed in UML; Structurizr's is typed in C4.
| Converter | Notation emitted | Default shapes | Extra names accepted |
|---|---|---|---|
archimate2linkedarchi |
ArchiMate 3.2 (am:) |
relationships, elements, core-shapes |
all |
bpmn2linkedarchi |
BPMN 2.0.2 (bpmn:) |
bpmn-shapes, bpmn-infra-shapes, core-shapes |
di, all, the individual bpmn-*-shapes names |
plantuml2linkedarchi |
UML 2.5.1 (uml:) |
uml-shapes, core-shapes |
all, any registered name |
structurizr2linkedarchi |
C4 (c4:) |
c4-shapes, structurizr-shapes, core-shapes |
all, any registered name |
backstage2linkedarchi |
Backstage (bs:) |
backstage-shapes, core-shapes |
all |
leanix2linkedarchi |
LeanIX Meta Model v4 (lmm:) |
leanix-shapes, core-shapes |
all, attributes |
LeanIX loads four ontologies, and two of them are easy to miss
The default ontology set is leanix, leanix-v3, arch-processes and core. The last two are
load-bearing, and the published shapes document names the set it needs in its own header:
leanix-v3, becauseLeanIXFactSheetLabelShapetargetslmm3:FactSheetas well aslmm:FactSheet, even though the converter emits v4.arch-processes, because the lifecycle stage values a fact sheet carries are its individuals — the LeanIX ontology reusesap:atLifecycleStageand the five stages ofap:ApplicationLifecyclerather than minting its own. Without it,sh:in ( ap:Plan … )has nothing to compare against.
A shape whose target class or value set the graph cannot resolve validates nothing while reporting
nothing — which looks exactly like conformance. --list-assets shows the full set.
Twenty-one LeanIX shapes apply, in three groups: the naming rule; four attribute shapes covering
the status, subscription-type and lifecycle-stage value sets, the completion datatype and range,
and the cardinalities the ontology deliberately keeps out of OWL (nothing is an
owl:FunctionalProperty, because that would infer owl:sameAs between two conflicting values and
silently merge a data error instead of reporting it); and one endpoint shape per relationship.
The endpoint shapes are what make a converter's own type mapping checkable — a --type-mapping
entry pointing a relation at the wrong class is caught by name ("A lmm:Provision must run from an
ITComponent."), which validated clean before they existed. Note they make a dangling endpoint
report twice, once from core and once naming the type that failed to arrive; that is
deliberate, and the second message is the one that tells you which type to add to the pull.
--without-shape attributes turns the four attribute shapes off together, for a graph converted
before that vocabulary was published.
A multi-repo Backstage run reports its cross-repo references
A single catalog conforms. A run spanning repositories does not, and the shapes are right to
say so. spec.owner: team-payments in one repo mints
{model}/element/group/default/team-payments inside that model's namespace, while the
Group entity declared in another repo is {other-model}/element/group/default/team-payments
— two IRIs for one team. bssh:OwnershipShape reports a target that is not a bs:Group, and
merging the graphs does not reconcile them, because neither IRI is the other.
This is a cross-model identity question — which namespace owns an entity referenced from a repo that does not declare it — and ADR 0001 places cross-repository uniqueness with the aggregating repository, so it is not decided in the converters.
It is not a new failure. Those references already failed
core-shapes#QualifiedRelationshipShape, which was in the Backstage default set before the
Backstage shapes were: 6 violations before, 12 after, all one root cause. The Backstage
shapes name the problem precisely (Ownership target must be a Group) instead of generically.
Each converter also loads the ontologies its shapes depend on, so rdfs:subClassOf
reasoning can resolve the class hierarchy. For BPMN that is the five BPMN ontologies plus
core; for ArchiMate, archimate3 plus core.
BPMN includes core-shapes in its default set because the BPMN shapes deliberately do not
restate what core already covers: bpmn/onto declares SequenceFlow, MessageFlow and
the other connectors to be arch:QualifiedRelationship, and
core-shapes#QualifiedRelationshipShape is what then checks each one has exactly one
arch:source and arch:target.
bpmn-suite is registered but not loaded by default. It is an owl:imports-only
aggregator with no axioms of its own, and SHACL processors do not dereference
owl:imports, so loading it would cost a fetch to add nothing. The BPMN-to-core alignment
it used to carry now lives in bpmn/onto, on the classes it describes.
BPMN diagram-interchange shapes
The bpmn-di-shapes, bpmn-di-core-shapes and bpmn-dc-shapes documents constrain
raw DI structures (di:bounds pointing at a dc:Bounds node, di:waypoint).
By default the converter flattens geometry into archvis:bounds-*, so those shapes are
not part of the BPMN default set — select them with --shapes di only for output
produced with --emit-raw-di-geometry. Applying them to normalized output reports a
missing di:bounds per shape, which means the shape set does not match the data.
Selecting your own shapes¶
--shapes replaces the converter's default set and is repeatable. Each value is one of:
- a published asset name from
--list-assets(fetched at runtime) - a URL (fetched at runtime)
- a path to a local file
# A subset of the published BPMN shapes
bpmn2linkedarchi validate -i out.trig --shapes bpmn-shapes
# Two published ArchiMate sets at once
archimate2linkedarchi validate -i out.trig --shapes relationships --shapes elements
# Your own rules, from a local file
plantuml2linkedarchi validate -i out.ttl --shapes ./shapes/my-rules.ttl
# Mix published and local
bpmn2linkedarchi validate -i out.trig --shapes bpmn-shapes --shapes ./shapes/house-rules.ttl
--ontology works the same way for the reasoning ontologies, and --extra-ontology adds
to the defaults rather than replacing them.
Where fetched files are stored¶
Fetched documents are written to a per-converter subdirectory so that converters never read each other's files or collide on a name:
<asset-root>defaults tolinked-archi-assetsinside the JVM temp directory (java.io.tmpdir), so nothing lands in your project tree. On macOS that resolves to something like/var/folders/.../T/linked-archi-assets, and on Linux/tmp.<converter>isarchimate,bpmn,plantuml,structurizr,backstageorleanix.- The file name comes from the whole URL path, not just its last segment, so
/coreand/core-shapes(or/archimate3/ontoand/archimate3/shapes) cannot overwrite each other.
The exact directory is logged on every run and shown by --list-assets:
Point it somewhere else with --asset-dir. The per-converter subdirectory is always
appended, so even a deliberately shared root stays isolated:
plantuml2linkedarchi validate -i out.ttl --asset-dir ./.assets
# -> ./.assets/plantuml/core-shapes.ttl
A document is downloaded once and reused on later runs. To keep nothing at all, use
--no-store, which deletes the fetched files when the run finishes:
Offline and air-gapped runs¶
--offline never touches the network: every required document must already be present in
the asset directory. If one is missing, the run fails with the exact path it expected, so
you know what to seed.
# First run populates the directory
plantuml2linkedarchi validate -i out.ttl --asset-dir ./.assets
# Later runs need no network
plantuml2linkedarchi validate -i out.ttl --asset-dir ./.assets --offline
Validation error: Offline mode: required asset is not present at
./.assets/plantuml/core-shapes.ttl. Run once without --offline to fetch it, or point
--asset-dir at a populated directory.
--no-download is accepted as an alias for --offline, and --cache-dir for
--asset-dir, so existing scripts keep working.
Shape coverage: avoiding a vacuous pass¶
A SHACL run reports "no violations" both when your data is genuinely correct and when
the shapes simply did not apply to any node in your data. The second case is a vacuous
pass and is misleading, so every run reports how many of the shapes' sh:targetClass
declarations actually matched something:
If nothing matched, the run warns explicitly instead of quietly claiming success:
Shape coverage: 0 of 156 target class(es) matched the data — no constraint was
actually checked. The shapes and the data most likely use different namespaces.
Validation reported no violations, but the result is VACUOUS (see coverage above).
Treat a vacuous result as "not validated". The usual cause is a namespace or version mismatch between the shape set and the converter output, so check that the shapes target the same namespaces the converter emits.
Exit codes¶
Every converter uses the same exit codes, so a validate step fails a CI job on its own.
| Code | Meaning |
|---|---|
0 |
Data conforms to the shapes |
1 |
SHACL violations found (report written to stdout or --report) |
2 |
Execution error (missing input, unparseable RDF, missing cached asset) |
Log output goes to stderr, so stdout carries only the report and can be redirected safely:
Options¶
The same options are available on every converter:
| Option | Default | Description |
|---|---|---|
-i, --input |
(required) | Input RDF file to validate |
-s, --shapes |
converter default | Shape source: published asset name, URL, or local file. Repeatable |
-o, --ontology |
converter default | Reasoning ontology source, same forms as --shapes. Repeatable |
-e, --extra-ontology |
none | Additional ontology on top of the defaults. Repeatable |
--data-format |
inferred from extension | Override input format: TURTLE, TRIG, JSONLD, RDFXML, NTRIPLES, NQUADS |
-r, --report |
stdout | Write the SHACL report to this file |
--asset-dir |
<java.io.tmpdir>/linked-archi-assets |
Parent directory for fetched documents. Alias: --cache-dir |
--no-store |
false |
Delete fetched documents when the run finishes |
--offline |
false |
Never fetch; require documents already present. Alias: --no-download |
--no-rdfs-reasoning |
false |
Disable rdfs:subClassOf reasoning |
--without-shape |
converter default | Do not enforce a specific shape. Alias or full IRI. Repeatable |
--list-assets |
— | List the published documents and exit |
Excluding a rule that does not apply¶
Some rules in a shared shape set do not make sense for every notation. Rather than
weakening the shapes, --without-shape switches off individual shapes for a run using
SHACL's own sh:deactivated, leaving every other constraint enforced.
# Turn off BPMN's naming rule only
bpmn2linkedarchi validate -i out.ttl --without-shape required-names
# Same thing by full IRI
bpmn2linkedarchi validate -i out.ttl \
--without-shape https://meta.linked.archi/bpmn/onto-shapes#RequiredNameShape
Aliases are per converter, because naming rules are owned by each notation's own shapes
graph. Every converter accepts labels and element-labels for its own naming rule, so a
script written against the former core aliases keeps working:
| Converter | Alias | Disables |
|---|---|---|
bpmn2linkedarchi |
required-names, labels, element-labels |
bpmn/onto-shapes#RequiredNameShape |
plantuml2linkedarchi |
labels, element-labels |
uml/shapes#NamedElementShape |
structurizr2linkedarchi |
labels, element-labels |
c4/shapes#C4ElementLabelShape and c4/structurizr-shapes#StructurizrElementLabelShape |
archimate2linkedarchi |
labels, element-labels |
archimate3/element-shapes#ArchiMateElementShape |
backstage2linkedarchi |
labels, element-labels |
backstage/shapes#BackstageElementLabelShape |
Passing a name a converter does not know lists the ones it does. The SHACL report always
names the shape you would need as sh:sourceShape, so a full IRI works everywhere.
Disabled shapes are logged and excluded from the coverage count, so the reported coverage still reflects only rules that were actually enforced.
The core label shapes were withdrawn
core-shapes#ElementLabelShape and #ViewLabelShape used to require a
skos:prefLabel on every arch:Element and arch:View, and the aliases
element-labels / view-labels / labels disabled them. Both shapes are gone:
requiring a label on every element is not a sound cross-notation rule, because routing
and connector constructs — BPMN gateways and events, ArchiMate junctions — are genuine
elements that their notations permit to be unnamed.
Naming is now stated by each metamodel, over the classes it actually requires to be
named, and every converter's default shape set includes its own notation's rule. For
BPMN that is bpmn/onto-shapes#RequiredNameShape, which does not target gateways,
events or sequence flows — so those no longer produce violations and there is normally
nothing to switch off. No converter disables a shape by default.
There is no core arch:View label rule either, and no notation currently supplies a
replacement, so a diagram's own title is not checked. bpmndi:BPMNDiagram makes it
optional, which is why the core rule could not stand.
The input format is inferred from the file extension (.ttl, .trig, .jsonld, .nt,
.rdf), defaulting to Turtle. Use --data-format when the extension does not match the
actual serialization.
Adding or changing a published document¶
Shapes and ontologies live in the meta.linked.archi repository, not here. To change a
rule, publish the corrected document there — converters pick it up on their next run, with
no code change or release.
To register a new document so it gets a short name, add an entry to
PublishedAssets in the core module and reference it from the converter's
defaultShapeAssets() / defaultOntologyAssets(). That is the only place URLs are
declared.
Until a document is published, validate fails with a message naming the missing URL:
Validation error: Published asset not found: https://meta.linked.archi/bpmn/onto-shapes.
It may not be published yet on https://meta.linked.archi — see
publish/meta.linked.archi/README.md.
In the meantime you can validate against local copies with --shapes ./path/to/shapes.ttl.
Reading the report¶
The report is a standard SHACL validation report. Each violation identifies the offending node, the property path and which constraint failed:
[] a sh:ValidationReport;
sh:conforms false;
sh:result [ a sh:ValidationResult;
sh:focusNode <https://example.org/la/element/order-service>;
sh:resultPath <https://meta.linked.archi/core#source>;
sh:sourceConstraintComponent sh:MinCountConstraintComponent;
sh:resultSeverity sh:Violation
] .
sh:focusNode— the resource that failedsh:resultPath— the property involvedsh:sourceConstraintComponent— the kind of constraint (MinCount,Datatype,Class, ...)
Structural checks vs SHACL¶
Separate from SHACL, the core library provides ConversionVerifier — lightweight
in-memory structural checks (relationships have both endpoints, view nodes reference a
view, no leftover generated IRIs) that operate on a model without a SHACL engine. These
are used in converter unit tests rather than exposed as a CLI command; validate is the
user-facing entry point.