Docker & CI Usage¶
Adopting this in-house?
For the full flow — hosting the converters on your GitLab, building the image into your own registry, and reusing it across architecture projects and the architecture graph — see Self-Hosting the Converters. This page focuses on the per-notation CI job snippets.
Docker image¶
The project builds a single Docker image containing all converters:
# Build locally. Fine as-is for development; the image just cannot name itself in the
# provenance of the graphs it converts. See "Recording provenance" below to build one that can.
docker build -t linkedarchi-converters .
# List available converters
docker run --rm linkedarchi-converters
# Run a converter
docker run --rm -v "$PWD:/work" -w /work linkedarchi-converters \
bpmn2linkedarchi convert process.bpmn \
--base-iri https://example.org/la/ --format TRIG -o out.trig
Available commands in the image:
| Command | Converter |
|---|---|
archimate2linkedarchi |
ArchiMate Exchange XML → RDF |
bpmn2linkedarchi |
BPMN 2.0 XML → RDF |
plantuml2linkedarchi |
PlantUML diagrams → RDF |
structurizr2linkedarchi |
Structurizr workspace JSON → RDF |
backstage2linkedarchi |
Backstage catalog YAML → RDF |
leanix2linkedarchi |
LeanIX fact sheet export → RDF, plus --diagrams-export for the diagrams, which become arch:View. mapping-report --fail-on-unmapped gates a pipeline on the workspace's types and relations all reaching a published class |
leanix-pull |
LeanIX workspace → fact sheet export over GraphQL; introspect writes a pull configuration from the tenant's own schema. Not a converter — it produces no RDF, and runs as its own job so that fetching and converting fail, and are scheduled, independently. See Pulling a workspace |
Using in GitLab CI¶
Publishing both TriG and Turtle
The snippets below emit a single format for clarity. To publish both (as the
example-architecture-project template does), loop over the formats and map each to
its extension:
variables:
FORMATS: "TRIG TURTLE"
script:
- |
for fmt in $FORMATS; do
[ "$fmt" = "TURTLE" ] && ext=ttl || ext=trig
bpmn2linkedarchi convert models/*.bpmn \
--base-iri $BASE_IRI --format $fmt --include-di \
-o out/bpmn.$ext
done
TriG keeps named graphs; Turtle is the flat union. See the format tradeoff.
Convert BPMN in a downstream pipeline¶
# .gitlab-ci.yml in your BPMN repository
convert-to-rdf:
stage: build
image: $CI_REGISTRY/linked-archi/linked-archi-tools/converters/converters:latest
script:
- bpmn2linkedarchi convert
models/*.bpmn
--base-iri https://archi.example.com/graph/
--format TRIG
--include-di
--diagrams-root models/
--diagrams-index models/diagram-index.yaml
-o out/bpmn-models.trig
artifacts:
paths:
- out/*.trig
Convert Backstage catalog (multi-repo collection)¶
# .gitlab-ci.yml in a central catalog aggregation repo
convert-catalog:
stage: build
image: $CI_REGISTRY/linked-archi/linked-archi-tools/converters/converters:latest
script:
- backstage2linkedarchi convert
collected-catalogs/*.yaml
--base-iri https://archi.example.com/graph/
--format TRIG
--diagrams-root collected-catalogs/
--diagrams-index collected-catalogs/catalog-index.yaml
-o out/backstage-catalog.trig
artifacts:
paths:
- out/*.trig
Convert PlantUML diagrams from docs repo¶
# .gitlab-ci.yml in your architecture docs repository
convert-diagrams:
stage: build
image: $CI_REGISTRY/linked-archi/linked-archi-tools/converters/converters:latest
script:
- plantuml2linkedarchi convert
diagrams/*.puml
--base-iri https://archi.example.com/graph/
--format TRIG
--diagrams-root diagrams/
--diagrams-index diagrams/diagram-index.yaml
-o out/uml-diagrams.trig
artifacts:
paths:
- out/*.trig
Convert Structurizr workspace¶
convert-c4:
stage: build
image: $CI_REGISTRY/linked-archi/linked-archi-tools/converters/converters:latest
script:
- structurizr2linkedarchi convert
workspace.json
--base-iri https://archi.example.com/graph/
--model-id my-system
--format TRIG
-o out/c4-model.trig
artifacts:
paths:
- out/*.trig
Full pipeline: convert + validate + publish¶
stages:
- convert
- validate
- publish
convert:
stage: convert
image: $CI_REGISTRY/linked-archi/linked-archi-tools/converters/converters:latest
script:
- bpmn2linkedarchi convert
models/*.bpmn
--base-iri https://archi.example.com/graph/
--format TRIG
--include-di
-o out/models.trig
artifacts:
paths:
- out/
validate:
stage: validate
image: $CI_REGISTRY/linked-archi/linked-archi-tools/converters/converters:latest
script:
# Exits non-zero on violations, so the job fails on its own.
- archimate2linkedarchi validate
--input out/models.trig
--shapes relationships
--report shacl-report.ttl
artifacts:
when: always
paths:
- shacl-report.ttl
needs: [convert]
publish:
stage: publish
image: curlimages/curl:latest
script:
- curl -X PUT
-H "Content-Type: application/trig"
--data-binary @out/models.trig
https://triplestore.example.com/repositories/architecture/statements
needs: [validate]
only:
- main
Validation in CI¶
Every converter uses the same exit codes (0 conforms, 1 violations, 2 error), so a
validate step fails the job without extra shell guards:
Logs go to stderr and the report to stdout, so you can redirect the report instead of
using --report:
Shapes are fetched at runtime¶
Every converter downloads its ontologies and SHACL shapes from meta.linked.archi on
each run, so the job needs outbound network access unless you seed them first.
To keep the job off the network, point --asset-dir at a cached directory and pass
--offline once it is warm. A per-converter subdirectory is created inside it, so one
cache entry serves all converters without them sharing files:
script:
# Warm on the first pipeline, then reuse
- plantuml2linkedarchi validate -i out/models.ttl --asset-dir .assets
- bpmn2linkedarchi validate -i out/models.trig --asset-dir .assets --offline
cache:
key: linked-archi-assets
paths:
- .assets
If a document is missing in --offline mode the job fails with the exact path it
expected. Without --asset-dir, fetched files go to a temp directory
(<java.io.tmpdir>/linked-archi-assets/<converter>) and are not preserved between jobs.
Add --no-store if you want them removed as soon as the run ends.
Check the reported shape coverage line in the job log. A 0 of N target class(es)
matched warning means nothing was actually validated — see
Validation.
Image versioning¶
The CI pipeline pushes two tags:
| Tag | When |
|---|---|
$CI_REGISTRY_IMAGE:$CI_COMMIT_SHORT_SHA |
Every push to main |
$CI_REGISTRY_IMAGE:latest |
Every push to main (overwritten) |
For pinned versions, push a Git tag — it will also trigger the docker job.
Recording provenance¶
Every conversion records how it was made: a prov:Activity for the run and a prov:SoftwareAgent
for the image that performed it. Two things make that record precise, and both are the caller's to
supply. See ADR 0008.
Build an image that can name itself¶
The pipeline passes four build args, and any pipeline building this image should:
docker build \
--build-arg "IMAGE_REF=$IMAGE_TAG" \
--build-arg "SOURCE_REVISION=$CI_COMMIT_SHA" \
--build-arg "IMAGE_VERSION=${CI_COMMIT_TAG:-$CI_COMMIT_SHORT_SHA}" \
--build-arg "SOURCE_URL=$CI_PROJECT_URL" \
-t "$IMAGE_TAG" -t "$IMAGE_LATEST" .
All four are optional and the build succeeds without them, so nothing breaks for an existing custom image — but its graphs will name no image and no source revision, and nothing warns about it. A converter that cannot identify itself says nothing rather than guessing.
SOURCE_URL has no default on purpose: pairing your revision with another project's URL would
publish a commit link that resolves to nothing. Self-hosted builds should use
.gitlab-ci.custom-registry.yml, which passes all four.
When you build from downloaded JARs, not from source¶
A pipeline building the image out of this repository has CI_COMMIT_SHA and CI_PROJECT_URL to hand.
One that downloads published JARs and packages them itself does not: its own CI_COMMIT_SHA is its
commit, not the converter's, and baking that would attribute the graph to code that never ran.
So both values are published next to the artifacts — SOURCE_REVISION and SOURCE_URL, one line
each, in the Pages downloads directory and under every Package
Registry version path:
BASE_URL=https://linked-archi.gitlab.io/linked-archi-tools/converters/converters/downloads
docker build \
--build-arg "IMAGE_REF=$IMAGE_TAG" \
--build-arg "IMAGE_VERSION=$(curl -fsSL "$BASE_URL/VERSION")" \
--build-arg "SOURCE_REVISION=$(curl -fsSL "$BASE_URL/SOURCE_REVISION")" \
--build-arg "SOURCE_URL=$(curl -fsSL "$BASE_URL/SOURCE_URL")" \
-t "$IMAGE_TAG" .
Do not substitute a version for a revision. Default-branch builds all publish the same -SNAPSHOT
version, so it does not identify a build, and LA_SOURCE_REVISION is read as a bare sha to compose a
commit URL from — a version there produces a link to nothing. Passing neither is the correct fallback
when they are genuinely unavailable: the image then makes no source-commit claim, which is what
"degrade honestly" means here.
Pin the image by digest¶
A container cannot discover its own tag or digest, so it reports what it was told. Running :latest
means the graph can only name a moving tag; pinning by digest names the exact build:
convert:
# Not `:latest` — a graph produced by "whatever latest was that day" cannot be traced back to a
# build. The digest is reported as the agent's dct:identifier.
image: registry.gitlab.com/linked-archi/linked-archi-tools/converters/converters@sha256:89ab…
script:
- bpmn2linkedarchi convert process.bpmn --base-iri https://example.org/la/ \
--format TRIG -o out.trig --git-provenance auto
GitLab exposes the reference the job asked for as CI_JOB_IMAGE, which the converter reads
automatically. The converter's own source revision is recorded either way, baked in at image build
time, so even an unpinned run identifies the code behind the graph.
Ask for the source commit¶
--git-provenance auto reads the variables below and falls back to git outside CI. It is off by
default because it changes published output.
| Variable | Becomes |
|---|---|
CI_COMMIT_SHA |
the commit the source was read at, as dct:identifier |
CI_COMMIT_TIMESTAMP |
prov:endedAtTime on the commit activity |
CI_COMMIT_AUTHOR |
a prov:Person: the address as schema:email, which keys the node, and the name as schema:name. Keyed on the address so a rename does not mint a second person — see ADR 0008 §5 |
CI_PROJECT_URL |
the repository as dct:isPartOf; the commit-pinned blob as prov:alternateOf on the source and as dct:source on the model or view; a browsable commit link as schema:url on the commit activity |
CI_PROJECT_DIR |
the repository root, so schema:name is a repo-relative path |
CI_JOB_IMAGE |
the agent's dct:identifier |
CI_JOB_URL |
the run's dct:identifier, so several models from one job are correlatable |
prov:wasDerivedFrom then points at the source file at that commit, so a published IRI can be
traced to the exact bytes it came from. The source entity states the repository (dct:isPartOf), the
path within it (schema:name) and the commit (dct:identifier) as three separate facts, since a path
identifies a file only relative to a repository — and a run whose inputs are pulled from elsewhere reads
many. See Collecting a Backstage catalog.
Reproducible output¶
--run-timestamp commit takes the run's times from the commit instead of the clock, which makes the
provenance graph byte-identical for an unchanged model at an unchanged commit. Useful when generated
graphs are committed and git diff should show only real changes. Note that the rest of the file is
not yet byte-stable: folder membership uses blank nodes whose identifiers are generated per run.
Local Docker usage¶
# Build
docker build -t linkedarchi-converters .
# Convert a local file
docker run --rm -v "$PWD:/work" -w /work linkedarchi-converters \
backstage2linkedarchi convert catalog/ \
--base-iri https://example.org/la/ --format TRIG -o out.trig
# Use with type mapping (mount the config)
docker run --rm -v "$PWD:/work" -w /work linkedarchi-converters \
bpmn2linkedarchi convert process.bpmn \
--base-iri https://example.org/la/ --format TRIG \
--type-mapping /work/config/type-mapping-bpmn-lite.yml \
-o out.trig