Skip to content

Architecture-as-Code Pipeline Overview

This document ties together the two ends of the Linked.Archi pipeline: the repos that produce RDF from architecture models, and the repo that aggregates that RDF into a single validated knowledge graph.

It is the end-to-end narrative that ties the converters into a full architecture-as-code pipeline. The converters documented in this site are the engine in the middle; this page shows where they sit and what feeds and consumes them.

The two reference implementations are:


The big picture

flowchart TB
    subgraph producers["Producer repos (example-architecture-project, team repos)"]
        models["models/<br/>BPMN, PlantUML, Backstage,<br/>C4, ArchiMate"]
        convert["convert-*.sh<br/>(Linked.Archi converters)"]
        out["out/*.trig<br/>(per-notation RDF,<br/>published as CI artifacts)"]
        models --> convert --> out
    end

    subgraph aggregation["Aggregation repo (example-archi-graph)"]
        graphdir["graph/**/*.trig<br/>(pulled RDF)"]
        validate["validate.sh<br/>(SHACL cross-model)"]
        merged["merge.sh<br/>→ merged-graph.trig"]
        graphdir --> validate --> merged
    end

    subgraph consumers["Consumers"]
        ts["Triplestore"]
        docs["rdf2docs"]
        sparql["SPARQL"]
        ai["AI"]
    end

    out -->|"pull-sources.sh reads<br/>sources-index.yaml"| graphdir
    merged --> ts & docs & sparql & ai

The core idea: many notations, converted independently, merged into one graph. RDF (specifically TriG with named graphs) is the common interchange format that makes the two sides interoperable without either knowing about the other's tooling.


Stage 1 — Produce RDF (example-architecture-project)

Role: A template for an engineering team that authors architecture models as code and converts them to RDF on every push.

Authored input lives in models/, one directory per notation:

Directory Notation Sample content
models/bpmn/ BPMN 2.0 order-fulfillment.bpmn + diagram-index.yaml
models/plantuml/ PlantUML inventory-domain.puml, checkout-sequence.puml + index
models/backstage/ Backstage catalog catalog-info.yaml + catalog-index.yaml
models/structurizr/ Structurizr (C4) workspace.json
models/archimate/ ArchiMate Exchange XML (placeholder — empty in the example)

Generated output: per-notation RDF in TriG format, written to out/ (e.g. out/bpmn.trig, out/plantuml.trig). Produced by the convert-*.sh scripts, which run the Linked.Archi converter Docker image, or by GitLab CI on every push.

out/ is gitignored. In a fresh checkout it is empty — the RDF only appears after running ./convert-all.sh locally or after a CI run.

Configuration knobs: a shared BASE_IRI, optional --type-mapping overrides (e.g. config/type-mapping-bpmn-lite.yml), and per-notation diagram-index.yaml files that control which files are processed, their stable model IDs, and publish status.

Other team repos (payments C4, catalog Backstage, etc.) play the same role — they are all producers that emit .trig artifacts.


Stage 2 — Aggregate RDF (example-archi-graph)

Role: The central repo that collects pre-converted RDF from many producer repos and combines it into one validated knowledge graph. It holds no source models — only RDF.

How sources are declared: sources-index.yaml lists every producer repo, the CI job and artifact path that produces its .trig, the local target path to store it, and a validation tier.

Landing zone: pulled .trig files land in graph/, organised by notation (archimate/, bpmn/, plantuml/, backstage/, c4/).

Like out/ upstream, graph/ is empty until pull-sources.sh (or CI) runs.

Scripts:

  • scripts/pull-sources.sh — downloads the latest .trig artifact from each indexed source
  • scripts/validate.sh — runs SHACL cross-model validation (shapes/cross-model-rules.ttl)
  • scripts/merge.sh — merges into merged-graph.trig (and merged-graph.ttl when sources publish Turtle)

Generated output: merged-graph.trig (and optionally merged-graph.ttl) — the whole organisation's architecture as one graph, ready to publish to a triplestore, feed to rdf2docs, or query via SPARQL.


The connective tissue

Three things make the two sides fit together:

  1. TriG with named graphs — the shared interchange format. Each producer emits semantic / views / provenance named graphs with per-model IRIs; the aggregator can concatenate them without collisions.
  2. sources-index.yaml — the contract that maps a producer's CI artifact to a slot in the aggregated graph. Adding a new producer is a matter of adding one entry.
  3. Two-tier validation — Tier 1 (structural SHACL) runs in the producer repo on every MR; Tier 2 (cross-model integration constraints) runs in the aggregation repo on pull/merge. Nothing reaches main without passing its tier.

Where the documentation lives

Topic Location
End-to-end pipeline (this doc) workflow/overview.md
Self-hosting: build the image in-house & reuse it workflow/self-hosting.md
Authoring models in a source repo workflow/source-project.md
Aggregation flow, governance, validation tiers workflow/aggregation-graph.md
Per-notation converter usage converters/*.md
Type mapping & ontology alignment config/type-mapping.md, config/ontology-alignment.md
Docker & CI usage ci-usage.md
Building a new converter development/converter-development-guide.md
Converter approach trade-offs development/converter-alternatives.md