Skip to content

Aggregating into a Knowledge Graph

This page describes the consumer side of the pipeline: the central repository that collects pre-converted RDF from many source repos and combines it into one validated knowledge graph. It is distilled from the example-archi-graph template.

For the producer side, see Authoring Models in a Source Repo. For the whole picture, see the Pipeline Overview.


Role

The aggregation repo holds no source models — only RDF. It:

  • Declares all known RDF sources in a sources index
  • Pulls the latest converted RDF from each source repo's CI artifacts
  • Runs cross-model (Tier 2) validation on the combined graph
  • Merges everything into merged-graph.trig (and merged-graph.ttl when sources publish Turtle)
  • Publishes to a triplestore, docs site, or artifact registry

It never runs a converter. The converted RDF file (.ttl or .trig) produced by each source repo is the interface contract; the aggregation repo only downloads, validates the combination, and stores.

Two RDF formats: Turtle (.ttl) and TriG (.trig)

Sources can publish in either format, and the aggregation repo handles both.

  • Turtle (.ttl) — flat triples in a single default graph. Simplest and most widely supported; easy to read and to load into any RDF tool. But it has no concept of named graphs, so a model's semantic / views / provenance triples all sit together, and merging many .ttl files collapses everything into one graph.
  • TriG (.trig) — a superset of Turtle that adds named graphs. Each converted model keeps its graphs separate (semantic / model / views / provenance) with per-model IRIs, so provenance stays attached and many sources can be cat into one merged-graph.trig without losing that separation. Where a model was built from several input files its semantic graph is split one per input, which is the only form in which "which input asserted this fact" has an answer.

The tradeoff: Turtle is simpler and more portable but loses graph-level structure and provenance boundaries; TriG preserves per-model isolation and is safer to merge, at the cost of being slightly less universally supported. Pick Turtle when a source is a single self-contained model consumed as one blob; pick TriG when provenance separation or clean multi-source merging matters.

The example-architecture-project template publishes both by default (its convert job emits out/*.trig and out/*.ttl via the FORMATS variable). The graph pulls whichever format(s) you point its sources-index.yaml entries at, and merge.sh builds one merged file per format present:

  • Point artifact + target at .trig (recommended) → merged-graph.trig, named graphs preserved.
  • Add a second entry per source targeting the .ttl → also merged-graph.ttl, the flat union.

merge.sh never mixes formats within one file (.trig files → merged-graph.trig, .ttl files → merged-graph.ttl), so both stay individually valid.


System landscape

flowchart TB
    subgraph sources["Source repos (distributed ownership)"]
        A["archi-group/archimate-models"]
        B["archi-group/bpmn-processes"]
        C["archi-group/plantuml-diagrams"]
        D["teams/catalog-service"]
        E["teams/payments"]
    end

    subgraph graph["archi-graph (central aggregation)"]
        pull["Scheduled pull\n(GitLab Artifacts API)"]
        validate["Tier 2 validation\n(cross-model SHACL)"]
        merge["merged-graph.trig"]
    end

    subgraph consumers["Consumers"]
        sparql["SPARQL endpoint"]
        docs["rdf2docs site"]
        ai["AI agents (RAG)"]
        dev["Developers (git clone)"]
    end

    A & B & C & D & E -->|"TriG artifact"| pull
    pull --> validate --> merge
    merge --> sparql & docs & ai & dev

Repository layout

archi-graph/
├── sources-index.yaml           Lists all known TriG sources
├── graph/                       Pulled TriG, per notation (committed to git)
│   ├── archimate/  bpmn/  plantuml/  backstage/  c4/
├── shapes/
│   └── cross-model-rules.ttl    SHACL shapes for cross-model governance
├── config/
│   └── namespaces.yaml          Shared namespace prefixes
├── scripts/
│   ├── pull-sources.sh          Download TriG from indexed sources
│   ├── validate.sh              Run cross-model SHACL validation
│   └── merge.sh                 Merge graph files → merged-graph.{trig,ttl}
└── .gitlab-ci.yml               pull → validate → merge → publish

graph/ is empty until the pull runs

Like out/ in a source repo, graph/ only fills up after pull-sources.sh (or CI) downloads artifacts.


Quick start

# 1. Set your GitLab token (needs read_api on the source projects)
export GITLAB_TOKEN="glpat-..."

# 2. Pull latest TriG from all indexed sources
./scripts/pull-sources.sh

# 3. Validate the combined graph
./scripts/validate.sh

# 4. Merge into a single file per format present
./scripts/merge.sh
# → merged-graph.trig  (and merged-graph.ttl if sources publish Turtle)

Sources index

sources-index.yaml is the single coordination point. Each entry maps a source repo's CI artifact to a slot in the graph:

sources:
  - id: enterprise-archimate
    repo: archi-group/archimate-models
    job: convert                    # CI job that produces the artifact
    artifact: out/archimate.trig    # path within the artifact archive
    target: graph/archimate/enterprise-model.trig
    tier: 1
    ref: main

Adding a new source is one entry. pull-sources.sh reads this file and fetches each artifact via the GitLab Jobs Artifacts API:

GET /api/v4/projects/:id/jobs/artifacts/:ref/raw/:artifact_path?job=:job_name

GitLab resolves the latest successful pipeline on :ref automatically — no pipeline ID needed. The source repo's convert job must declare artifacts: paths: [out/] and the artifact must not be expired.


Validation tiers

Validation is split across the two repos so problems are caught as early as possible.

Tier Where When Scope
1 Source repo CI On MR Structural SHACL — element types, relationship endpoints, no _gen_ fallback IRIs, view nodes linked to a diagram
2 This repo CI On pull/merge Cross-model SHACL — every element has a skos:prefLabel, no orphan relationships, cross-notation references resolve, org governance rules

Tier 1 runs at the source and blocks the MR if the RDF is structurally invalid. By the time the RDF reaches the aggregation repo it is already structurally sound.

Tier 2 runs here, on the combined graph — it catches integration problems that only appear when models from different notations are loaded together (e.g. a Backstage Component claims system: X, but no System X exists in any graph).

Custom governance rules

Cross-model rules live in shapes/cross-model-rules.ttl:

@prefix sh: <http://www.w3.org/ns/shacl#> .

<#ComponentMustHaveOwner>
    a sh:NodeShape ;
    sh:targetClass <https://meta.linked.archi/archimate3/onto#ApplicationComponent> ;
    sh:property [
        sh:path <https://meta.linked.archi/backstage/onto#ownedBy> ;
        sh:minCount 1 ;
        sh:severity sh:Warning ;
        sh:message "ApplicationComponent has no owner relationship." ;
    ] .

End-to-end process

sequenceDiagram
    participant Dev as Developer
    participant Src as Source Repo CI
    participant API as GitLab Artifacts API
    participant AG as archi-graph CI
    participant Main as archi-graph main
    participant TS as Triplestore / rdf2docs

    Dev->>Src: push .puml / .bpmn / .yaml
    Src->>Src: convert → TriG artifact
    Src->>Src: validate (Tier 1)
    Src-->>Dev: ✓ MR mergeable
    Dev->>Src: merge to main
    Note over Src: artifact stored on main pipeline

    AG->>API: GET .../artifacts/main/raw/out/model.trig?job=convert
    API-->>AG: raw TriG bytes
    AG->>AG: git diff → changes detected
    AG->>AG: branch + MR opened
    AG->>AG: validate combined graph (Tier 2)
    AG->>Main: merge (auto for Tier 1, manual for Tier 2)
    Main->>TS: publish merged-graph.trig

The scheduled aggregation pipeline creates a branch update/graph-{timestamp}, commits the pulled TriG, and opens an MR. Nothing reaches main without passing Tier 2.


The store's default graph must be a union

A converted model is not one graph. Its facts are spread across graph/semantic — one per input file, where a model was built from several — plus graph/model for the curated part, graph/views for geometry and graph/provenance for how it was made. They compose into a model only if the store treats its default graph as the union of all named graphs.

Most stores need telling. In Fuseki it is tdb2:unionDefaultGraph true on the dataset; in GraphDB the sesame:nil / "include inferred and all graphs" setting; in rdflib Dataset(default_union=True). Without it a query with no GRAPH clause matches nothing, and the queries below silently return no rows rather than failing.

Two things make the split safe to query across:

Every concept states its model. arch:inModel travels in the same graph as the concept, so the model reassembles without depending on graph boundaries — and therefore survives a merged-graph.ttl as well as a merged-graph.trig.

SELECT ?concept WHERE { ?concept arch:inModel <…/backstage/service-catalog> }

Every graph states where it came from. Each is a prov:Bundle in that model's provenance graph.

If you already have the graph IRI — you read it out of the output, or a previous query returned it — name it directly. It is one clause, it needs no provenance, and nothing beats it:

SELECT ?s ?p ?o WHERE {
  GRAPH <…/backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml> { ?s ?p ?o }
}

The provenance route answers the question you cannot answer that way: which graph holds this file?

# everything one input produced, named by the file rather than by the graph
SELECT ?s ?p ?o WHERE {
  ?g prov:wasDerivedFrom ?src .
  ?src schema:name "catalog-info.yaml" ; dct:isPartOf <https://git.example.org/group/orders> .
  GRAPH ?g { ?s ?p ?o }
}

# which input asserted a fact — more than one answer where more than one input asserts it
SELECT ?path WHERE {
  GRAPH ?g { <…/element/component/default/order-service> bs:lifecycleState bs:Production }
  ?g prov:wasDerivedFrom ?src . ?src schema:name ?path .
}

Do not compose a graph IRI from a file path

The direct form above is fine when you have the IRI. Building one from a path is not, and it is the reason the provenance route exists:

  • The slug is lossy. catalog/checkout.yaml becomes catalog-checkout-yaml; so does catalog-checkout.yaml. You cannot invert it.
  • The shape depends on the run. With a repository known the graph is …/graph/semantic/group-orders/catalog-info-yaml; without one it is …/graph/semantic/catalog-info-yaml. The same file gives two different IRIs depending on whether CI_PROJECT_URL was set.
  • A single-input conversion is not partitioned at all — its facts are in plain …/graph/semantic.
  • ADR 0008 §3 declares these IRIs opaque, which is what makes all of the above safe to change.

In an aggregate, pin dct:isPartOf alongside schema:name as the query above does. Paths repeat across repositories — every Backstage descriptor is catalog-info.yaml — so the name alone will match graphs from every repository that has one.

A re-pull replaces a graph, it does not accumulate one. A per-source graph is named after the file, not after the file at a commit, so a descriptor whose commit moved lands in the same graph as before. The prov:Entity describing that file does carry the commit, so the revisions accumulate there — which is where the history belongs. Loading two runs of the same source therefore gives one set of graphs and one current value per property, alongside one source entity per revision read.


Versioning & overwrites

Each pull overwrites the target file in place. pull-sources.sh downloads every source to its fixed target path (curl -o), so graph/.../model.trig always holds the latest successful pipeline output from the configured ref. There is no version suffix and no accumulation of old files.

Git is the version history — not the filenames. Because the TriG is committed:

  • git diff on graph/ shows exactly which triples changed on each pull
  • git log / git blame trace when a change landed and which MR introduced it
  • reverting a bad graph state is a normal git revert

Overwriting is safe because IRIs are stable and deterministic — they are minted from source identifiers and the model-id, never from filenames. Re-pulling the same model therefore produces a clean diff (changed triples only) rather than duplicate resources, and each model's named-graph IRIs are unique, so merge.sh never collides across sources.

On the source side: artifacts are regenerated, models are versioned

The RDF artifact is a build output, not a versioned asset. In the source repo out/ is gitignored, and the convert job rewrites it from scratch on every pipeline. So the artifact itself is disposable — the real version history lives one level up:

  • The source models are the versioned truth. The .puml / .bpmn / catalog-info.yaml files are committed to the source repo's git. Their history (who changed the model, when, in which MR) is the authoritative record. The artifact is just a deterministic projection of whatever model sits on a given ref.
  • Every pipeline regenerates the artifact. On each MR and each main commit the convert job produces a fresh out/*.trig, overwriting the previous run. GitLab retains one artifact set per pipeline (bounded by expire_in), so historical versions exist only as long as their pipeline's artifact is retained.
  • "Latest" is resolved by GitLab. The pull hits /jobs/artifacts/{ref}/raw/...?job=convert, which returns the artifact from the latest successful pipeline on that ref. You never address a pipeline ID — moving the ref (e.g. a new commit on main) automatically changes what gets pulled next.

Two consequences worth designing for:

  • Reproducibility depends on the converter version, not just the model. The same model converted with a different converter image can yield different triples. Pin CONVERTER_IMAGE to a fixed tag (not :latest) in the source repo if you need a regenerated artifact to match a past one byte-for-byte.
  • Pulling by tag needs the tagged artifact to still exist. ref: v2.3.0 only works if the pipeline that built the v2.3.0 tag still has its artifact (use expire_in: never on tag/main pipelines). If the artifact expired, re-run the tagged pipeline to regenerate it — deterministic IRIs mean the output is equivalent.

Pinning a specific version

To freeze one source at a released version, set its ref to a tag (or branch) instead of main:

  - id: enterprise-archimate
    repo: archi-group/archimate-models
    job: convert
    artifact: out/archimate.trig
    target: graph/archimate/enterprise-model.trig
    tier: 1
    ref: v2.3.0        # pull from this tag instead of main

Keeping multiple versions live at once

Overwrite means only one version of a given source lives in the graph at a time. If you need v1 and v2 present simultaneously (e.g. to model a migration), give them distinct identities on two axes:

  1. Distinct index entries — separate id + target, each pulling from its own ref:

      - id: order-bpmn-v1
        repo: archi-group/bpmn-processes
        job: convert-bpmn
        artifact: out/bpmn.trig
        target: graph/bpmn/order-v1.trig
        ref: v1.2.0
      - id: order-bpmn-v2
        repo: archi-group/bpmn-processes
        job: convert-bpmn
        artifact: out/bpmn.trig
        target: graph/bpmn/order-v2.trig
        ref: main
    
  2. Distinct IRIs at conversion time — the two versions must be converted with a different model-id (or base IRI) in the source repo, so their resource and named-graph IRIs don't overwrite each other when merged. Two files pulled with the same model-id would emit the same graph IRIs and the second would shadow the first in the merged output.

In short: versioning across time is free via git; versioning side-by-side requires separate identities in both the index and the minted IRIs.


Why this design

  • Convert at source — the team that owns a model validates it, gets fast MR feedback, and conversion compute is distributed instead of bottlenecked centrally.
  • The TriG artifact is the contract — source repos can use any tool or language, as long as the output conforms to the Linked.Archi foundational ontology.
  • The graph IS the repository — committed TriG means git clone gives the full graph (no triplestore required), git diff shows what changed, and git blame traces who introduced which triples.
  • Merge via MR (never direct push) — every change to the graph is reviewable, validated, and auditable; auto-merge is allowed for trusted Tier 1 sources.
  • Scheduled pulls — simpler and more resilient than per-repo webhooks; naturally batches one MR per run. A downstream trigger can be added for latency-sensitive Tier 1 sources.

This scales to hundreds of source repos: each is independent, sources-index.yaml is the only coordination point, and the merge step is O(files) concatenation.