Aggregating into a Knowledge Graph¶
This page describes the consumer side of the pipeline: the central repository that
collects pre-converted RDF from many source repos and combines it into one validated
knowledge graph. It is distilled from the example-archi-graph template.
For the producer side, see Authoring Models in a Source Repo. For the whole picture, see the Pipeline Overview.
Role¶
The aggregation repo holds no source models — only RDF. It:
- Declares all known RDF sources in a sources index
- Pulls the latest converted RDF from each source repo's CI artifacts
- Runs cross-model (Tier 2) validation on the combined graph
- Merges everything into
merged-graph.trig(andmerged-graph.ttlwhen sources publish Turtle) - Publishes to a triplestore, docs site, or artifact registry
It never runs a converter. The converted RDF file (.ttl or .trig) produced by each
source repo is the interface contract; the aggregation repo only downloads, validates
the combination, and stores.
Two RDF formats: Turtle (.ttl) and TriG (.trig)
Sources can publish in either format, and the aggregation repo handles both.
- Turtle (
.ttl) — flat triples in a single default graph. Simplest and most widely supported; easy to read and to load into any RDF tool. But it has no concept of named graphs, so a model's semantic / views / provenance triples all sit together, and merging many.ttlfiles collapses everything into one graph. - TriG (
.trig) — a superset of Turtle that adds named graphs. Each converted model keeps its graphs separate (semantic / model / views / provenance) with per-model IRIs, so provenance stays attached and many sources can becatinto onemerged-graph.trigwithout losing that separation. Where a model was built from several input files its semantic graph is split one per input, which is the only form in which "which input asserted this fact" has an answer.
The tradeoff: Turtle is simpler and more portable but loses graph-level structure and provenance boundaries; TriG preserves per-model isolation and is safer to merge, at the cost of being slightly less universally supported. Pick Turtle when a source is a single self-contained model consumed as one blob; pick TriG when provenance separation or clean multi-source merging matters.
The example-architecture-project template publishes both by default (its
convert job emits out/*.trig and out/*.ttl via the FORMATS variable). The
graph pulls whichever format(s) you point its sources-index.yaml entries at, and
merge.sh builds one merged file per format present:
- Point
artifact+targetat.trig(recommended) →merged-graph.trig, named graphs preserved. - Add a second entry per source targeting the
.ttl→ alsomerged-graph.ttl, the flat union.
merge.sh never mixes formats within one file (.trig files → merged-graph.trig,
.ttl files → merged-graph.ttl), so both stay individually valid.
System landscape¶
flowchart TB
subgraph sources["Source repos (distributed ownership)"]
A["archi-group/archimate-models"]
B["archi-group/bpmn-processes"]
C["archi-group/plantuml-diagrams"]
D["teams/catalog-service"]
E["teams/payments"]
end
subgraph graph["archi-graph (central aggregation)"]
pull["Scheduled pull\n(GitLab Artifacts API)"]
validate["Tier 2 validation\n(cross-model SHACL)"]
merge["merged-graph.trig"]
end
subgraph consumers["Consumers"]
sparql["SPARQL endpoint"]
docs["rdf2docs site"]
ai["AI agents (RAG)"]
dev["Developers (git clone)"]
end
A & B & C & D & E -->|"TriG artifact"| pull
pull --> validate --> merge
merge --> sparql & docs & ai & dev
Repository layout¶
archi-graph/
├── sources-index.yaml Lists all known TriG sources
├── graph/ Pulled TriG, per notation (committed to git)
│ ├── archimate/ bpmn/ plantuml/ backstage/ c4/
├── shapes/
│ └── cross-model-rules.ttl SHACL shapes for cross-model governance
├── config/
│ └── namespaces.yaml Shared namespace prefixes
├── scripts/
│ ├── pull-sources.sh Download TriG from indexed sources
│ ├── validate.sh Run cross-model SHACL validation
│ └── merge.sh Merge graph files → merged-graph.{trig,ttl}
└── .gitlab-ci.yml pull → validate → merge → publish
graph/ is empty until the pull runs
Like out/ in a source repo, graph/ only fills up after pull-sources.sh
(or CI) downloads artifacts.
Quick start¶
# 1. Set your GitLab token (needs read_api on the source projects)
export GITLAB_TOKEN="glpat-..."
# 2. Pull latest TriG from all indexed sources
./scripts/pull-sources.sh
# 3. Validate the combined graph
./scripts/validate.sh
# 4. Merge into a single file per format present
./scripts/merge.sh
# → merged-graph.trig (and merged-graph.ttl if sources publish Turtle)
Sources index¶
sources-index.yaml is the single coordination point. Each entry maps a source repo's
CI artifact to a slot in the graph:
sources:
- id: enterprise-archimate
repo: archi-group/archimate-models
job: convert # CI job that produces the artifact
artifact: out/archimate.trig # path within the artifact archive
target: graph/archimate/enterprise-model.trig
tier: 1
ref: main
Adding a new source is one entry. pull-sources.sh reads this file and fetches each
artifact via the GitLab Jobs Artifacts API:
GitLab resolves the latest successful pipeline on :ref automatically — no pipeline
ID needed. The source repo's convert job must declare artifacts: paths: [out/] and
the artifact must not be expired.
Validation tiers¶
Validation is split across the two repos so problems are caught as early as possible.
| Tier | Where | When | Scope |
|---|---|---|---|
| 1 | Source repo CI | On MR | Structural SHACL — element types, relationship endpoints, no _gen_ fallback IRIs, view nodes linked to a diagram |
| 2 | This repo CI | On pull/merge | Cross-model SHACL — every element has a skos:prefLabel, no orphan relationships, cross-notation references resolve, org governance rules |
Tier 1 runs at the source and blocks the MR if the RDF is structurally invalid. By the time the RDF reaches the aggregation repo it is already structurally sound.
Tier 2 runs here, on the combined graph — it catches integration problems that
only appear when models from different notations are loaded together (e.g. a Backstage
Component claims system: X, but no System X exists in any graph).
Custom governance rules¶
Cross-model rules live in shapes/cross-model-rules.ttl:
@prefix sh: <http://www.w3.org/ns/shacl#> .
<#ComponentMustHaveOwner>
a sh:NodeShape ;
sh:targetClass <https://meta.linked.archi/archimate3/onto#ApplicationComponent> ;
sh:property [
sh:path <https://meta.linked.archi/backstage/onto#ownedBy> ;
sh:minCount 1 ;
sh:severity sh:Warning ;
sh:message "ApplicationComponent has no owner relationship." ;
] .
End-to-end process¶
sequenceDiagram
participant Dev as Developer
participant Src as Source Repo CI
participant API as GitLab Artifacts API
participant AG as archi-graph CI
participant Main as archi-graph main
participant TS as Triplestore / rdf2docs
Dev->>Src: push .puml / .bpmn / .yaml
Src->>Src: convert → TriG artifact
Src->>Src: validate (Tier 1)
Src-->>Dev: ✓ MR mergeable
Dev->>Src: merge to main
Note over Src: artifact stored on main pipeline
AG->>API: GET .../artifacts/main/raw/out/model.trig?job=convert
API-->>AG: raw TriG bytes
AG->>AG: git diff → changes detected
AG->>AG: branch + MR opened
AG->>AG: validate combined graph (Tier 2)
AG->>Main: merge (auto for Tier 1, manual for Tier 2)
Main->>TS: publish merged-graph.trig
The scheduled aggregation pipeline creates a branch update/graph-{timestamp}, commits
the pulled TriG, and opens an MR. Nothing reaches main without passing Tier 2.
The store's default graph must be a union¶
A converted model is not one graph. Its facts are spread across graph/semantic — one per input file,
where a model was built from several — plus graph/model for the curated part, graph/views for
geometry and graph/provenance for how it was made. They compose into a model only if the store treats
its default graph as the union of all named graphs.
Most stores need telling. In Fuseki it is tdb2:unionDefaultGraph true on the dataset; in GraphDB the
sesame:nil / "include inferred and all graphs" setting; in rdflib Dataset(default_union=True).
Without it a query with no GRAPH clause matches nothing, and the queries below silently return no rows
rather than failing.
Two things make the split safe to query across:
Every concept states its model. arch:inModel travels in the same graph as the concept, so the
model reassembles without depending on graph boundaries — and therefore survives a merged-graph.ttl
as well as a merged-graph.trig.
Every graph states where it came from. Each is a prov:Bundle in that model's provenance graph.
If you already have the graph IRI — you read it out of the output, or a previous query returned it — name it directly. It is one clause, it needs no provenance, and nothing beats it:
SELECT ?s ?p ?o WHERE {
GRAPH <…/backstage/service-catalog/graph/semantic/group-orders/catalog-info-yaml> { ?s ?p ?o }
}
The provenance route answers the question you cannot answer that way: which graph holds this file?
# everything one input produced, named by the file rather than by the graph
SELECT ?s ?p ?o WHERE {
?g prov:wasDerivedFrom ?src .
?src schema:name "catalog-info.yaml" ; dct:isPartOf <https://git.example.org/group/orders> .
GRAPH ?g { ?s ?p ?o }
}
# which input asserted a fact — more than one answer where more than one input asserts it
SELECT ?path WHERE {
GRAPH ?g { <…/element/component/default/order-service> bs:lifecycleState bs:Production }
?g prov:wasDerivedFrom ?src . ?src schema:name ?path .
}
Do not compose a graph IRI from a file path
The direct form above is fine when you have the IRI. Building one from a path is not, and it is the reason the provenance route exists:
- The slug is lossy.
catalog/checkout.yamlbecomescatalog-checkout-yaml; so doescatalog-checkout.yaml. You cannot invert it. - The shape depends on the run. With a repository known the graph is
…/graph/semantic/group-orders/catalog-info-yaml; without one it is…/graph/semantic/catalog-info-yaml. The same file gives two different IRIs depending on whetherCI_PROJECT_URLwas set. - A single-input conversion is not partitioned at all — its facts are in plain
…/graph/semantic. - ADR 0008 §3 declares these IRIs opaque, which is what makes all of the above safe to change.
In an aggregate, pin dct:isPartOf alongside schema:name as the query above does. Paths repeat
across repositories — every Backstage descriptor is catalog-info.yaml — so the name alone will match
graphs from every repository that has one.
A re-pull replaces a graph, it does not accumulate one. A per-source graph is named after the file,
not after the file at a commit, so a descriptor whose commit moved lands in the same graph as before.
The prov:Entity describing that file does carry the commit, so the revisions accumulate there — which
is where the history belongs. Loading two runs of the same source therefore gives one set of graphs and
one current value per property, alongside one source entity per revision read.
Versioning & overwrites¶
Each pull overwrites the target file in place. pull-sources.sh downloads every
source to its fixed target path (curl -o), so graph/.../model.trig always holds
the latest successful pipeline output from the configured ref. There is no version
suffix and no accumulation of old files.
Git is the version history — not the filenames. Because the TriG is committed:
git diffongraph/shows exactly which triples changed on each pullgit log/git blametrace when a change landed and which MR introduced it- reverting a bad graph state is a normal
git revert
Overwriting is safe because IRIs are stable and deterministic — they are minted from
source identifiers and the model-id, never from filenames. Re-pulling the same model
therefore produces a clean diff (changed triples only) rather than duplicate resources,
and each model's named-graph IRIs are unique, so merge.sh never collides across sources.
On the source side: artifacts are regenerated, models are versioned¶
The RDF artifact is a build output, not a versioned asset. In the source repo out/
is gitignored, and the convert job rewrites it from scratch on every pipeline. So the
artifact itself is disposable — the real version history lives one level up:
- The source models are the versioned truth. The
.puml/.bpmn/catalog-info.yamlfiles are committed to the source repo's git. Their history (who changed the model, when, in which MR) is the authoritative record. The artifact is just a deterministic projection of whatever model sits on a given ref. - Every pipeline regenerates the artifact. On each MR and each
maincommit theconvertjob produces a freshout/*.trig, overwriting the previous run. GitLab retains one artifact set per pipeline (bounded byexpire_in), so historical versions exist only as long as their pipeline's artifact is retained. - "Latest" is resolved by GitLab. The pull hits
/jobs/artifacts/{ref}/raw/...?job=convert, which returns the artifact from the latest successful pipeline on that ref. You never address a pipeline ID — moving the ref (e.g. a new commit onmain) automatically changes what gets pulled next.
Two consequences worth designing for:
- Reproducibility depends on the converter version, not just the model. The same
model converted with a different converter image can yield different triples. Pin
CONVERTER_IMAGEto a fixed tag (not:latest) in the source repo if you need a regenerated artifact to match a past one byte-for-byte. - Pulling by tag needs the tagged artifact to still exist.
ref: v2.3.0only works if the pipeline that built thev2.3.0tag still has its artifact (useexpire_in: neveron tag/main pipelines). If the artifact expired, re-run the tagged pipeline to regenerate it — deterministic IRIs mean the output is equivalent.
Pinning a specific version¶
To freeze one source at a released version, set its ref to a tag (or branch) instead
of main:
- id: enterprise-archimate
repo: archi-group/archimate-models
job: convert
artifact: out/archimate.trig
target: graph/archimate/enterprise-model.trig
tier: 1
ref: v2.3.0 # pull from this tag instead of main
Keeping multiple versions live at once¶
Overwrite means only one version of a given source lives in the graph at a time. If you
need v1 and v2 present simultaneously (e.g. to model a migration), give them
distinct identities on two axes:
-
Distinct index entries — separate
id+target, each pulling from its ownref: -
Distinct IRIs at conversion time — the two versions must be converted with a different
model-id(or base IRI) in the source repo, so their resource and named-graph IRIs don't overwrite each other when merged. Two files pulled with the samemodel-idwould emit the same graph IRIs and the second would shadow the first in the merged output.
In short: versioning across time is free via git; versioning side-by-side requires separate identities in both the index and the minted IRIs.
Why this design¶
- Convert at source — the team that owns a model validates it, gets fast MR feedback, and conversion compute is distributed instead of bottlenecked centrally.
- The TriG artifact is the contract — source repos can use any tool or language, as long as the output conforms to the Linked.Archi foundational ontology.
- The graph IS the repository — committed TriG means
git clonegives the full graph (no triplestore required),git diffshows what changed, andgit blametraces who introduced which triples. - Merge via MR (never direct push) — every change to the graph is reviewable, validated, and auditable; auto-merge is allowed for trusted Tier 1 sources.
- Scheduled pulls — simpler and more resilient than per-repo webhooks; naturally batches one MR per run. A downstream trigger can be added for latency-sensitive Tier 1 sources.
This scales to hundreds of source repos: each is independent, sources-index.yaml is
the only coordination point, and the merge step is O(files) concatenation.