Skip to content

3.3 Tools & enablers for FAIR-aligned data management

Together, these tools turn FAIR‑aligned principles into operational capabilities that support scalable, automated, and trustworthy AI and HPC workflows.

3.3.1 Data catalogs

Data catalogs act as a central inventory for datasets, models, and documentation across an organization. They help teams discover existing assets, understand lineage and ownership, and avoid duplicating effort. Catalogs also improve transparency by showing how data connects to ongoing projects and how it is used across teams.

Examples of open‑source data catalog solutions include Amundsen, DataHub, and OpenMetadata.

Catalogs can integrate with APIs so AI and HPC pipelines can query metadata automatically and retrieve the correct datasets without manual intervention. In large industrial environments, a data catalog often becomes the navigation layer that anchors efficient data discovery, reuse, and governance across AI and HPC workflows.

3.3.2 Metadata schemas, controlled vocabularies & ontologies

Metadata schemas, controlled vocabularies, and ontologies ensure that datasets are described and interpreted consistently across tools and teams. Schemas define how data is structured, while vocabularies and ontologies define the meaning of key terms, providing semantic clarity across workflows. This uniformity reduces integration friction, supports automation, and helps avoid costly misinterpretations.

Common metadata schemas and semantic foundations

  • DCAT – for structured dataset descriptions
  • schema.org – for general‑purpose metadata
  • Domain‑specific metadata standards defined by research or industry communities
  • Controlled vocabularies and ontologies that provide shared semantic meaning across systems

By adopting shared schemas and semantic standards, organizations ensure that AI and HPC systems can parse and process data with minimal manual intervention. This becomes increasingly important as datasets grow in size, complexity, and frequency of reuse.

Semantic standards for automated workflows

While general schemas and vocabularies define structure and meaning, some semantic standards extend this foundation by describing how data is produced, packaged, and governed in automated AI and HPC workflows. These standards extend foundational metadata by encoding provenance, context, and usage conditions in a way that supports reproducibility, traceability, and controlled reuse at scale.

  • RO‑Crate packages datasets together with structured, JSON‑LD‑based metadata describing content, context, and provenance. It is well suited for sharing and preserving AI training datasets, simulation outputs, and benchmark data across compute environments. 🔗 https://www.researchobject.org/ro-crate/
  • PROV‑O is a W3C ontology for expressing provenance and lineage, enabling traceability, auditability, and reproducibility across complex, multi‑step AI and HPC pipelines. 🔗 https://www.w3.org/TR/prov-o/
  • ODRL is a W3C‑standardized language for expressing usage permissions and constraints in a machine‑readable form, supporting governed reuse of proprietary or restricted data in industrial environments. 🔗 https://www.w3.org/TR/odrl-model/

Together with metadata schemas, vocabularies, and ontologies, these standards ensure that meaning, provenance, and usage conditions travel with the data. This reduces manual interpretation, strengthens governance, and supports scalable, trustworthy AI and HPC workflows in highly automated environments.

3.3.3 Persistent Identifiers (PIDs)

Persistent identifiers (PIDs) provide stable, long‑term identities for datasets, models, and versions, enabling teams to trace exactly which inputs were used in specific simulations or training runs. They help prevent confusion between dataset versions and make it easier to manage large, evolving data assets. In industrial AI and HPC workflows, PIDs are essential for reproducibility, auditability, and long‑term traceability, especially when results must be verified months or years later.

In practice, what matters is that identifiers are globally unique, resolvable, and persistent over time, even as storage locations or infrastructures change.

PID examples

  • DOIs (Digital Object Identifiers) for datasets, simulation outputs, or benchmark data. Widely used for long‑term identification and citation of datasets. 🔗 https://www.doi.org 🔗 https://datacite.org
  • Handle identifiers for internal or restricted datasets. Commonly used in enterprise and research infrastructures. 🔗 https://handle.net
  • Model or dataset version identifiers embedded in ML workflows. Used to trace which exact version of data or model was used in training or inference (e.g. dataset‑v1.2.3, model‑hash‑ID)
  • Software Heritage identifiers (SWHIDs) for code tied to data processing and training. Useful for linking datasets to the exact code used to generate or transform them. 🔗 https://www.softwareheritage.org
  • Internal PID systems

Many organizations use internal resolvable identifiers (often backed by catalog or metadata systems) for proprietary data while keeping them FAIR‑aligned.

Using PIDs consistently ensures that teams can confidently revisit past experiments, understand dependencies between data, models, and code, and maintain a clear chain of custody across the data lifecycle without relying on fragile file paths or ad‑hoc naming conventions.

3.3.4 Knowledge graphs

Knowledge graphs connect datasets, models, processes, metadata, and documentation into a unified network of relationships. They make dependencies explicit such as which dataset feeds which model, how data versions relate to experiments, or which preprocessing steps were applied helping teams navigate complex AI and HPC data landscapes. Knowledge graphs also enable richer discovery, allowing users and systems to find related data even when terminology or labels differ between teams.

By capturing relationships explicitly, knowledge graphs help prevent silos and improve semantic consistency across the organization. This supports better decision‑making, smoother collaboration, and more reliable automated pipelines in large industrial AI and HPC environments, where understanding dependencies is critical for reproducibility, governance, and scalable operations.

Figure: Knowledge graphs represent networks of relationships where knowledge emerges through connections between data, models, processes, and metadata

Open‑source technologies commonly used to build knowledge graphs include graph databases and semantic‑web frameworks such as Neo4j, JanusGraph, RDF stores, and related tooling.

Check your understanding
Question 1 of 3

How do data catalogs support FAIR-aligned data management?

Select all that apply.

© 2026 LUMI AI FactoryContent licensed under CC BY 4.0Code licensed under the MIT Licence

The LUMI AI Factory Service Center is funded jointly by the EuroHPC Joint Undertaking and the Participating States FI, CZ, DK, EE, NO, PL.