26
2 Research Data Infrastructures and Engineering Metadata
automated toolchain (as an external tool) was implemented using the Dataverse API
for the specific use case of thermodynamics: after a simulation run, an automated
metadata extraction is triggered. Then, the extracted metadata altogether with the
data is automatically ingested into the DaRUS repository [19].
2.2.3.2 NOMAD
In contrast to DaRUS, the Novel Materials Discovery (NOMAD) laboratory
9 (or
Novel Materials Discovery Center of Excellence (NOMAD CoE)) is a prime example of a domain-specific data infrastructure which is highly integrated [20] in a virtual research environment. The repository part is complemented with the NOMAD
Archive, the NOMAD Encylopedia, the NOMAD Visualization Tools and the
NOMAD Analytics Toolkit. NOMAD is recommended by Nature
10 for depositing
supplementary data when submitting a research article on materials modelling.
The NOMAD repository is the central component of the laboratory and holds
input and output data from material simulations with a retention period of 10 years for
free. The NOMAD archive holds the open-access data from the repository which was
converted into a code-independent format. To accomplish this, developing a metadata
definition and a metadata component was crucial for this. It serves, just as proposed
in Sect. 2.1.1.1, as a common understanding
11 and, as the overall outline of this book,
for making data semantically interoperable. The metadata definition uses 168 aligned
and 2,360 code-specific metadata keys. For example, the different terms for quantities
had to be mapped to one aligned term. According to [20], the development of this
component of the data infrastructure was a challenge. The NOMAD encylopedia is
the part of the NOMAD data infrastructure which provides millions of calculations
via a web GUI with a materials-oriented view and therefore serves as knowledge
base and a material classification system. The NOMAD visualization tools are a
centralized service for data visualization within the data infrastructure allowing users
interactive graphical analysis in materials modelling. Additionally, the NOMAD
Analytics Toolkit is a big data analytics approach to support data evaluation, for
example, scanning for specific thermoelectric materials or finding suitable materials
for heterogeneous catalysis.
In the NOMAD laboratory, the archive and the repository components correspond
to the storage layer (l1) and the object layer (l2), whereas the encyclopedia, the
analytics toolkit and the visualization tools correspond to the service layer (l3),
which is strongly coupled to the base layers.
As of February 2020, the NOMAD data infrastructure holds 49TB of raw data
in the repository and 19TB of the archive in normalized, annotated form in 758
datasets.
12
9 https://www.nomad-coe.eu.
10 https://www.nature.com/sdata/policies/repositories.
11 https://www.nomad-coe.eu/the-project/nomad-archive/archive-meta-info.
12 https://metainfo.nomad-coe.eu/nomadmetainfo_public/archive.html.
Précédent

- 35/101

Suivant