2.2 Research Data Infrastructures
25
Fraunhofer-Gesellschaft is using it to store research data in its institutional repository
Fordatis
7 [16].
Research data infrastructures can be classified as institutional and domain-specific
infrastructures. Institutional data infrastructures resemble to research data management on an institutional level and are not bound to a specific discipline. An example
of this type is DaRUS, which will be discussed in Sect. 2.2.3.1 of this chapter. A
domain-specific data infrastructure serves as an approach which is bound to a specific
discipline and can span across multiple institutions. An example of a domain-specific
data infrastructure for materials modelling is NOMAD, which will be discussed in
Sect. 2.2.3.2.
2.2.3 Examples of Research Data Infrastructures in
Materials Modelling
2.2.3.1 DaRUS
Even though the Data Repository of the University of Stuttgart (DaRUS)
8 is an institutional repository and not limited to materials modelling, it will be discussed here
since its development was strongly driven by the EngMeta metadata model. Moreover, it is an example of a loosely coupled data infrastructure. Its overall development
was urged by the need of a sustainable repository for the University of Stuttgart and,
in particular, the materials modelling community at the university as well as by the
precursory design of EngMeta. Within the repository, EngMeta serves as the semantic core and the repository is built around the metadata model, which is also deemed
metadata-driven repository development. The requirements, such as handling large
datasets, were stemming from aerodynamics and molecular dynamics [17].
DaRUS is based on Dataverse, and the driving factors for choosing this repository
software were its design for research data management, its integration with the DOI
persistent identifier infrastructure, its adaptability with metadata standards and its
monolithic design. In the Dataverse repository software package, all the data is organized in Dataverses (organizational structure), datasets and files [18]. A Dataverse is
the highest element in the hierarchical data organization structure in the repository
and typically represents an institute or a research project. A dataset in the Dataverse
terminology resembles to a directory or a collection of files. As of July 2020, DaRUS
holds almost 600 files in 49 datasets, which are organized in 60 Dataverses, mainly
from the fields of engineering, computer science and physics.
As DaRUS is an institutional repository, it is only loosely coupled to the research
infrastructure since it is generic. This means that the service layer (l3) is basically
the generic Dataverse web GUI. Additional services can be integrated by using
one of the APIs that Dataverse offers, such as REST or SWORD. For example, an
7 https://fordatis.fraunhofer.de/.
8 https://darus.uni-stuttgart.de/.
25
Fraunhofer-Gesellschaft is using it to store research data in its institutional repository
Fordatis
7 [16].
Research data infrastructures can be classified as institutional and domain-specific
infrastructures. Institutional data infrastructures resemble to research data management on an institutional level and are not bound to a specific discipline. An example
of this type is DaRUS, which will be discussed in Sect. 2.2.3.1 of this chapter. A
domain-specific data infrastructure serves as an approach which is bound to a specific
discipline and can span across multiple institutions. An example of a domain-specific
data infrastructure for materials modelling is NOMAD, which will be discussed in
Sect. 2.2.3.2.
2.2.3 Examples of Research Data Infrastructures in
Materials Modelling
2.2.3.1 DaRUS
Even though the Data Repository of the University of Stuttgart (DaRUS)
8 is an institutional repository and not limited to materials modelling, it will be discussed here
since its development was strongly driven by the EngMeta metadata model. Moreover, it is an example of a loosely coupled data infrastructure. Its overall development
was urged by the need of a sustainable repository for the University of Stuttgart and,
in particular, the materials modelling community at the university as well as by the
precursory design of EngMeta. Within the repository, EngMeta serves as the semantic core and the repository is built around the metadata model, which is also deemed
metadata-driven repository development. The requirements, such as handling large
datasets, were stemming from aerodynamics and molecular dynamics [17].
DaRUS is based on Dataverse, and the driving factors for choosing this repository
software were its design for research data management, its integration with the DOI
persistent identifier infrastructure, its adaptability with metadata standards and its
monolithic design. In the Dataverse repository software package, all the data is organized in Dataverses (organizational structure), datasets and files [18]. A Dataverse is
the highest element in the hierarchical data organization structure in the repository
and typically represents an institute or a research project. A dataset in the Dataverse
terminology resembles to a directory or a collection of files. As of July 2020, DaRUS
holds almost 600 files in 49 datasets, which are organized in 60 Dataverses, mainly
from the fields of engineering, computer science and physics.
As DaRUS is an institutional repository, it is only loosely coupled to the research
infrastructure since it is generic. This means that the service layer (l3) is basically
the generic Dataverse web GUI. Additional services can be integrated by using
one of the APIs that Dataverse offers, such as REST or SWORD. For example, an
7 https://fordatis.fraunhofer.de/.
8 https://darus.uni-stuttgart.de/.
