24
2 Research Data Infrastructures and Engineering Metadata
Table 2.2 Data repository software
Repository
Origin
Sample installation
(Type[,field])
Dataverse
Data management
University of Stuttgart/DaRUS
(institutional)
Dspace
Document management
Fraunhofer
Gesellschaft/Fordatis
(institutional)
Fedora
Document management
Saarland University/CLARIN
(domain-specific,
linguistics)
Invenio
Data management
Swiss National Computing
Centre/Materials Cloud
ARCHIVE (domain specific,
materials modelling)
basic functionality is metadata management. By this layer, data from the storage
layer is enriched with metadata and data objects become information objects with
a persistent identifier, whose purpose is to make the data citable. The third layer is
the service layer (l3) and includes the user interface and marks the visible part of
the data infrastructures. Moreover, this layer includes additional services, such as an
automated metadata extraction.
Basically, data infrastructures implement all three layers; however, they can operate or work in distributed environments. Usually the base layer (l1) is the hardware
part of the data infrastructure, whereas the layers (l2) and (l3) are the software part.
The functionalities of the layers (l2) and (l3) are usually covered by repository software. A repository is a store for data that organizes this data in some logical manner
and makes the data available for usage to a specified group of persons. It is important
to mention that a repository is not a filesystem, which means that its purpose is not
to manage the files in directory structures. In contrast, a repository must be imagined
as collections of files organized in sets (of some logical manner, for example, as
datasets, as linked data, in a loose hierarchical structure,...), which are described by
metadata, are search and retrievable, and are provided with a persistent identifier.
Out-of-the-box generic repository software packages are generally available and
serve different purposes. Some of those packages stem from document management,
whereas others have their origins in data/file management. Their origin has to be
taken into account when evaluating the repository for a specific use case or domain.
Table 2.2 gives an overview over typical data repository software. For example, Dataverse originates from the management of datasets, whereas Dspace stems from managing document files.
6 However, also Dspace is capable of managing datasets, and the
6 In the context of this chapter, iRods has to be mentioned. Even though it is not a classical repository
software package but offers a unified namespace, its functionalities include repository-style data
management on a filesystem level [15].
Précédent

- 33/101

Suivant