6
1 Introduction
Fig. 1.2 The four metadata categories and their level of specificity
Table 1.1 Examples of relevant metadata standards for the four categories
Metadata category
Standard
Description/Remarks
Descriptive
DataCite
DOI compatibility
DublinCore
Technical
PREMIS
Preservation MD
Process
PROV
Provenance MD
CodeMeta
For code and software
Domain
EngMeta
For computational engineering
TEI
For digital humanities
key might be part of two or more categories. The probability that suitable standards
exist for the four categories is decreasing with the categories’ specificity, as shown
in Fig. 1.2. Whereas for technical and descriptive metadata, many standards exist,
this does not hold for the latter two categories, which require a significant dedicated development effort. Regarding technical metadata, the semantic information
is similar in all research fields as long as the data are organized in files. A typical standard here is PREMIS [39]. Also descriptive metadata keys are similar (or
even the same) throughout all disciplines. Here, DataCite is the de facto standard
for a general description and citable data objects [40]. In contrast, process metadata are strongly related to the research process, where metadata standards only exit
for specific processes, e.g. CodeMeta [41] and the Citation File Format [42] for the
description of software and codes. For domain-specific metadata, only standards for
specific research objects exist. Some relevant standards for to the four categories are
shown in Table 1.1. Knowledge of the categories and existing standards enables the
metadata designer to use certain parts as building blocks when compiling a standard
for a certain area.
Moreover, the distinction between these four categories is crucial with respect to
automated extractability. It has been shown that some categories are easier to automatically extract than others in computational engineering [38]: Technical information is
Précédent

- 15/101

Suivant