14
2 Research Data Infrastructures and Engineering Metadata
2.1.1 How to Engineer Metadata
The art of engineering a metadata model includes several consecutive steps which are
described in this subsection. It may happen that this process or a single step has to be
iterated several times to come to a fine-grained, purposeful description of the research
asset. In short, the following steps are necessary to engineer a metadata model. First, a
consensus must be reached about what metadata actually serves in the single context.
Then, an object model has to be carved out of the research process. Last, the object
model has to be transferred to a formal representation and implemented and therefore
becomes a metadata model.
2.1.1.1 Definitions of Metadata and Metadata Models
However, in the beginning of designing metadata for a certain purpose, it first has to
be discussed how metadata is defined. Usually, metadata is defined as a structured
form of knowledge representation, or simply, as many authors put it, “data about
data” [2]. Edwards describes this as the holy grail of information science:
Extensive, highly structured metadata often are seen as a holy grail, a magic chalice both necessary and sufficient to render sharing and reusing data seamless, perhaps even automatic. [3,
p. 672]
However, metadata is always strongly context dependent. To tackle their context
dependence, metadata must serve as a mode of communication:
We propose an alternative view of metadata, focusing on its role in an ephemeral process of
scientific communication, rather than as an enduring outcome or product. [3, p. 667]
Following this, metadata takes the role of semantic technology: Its task is to relieve
the direct communication and negotiation of data producers and data consumers and
should therefore diminish “science friction” [3], which occurs in every process where
research data is exchanged. To illustrate science friction, imagine two researchers
exchanging a dataset, which is not properly described by metadata. The receiver
might suppose the variable t i as a data point in a time series. To provide clarification,
the receiver would have to contact the sender of the data, and also in this process can
be defective. This example shows the importance of metadata as semantic asset, and
therefore as a mode of fixed, negotiated communication.
Additionally, as Jane Greenberg puts it, metadata should semantically support the
specific workflow [4]. For example, metadata describes a data point with an error bar
and defines the form of the error. Thus, metadata would support the interpretation of
the data point.
Following the discussion of metadata, a metadata model then can be seen as the
middle ground of a non-formal model and a complete formalization of metadata
Précédent

- 23/101

Suivant