1.4 Perspective and Outline of the Book
9
words of Kuhn, scientific revolutions are characterized by a “change in several of the
taxonomic categories prerequisite to scientific descriptions and generalizations. That
change, furthermore, is an adjustment not only of criteria relevant to categorization,
but also of the way in which given objects and situations are distributed among preexisting categories” [50]. Insisting on the adoption of a single ontology by a whole
field of science makes that field perfectly interoperable, but at the price of scientific
stagnation. In this view, any field that is developing fast—as modelling and simulation technology today certainly is—will consider a plurality of paradigms, semantic
heterogeneity, as an indicator of its success, rather than as an obstacle, and pursue
the frequent renegotiation of alignments between multiple ontologies, taxonomies
and hierarchical schemas [26, 27, 51].
However, even if knowledge is formalized in a machine-readable format, besides
the design of the ontology, there will still be other steps where human intervention is
needed, notably, to classify and annotate the individuals. Whenever this occurs, we
rely on the fact that the semantics is really shared, i.e. that there is an actual common
understanding of all the concepts by all the users. In practice, colloquial language
and insufficient familiarity with concept definitions can give rise to common pitfalls
and false friends in using ontologies. By stating that something is generally true, we
could mean “all the time” or “most of the time;” in the EMMC context, the words
translation and mesoscopic have very specific meanings that would be misconstrued
intuitively by most domain experts in materials modelling. Hence, human error must
be anticipated (in addition to genuine disagreements on how an ontology should
be applied), and similarly, automated annotation tools based on natural language
processing are not free of error. All of this highlights the need for a community to
gather and share concepts in a continuous effort; it corroborates the conclusion that
there is a trade-off between the expressive power of ontologies and the associated
social and technological cost of development and maintenance.
The remainder of the book is structured as follows: Chapter 2 introduces the XSDbased engineering metadata schema EngMeta and the research data infrastructure
DaRUS, situating it in the context of the emerging environment of databases and
repositories for data from physics-based modelling and simulation. The system of
marketplace-level domain ontologies developed by VIMMP is presented in Chap. 3,
concerning data provenance, services and transactions at the marketplace level, and
in Chap. 4, concerning the description of solvers, associated aspects such as licenses
and software features, and the characterization of physical and other variables that
occur in modelling and simulation. On this basis, Chap. 5 addresses issues related
to the practical use of the metadata standards, including syntactic interoperability
and concrete scenarios from molecular modelling and simulation; it also discusses
challenges that arise from semantic heterogeneity, wherever multiple interoperability
standards are concurrently employed for identical or overlapping domains of knowledge, or where domain ontologies need to be matched to top-level ontologies such
as the EMMO.
9
words of Kuhn, scientific revolutions are characterized by a “change in several of the
taxonomic categories prerequisite to scientific descriptions and generalizations. That
change, furthermore, is an adjustment not only of criteria relevant to categorization,
but also of the way in which given objects and situations are distributed among preexisting categories” [50]. Insisting on the adoption of a single ontology by a whole
field of science makes that field perfectly interoperable, but at the price of scientific
stagnation. In this view, any field that is developing fast—as modelling and simulation technology today certainly is—will consider a plurality of paradigms, semantic
heterogeneity, as an indicator of its success, rather than as an obstacle, and pursue
the frequent renegotiation of alignments between multiple ontologies, taxonomies
and hierarchical schemas [26, 27, 51].
However, even if knowledge is formalized in a machine-readable format, besides
the design of the ontology, there will still be other steps where human intervention is
needed, notably, to classify and annotate the individuals. Whenever this occurs, we
rely on the fact that the semantics is really shared, i.e. that there is an actual common
understanding of all the concepts by all the users. In practice, colloquial language
and insufficient familiarity with concept definitions can give rise to common pitfalls
and false friends in using ontologies. By stating that something is generally true, we
could mean “all the time” or “most of the time;” in the EMMC context, the words
translation and mesoscopic have very specific meanings that would be misconstrued
intuitively by most domain experts in materials modelling. Hence, human error must
be anticipated (in addition to genuine disagreements on how an ontology should
be applied), and similarly, automated annotation tools based on natural language
processing are not free of error. All of this highlights the need for a community to
gather and share concepts in a continuous effort; it corroborates the conclusion that
there is a trade-off between the expressive power of ontologies and the associated
social and technological cost of development and maintenance.
The remainder of the book is structured as follows: Chapter 2 introduces the XSDbased engineering metadata schema EngMeta and the research data infrastructure
DaRUS, situating it in the context of the emerging environment of databases and
repositories for data from physics-based modelling and simulation. The system of
marketplace-level domain ontologies developed by VIMMP is presented in Chap. 3,
concerning data provenance, services and transactions at the marketplace level, and
in Chap. 4, concerning the description of solvers, associated aspects such as licenses
and software features, and the characterization of physical and other variables that
occur in modelling and simulation. On this basis, Chap. 5 addresses issues related
to the practical use of the metadata standards, including syntactic interoperability
and concrete scenarios from molecular modelling and simulation; it also discusses
challenges that arise from semantic heterogeneity, wherever multiple interoperability
standards are concurrently employed for identical or overlapping domains of knowledge, or where domain ontologies need to be matched to top-level ontologies such
as the EMMO.
