Provenance Data Models and Assertions: A Demonstrative Approach
109
After a careful study of OPM, the provenance working group proposed a formal
definition of provenance known as the PROV-DM data model which is considered
the universal standard provenance model after the OPM, that would specify rules
that would help in the creation of provenance [17].
4.2 PROV-DM Data Model
“Data provenance also known as pedigree or lineage is the description of a piece of
information from its origins” [18]. It specifies the processes by which the given piece
of data in a given database. We consider the example of a database to understand
the above scenario and to demonstrate the need for why and where model of data
provenance. Suppose a database D = Q(d) is constructed from a query Q over the
data d, and the user wishes to evaluate how and by which source and process the
given data is derived from the database or which are the tuples that contributed to the
formation of the given piece data item. The users of these databases are not able to
track the accuracy and the timeliness of data being delivered to them. The data in the
field of scientific research is derived from various public databases, however only
a handful of these databases provide experimental data, the remaining databases do
not provide for good quality data and are in some way or the other the dataset views
of the other databases. This is because these remaining databases carry additional
values added to them by experts by the means of corrections and annotations. These
databases although carry huge sources of information yet the users of these databases
are oblivious to how the provenance information can be tracked by them. These
problems are addressed in [18, 19] by stating their derivation in a relational database
using the tuples that led to the development of the piece of a data item. However,
questions like “from where did the piece of data of data item come” and “why is it in
the database” remain unanswered. The who and how assertions, with help entity and
activity of the provenance in a data model seek to find answers to such questions.
Contemporary researches have demonstrated the why of provenance models and
been reflected by authors in their articles [19, 20].
This chapter demonstrates a generic implementation of provenance over the
PROV-DM data model that allows us to compute, derive, and understand both the
why and where of a given piece of data item [21].
5 The Provenance Architecture
The PROV cake is the basic data model proposed to represent entities, human agents,
and objects that may be associated with generating data pieces. The provenance stack
is layered, and the details of each layer are elaborated in the following sections.
Précédent

- 126/424

Suivant