and their data products as a cyclic process. The PDCA approach can be viewed as a
proactive framework which continuously monitors and registers data, checks their
integrity, acts upon the checked datasets to feed information-based decision-making
and plans strategies, including proposition of improvements on the sensing/monitoring system which in turn closes the loop [26, 27].
2.2 The Role of Validation in Data Quality Control
Within the DQC context, validation techniques play a key role in connecting the
wealth of information obtained by raw data acquisition with decision-making and
planning. The acquired data (i.e. the result of a “Do” step in a PDCA cycle, see
Fig. 2) needs to be checked against errors and, in case faults are detected, needs to be
corrected before feeding any decision-making process (i.e. the steps of “Act” and
“Plan”). To complete this transition, a “Check” step is needed, which is better known
in information analysis as Data Validation [23].
As seen in Fig. 3, data validation can be further distinguished in three steps:
Collection, Detection and Correction. Data collection refers to the process of gathering data through data streams from each sensing device to a database, otherwise
known as a data warehouse. The step that follows is the detection of a subset of data
which could be deemed faulty. Detection techniques have to ensure that they can
safely distinguish between actual faulty data and data which appears doubtful but its
deviation could be attributed to something else than an error (Fig. 4). As a last step,
the data confirmed to be faulty need to be corrected (e.g. empty values filled, outliers
corrected based on other close values, etc.) before the data can be interpreted further
and used as a basis for decision-making. This stepwise process of identifying and
Fig. 3 The three steps
comprising identification of
faulty data
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
71
proactive framework which continuously monitors and registers data, checks their
integrity, acts upon the checked datasets to feed information-based decision-making
and plans strategies, including proposition of improvements on the sensing/monitoring system which in turn closes the loop [26, 27].
2.2 The Role of Validation in Data Quality Control
Within the DQC context, validation techniques play a key role in connecting the
wealth of information obtained by raw data acquisition with decision-making and
planning. The acquired data (i.e. the result of a “Do” step in a PDCA cycle, see
Fig. 2) needs to be checked against errors and, in case faults are detected, needs to be
corrected before feeding any decision-making process (i.e. the steps of “Act” and
“Plan”). To complete this transition, a “Check” step is needed, which is better known
in information analysis as Data Validation [23].
As seen in Fig. 3, data validation can be further distinguished in three steps:
Collection, Detection and Correction. Data collection refers to the process of gathering data through data streams from each sensing device to a database, otherwise
known as a data warehouse. The step that follows is the detection of a subset of data
which could be deemed faulty. Detection techniques have to ensure that they can
safely distinguish between actual faulty data and data which appears doubtful but its
deviation could be attributed to something else than an error (Fig. 4). As a last step,
the data confirmed to be faulty need to be corrected (e.g. empty values filled, outliers
corrected based on other close values, etc.) before the data can be interpreted further
and used as a basis for decision-making. This stepwise process of identifying and
Fig. 3 The three steps
comprising identification of
faulty data
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
71
