that are erroneously measured or that deviate from what is regularly measured in the
system (e.g. a water distribution network or a treatment plant).
Detection is performed by developing a validation technique for the object that
generates data (e.g. a sensor) at the time it generates the data, with various techniques
having been developed for this purpose in literature [3, 33, 34]. In typical operational
cases, data validation is carried out manually by expert judgment using both analytic
and visualization tools. The issue with that approach is that with current data streams
only a small amount of data can be validated by operatives from the utilities [2], and
as evidenced by several authors, there is always a human bias in the decision making.
This human bias is particularly important while determining whether data corresponds to anomalous/irregular behaviour (which can be explained, e.g. due to
extremes or failure of network parts, such as leakages and pipe bursts) or faulty
data (unexplained, e.g. due to sensor faults) [35], as explained in Fig. 4. In this text,
faulty data corresponds to data belonging to the two error categories explained in
Sect. 2.1., i.e. measurement and human errors only.
An inventory of faulty data validation techniques is presented in §3.2, based on a
literature review on the subject. In general, there are sequential steps for data
validation corresponding to:
(a) Input variable selection, which consists of the selection of a subset of interest
from the data warehouse that have to be validated [36].
(b) Pre-processing, which includes a number of statistical and modelling techniques
that help to identify anomalies, either by statistical analyses or by contrasting
real data to modelled data equivalents.
(c) Anomaly and faulty data detection, which can be performed with either simple
or advanced (statistical) methods.
The diagram of Fig. 5 also provides an indication of the amount of data required
(Data arrow) and the amount of time (Time arrow) for each technique. Evidently,
more advanced techniques require more data and time. In pre-processing techniques,
the use of models or meta-models may drastically increase the data and computational requirements, but it may lead to a significant reduction of the uncertainty in
faulty data identification.
3.2 Faulty Data Detection Techniques
Faulty data detection techniques are generally classifiers which divide the data in two
classes (correct and faulty/doubtful). While some detection techniques have been
applied for generic problems [29, 37, 38], water research has also developed domainspecific techniques for sewer systems [36], geo-hydrological systems [4, 6, 7, 33,
39], water quality sensing [40], automatic or real-time data validation in urban
systems [2] and specific problems such as the determination of leakages as anomalous data in water supply systems [41]. As depicted in Fig. 5, these techniques are
divided based on their complexity to two main categories: simple tests and statistical
tests.
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
77
system (e.g. a water distribution network or a treatment plant).
Detection is performed by developing a validation technique for the object that
generates data (e.g. a sensor) at the time it generates the data, with various techniques
having been developed for this purpose in literature [3, 33, 34]. In typical operational
cases, data validation is carried out manually by expert judgment using both analytic
and visualization tools. The issue with that approach is that with current data streams
only a small amount of data can be validated by operatives from the utilities [2], and
as evidenced by several authors, there is always a human bias in the decision making.
This human bias is particularly important while determining whether data corresponds to anomalous/irregular behaviour (which can be explained, e.g. due to
extremes or failure of network parts, such as leakages and pipe bursts) or faulty
data (unexplained, e.g. due to sensor faults) [35], as explained in Fig. 4. In this text,
faulty data corresponds to data belonging to the two error categories explained in
Sect. 2.1., i.e. measurement and human errors only.
An inventory of faulty data validation techniques is presented in §3.2, based on a
literature review on the subject. In general, there are sequential steps for data
validation corresponding to:
(a) Input variable selection, which consists of the selection of a subset of interest
from the data warehouse that have to be validated [36].
(b) Pre-processing, which includes a number of statistical and modelling techniques
that help to identify anomalies, either by statistical analyses or by contrasting
real data to modelled data equivalents.
(c) Anomaly and faulty data detection, which can be performed with either simple
or advanced (statistical) methods.
The diagram of Fig. 5 also provides an indication of the amount of data required
(Data arrow) and the amount of time (Time arrow) for each technique. Evidently,
more advanced techniques require more data and time. In pre-processing techniques,
the use of models or meta-models may drastically increase the data and computational requirements, but it may lead to a significant reduction of the uncertainty in
faulty data identification.
3.2 Faulty Data Detection Techniques
Faulty data detection techniques are generally classifiers which divide the data in two
classes (correct and faulty/doubtful). While some detection techniques have been
applied for generic problems [29, 37, 38], water research has also developed domainspecific techniques for sewer systems [36], geo-hydrological systems [4, 6, 7, 33,
39], water quality sensing [40], automatic or real-time data validation in urban
systems [2] and specific problems such as the determination of leakages as anomalous data in water supply systems [41]. As depicted in Fig. 5, these techniques are
divided based on their complexity to two main categories: simple tests and statistical
tests.
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
77
