data that cannot be explained by the underlying physical rules of the measured
system.
2 These anomalies, otherwise known as errors, can be further distinguished in
three types [10]:
1. Measurement errors (e.g. failure of data registration, maintenance problems,
drifts, bias, strong gradients, lack of redundancy, problems of coherence at both
local and global scale, duplication of data)
2. Human errors (e.g. sensor placement, sensor settings, faulty/inadequate calibration, unit conversions, round-off and data conversion errors)
3. Any occurrence of unexpected processes, modifications and events in the monitored urban water systems, either controlled or uncontrolled (i.e. pipe bursts,
flooded pump station, maintenance of a filter at a treatment plant)
Untreated or mismanaged data strongly impacts the service operation, as it
propagates deeper [11] into the decision-making process and leads to erroneous or
ill-informed decisions regarding system operation, organizational mistrust, reduced
service efficiency and, eventually, customer dissatisfaction [12]. The detection and
identification of the aforementioned errors can be carried out with a variety of
methods that include threshold, data-driven and model-based approaches, as further
discussed in Sect. 3.
1.2 Scope and Approach
This chapter aims at providing a bird’s-eye view of data validation in the drinking
water industry of the Netherlands towards better Data Quality Control (DQC)
policies in the drinking water sector, by providing insights on (raw) data validation
in two problem types, one of water quantity and one of water quality. The focus of
this chapter is on a specific aspect of the overall DQC chain, which deals with faulty
data detection and isolation (FDI). Furthermore, of interest are errors in the measurements, because sensing and human data editing process lead to raw data
distortion in the form of, e.g. drift, bias, precision degradation or sensor failure
[13]. Mapping this focal point to the typology of errors seen in Sect. 1.1, it becomes
evident that this chapter focuses only on errors of type (i) and type (ii),
i.e. measurement and human errors. Moreover, the focus lies on data validation to
determine faulty data and the identification techniques, without expanding further on
the decision-making process regarding to accepting or rejecting the faulty data.
As a first step, in order to identify the needs of the industry and its current
practice, an inventory of current applications regarding DQC within the water
companies was conducted. Visits or interviews with four water companies took
2 Given this definition, any outliers or anomalies in data owing to natural rare and/or extreme events,
including very low probability cases such as black swans [9], should not be considered as faulty
data due to errors that have to be corrected.
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
67
system.
2 These anomalies, otherwise known as errors, can be further distinguished in
three types [10]:
1. Measurement errors (e.g. failure of data registration, maintenance problems,
drifts, bias, strong gradients, lack of redundancy, problems of coherence at both
local and global scale, duplication of data)
2. Human errors (e.g. sensor placement, sensor settings, faulty/inadequate calibration, unit conversions, round-off and data conversion errors)
3. Any occurrence of unexpected processes, modifications and events in the monitored urban water systems, either controlled or uncontrolled (i.e. pipe bursts,
flooded pump station, maintenance of a filter at a treatment plant)
Untreated or mismanaged data strongly impacts the service operation, as it
propagates deeper [11] into the decision-making process and leads to erroneous or
ill-informed decisions regarding system operation, organizational mistrust, reduced
service efficiency and, eventually, customer dissatisfaction [12]. The detection and
identification of the aforementioned errors can be carried out with a variety of
methods that include threshold, data-driven and model-based approaches, as further
discussed in Sect. 3.
1.2 Scope and Approach
This chapter aims at providing a bird’s-eye view of data validation in the drinking
water industry of the Netherlands towards better Data Quality Control (DQC)
policies in the drinking water sector, by providing insights on (raw) data validation
in two problem types, one of water quantity and one of water quality. The focus of
this chapter is on a specific aspect of the overall DQC chain, which deals with faulty
data detection and isolation (FDI). Furthermore, of interest are errors in the measurements, because sensing and human data editing process lead to raw data
distortion in the form of, e.g. drift, bias, precision degradation or sensor failure
[13]. Mapping this focal point to the typology of errors seen in Sect. 1.1, it becomes
evident that this chapter focuses only on errors of type (i) and type (ii),
i.e. measurement and human errors. Moreover, the focus lies on data validation to
determine faulty data and the identification techniques, without expanding further on
the decision-making process regarding to accepting or rejecting the faulty data.
As a first step, in order to identify the needs of the industry and its current
practice, an inventory of current applications regarding DQC within the water
companies was conducted. Visits or interviews with four water companies took
2 Given this definition, any outliers or anomalies in data owing to natural rare and/or extreme events,
including very low probability cases such as black swans [9], should not be considered as faulty
data due to errors that have to be corrected.
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
67
