techniques was broad, a similar content is expected in the case of data correction
techniques.
Data Reconciliation
The decision which must be made by utilities about which data to trust and which
data not to trust is a constant struggle. Often, during the selection of data and case
studies, a common concern was voiced by operatives of utilities about not trusting
the data of a region or DMA. However, no quantification or metric was made
available to express this in any of the cases and to our knowledge such metrics are
not available in the Dutch context. This means that expert knowledge has been
applied by utility’s experts, with prior experience in the management of their system.
However, such expert knowledge is currently only encapsulated in the minds of
experts. There are several methodologies available to transform such qualitative
decision making into quantifiable rules for the determination of likelihood of data as
being faulty. This can lead to improvements in data collection and reduce dependence of utilities on specific experts if they are not available.
5.2.2 Selection of Faulty Detection Techniques
Regarding DQC, there is no one technique fits all purposes. Depending on the
monitored event/variable, different techniques with different parameters should be
applied, also according the objective of the validation.
Software tools cannot validate data by themselves. Expert knowledge is always
needed for a good determination of the parameters to perform the data validation,
especially in complex systems, such as drinking water systems.
Given the variability of datasets, number of records, timestamps, time resolution
and variables, only simple techniques for data validation have been applied for the
water utilities data presented as case studies in this chapter. So it was recommended
to the participating utilities that, in a future project, similar techniques to be specifically calibrated for subsets of data from the same variables. For example, in this
chapter a short analysis of data validation for water balance is presented, but an
extensive literature on the matter is also available. This means that the possibility to
increase the identification of faulty data can be explored with many more techniques
than the ones presented here on that subject. Once this is done, then a proper
selection of best practice techniques for water balance can be established for the
sector of Dutch drinking water companies.
5.2.3 Modelling for Anomaly Identification
In this manuscript, only a short portion of techniques for data validation was
explored, and only simple tests were implemented. One of the biggest obstacles in
104
M. Castro-Gama et al.
techniques.
Data Reconciliation
The decision which must be made by utilities about which data to trust and which
data not to trust is a constant struggle. Often, during the selection of data and case
studies, a common concern was voiced by operatives of utilities about not trusting
the data of a region or DMA. However, no quantification or metric was made
available to express this in any of the cases and to our knowledge such metrics are
not available in the Dutch context. This means that expert knowledge has been
applied by utility’s experts, with prior experience in the management of their system.
However, such expert knowledge is currently only encapsulated in the minds of
experts. There are several methodologies available to transform such qualitative
decision making into quantifiable rules for the determination of likelihood of data as
being faulty. This can lead to improvements in data collection and reduce dependence of utilities on specific experts if they are not available.
5.2.2 Selection of Faulty Detection Techniques
Regarding DQC, there is no one technique fits all purposes. Depending on the
monitored event/variable, different techniques with different parameters should be
applied, also according the objective of the validation.
Software tools cannot validate data by themselves. Expert knowledge is always
needed for a good determination of the parameters to perform the data validation,
especially in complex systems, such as drinking water systems.
Given the variability of datasets, number of records, timestamps, time resolution
and variables, only simple techniques for data validation have been applied for the
water utilities data presented as case studies in this chapter. So it was recommended
to the participating utilities that, in a future project, similar techniques to be specifically calibrated for subsets of data from the same variables. For example, in this
chapter a short analysis of data validation for water balance is presented, but an
extensive literature on the matter is also available. This means that the possibility to
increase the identification of faulty data can be explored with many more techniques
than the ones presented here on that subject. Once this is done, then a proper
selection of best practice techniques for water balance can be established for the
sector of Dutch drinking water companies.
5.2.3 Modelling for Anomaly Identification
In this manuscript, only a short portion of techniques for data validation was
explored, and only simple tests were implemented. One of the biggest obstacles in
104
M. Castro-Gama et al.
