Although only a small fraction of the data is validated, all of these data are
potentially still used as input for different types of models, e.g. reliability analysis of
predictions of systems performance under certain scenarios. However, this is dangerous as spurious data supplied to models will provide spurious results of model
simulations.
Significant differences in the resolution of different parameters were also
observed defined by the frequency with which they are stored in utility databases.
For example, the year of installation of pipes usually has a resolution of 1 year. On
the other hand, water quality data for turbidity can be stored with timestamps every
5 s in a database. However, each company has its own temporal resolutions for each
variable.
There is still a lack of knowledge about how to first select the appropriate
technique for validating a given dataset and secondly how to fine tune the parameters
of each validation technique, to minimize the effort between identification of
possible anomalies and extreme events in the system with high accuracy. All utilities
face the same challenges when defining flags for data validation:
• Too many flags, (false-positives) operatives become reluctant to perform verification and data validation, and the trust on the data validation diminishes.
• Too few flags and robustness is lost, as water companies would not be able to
differentiate between a regular event and extreme event and a real anomaly.
• Changes in the system is not always recorded or archived in the historic data. This
is a challenge when the system configurations change dynamically (e.g. in the
case of Company C).
5 Discussion and Recommendations for Future Work
5.1 Introduction
Data quality is a key consideration for the reliable functioning of drinking water
systems, as data are used to monitor and operate systems, to bill customers, to report
the performance of the company and to feed different types of models. Improving the
quality of the data and making it more accessible will benefit every department of a
company.
The following sections discuss a number of recommendations and future work in
the field of DQC for drinking water utilities.
5.2 Recommendation Regarding Future Work
During the survey of the water companies, it has been evident that most utilities
apply diverse methods of DQC. One of the features which is lacking across is a
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
101
potentially still used as input for different types of models, e.g. reliability analysis of
predictions of systems performance under certain scenarios. However, this is dangerous as spurious data supplied to models will provide spurious results of model
simulations.
Significant differences in the resolution of different parameters were also
observed defined by the frequency with which they are stored in utility databases.
For example, the year of installation of pipes usually has a resolution of 1 year. On
the other hand, water quality data for turbidity can be stored with timestamps every
5 s in a database. However, each company has its own temporal resolutions for each
variable.
There is still a lack of knowledge about how to first select the appropriate
technique for validating a given dataset and secondly how to fine tune the parameters
of each validation technique, to minimize the effort between identification of
possible anomalies and extreme events in the system with high accuracy. All utilities
face the same challenges when defining flags for data validation:
• Too many flags, (false-positives) operatives become reluctant to perform verification and data validation, and the trust on the data validation diminishes.
• Too few flags and robustness is lost, as water companies would not be able to
differentiate between a regular event and extreme event and a real anomaly.
• Changes in the system is not always recorded or archived in the historic data. This
is a challenge when the system configurations change dynamically (e.g. in the
case of Company C).
5 Discussion and Recommendations for Future Work
5.1 Introduction
Data quality is a key consideration for the reliable functioning of drinking water
systems, as data are used to monitor and operate systems, to bill customers, to report
the performance of the company and to feed different types of models. Improving the
quality of the data and making it more accessible will benefit every department of a
company.
The following sections discuss a number of recommendations and future work in
the field of DQC for drinking water utilities.
5.2 Recommendation Regarding Future Work
During the survey of the water companies, it has been evident that most utilities
apply diverse methods of DQC. One of the features which is lacking across is a
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
101
