Abstract During the last decades, the role of data as a vital resource that enhances
decision-making and supports efficient systems operation has become evident, with
a growing number of companies viewing data as a key organizational aspects that
has to be properly managed, instead of an operational side-product. At the same
time, drinking water systems increase in complexity and feature smart sensors,
which in turn leads to data-richer operation environments for the water services.
Given this challenging context, the often-overlooked factor of ensuring high data
quality and preventing errors in data streams becomes increasingly important. In this
chapter the current data validation techniques, challenges and best practices of the
Dutch drinking water companies is presented.
Keywords Anomaly detection, Best practices, Data quality control,
Hydroinformatics, Netherlands
1 Introduction
1.1 Background
Despite the emerging need for holistic, efficient data management policies,
implementing a proper Data Quality Control (DQC) strategy is generally a
non-trivial task, as the protocols and techniques used are process- and contextdependent. For the water sector, protocols to standardize data acquisition and
analysis are being developed for different parts of the water cycle, targeting the
data streams of specific processes. For example, management frameworks in the
context of urban hydrology and sewer systems have been developed [1, 2], as well as
initiatives for European ocean and sea data management [3].
In the Netherlands, since 2012 a protocol for data quality is being developed by
KWR Water Research Institute and the Netherlands Organisation for Applied
Scientific Research (TNO) for registration of groundwater levels and hydraulic
heads [4–6]. A consequence of such protocol is the development of a validation
tool named Menyantes1 [7] for which a new version of this software, named
HydroMonitor, is currently under development. In the drinking water sector, a
uniform failure registration database (USTORE) has been used by eight Dutch
drinking water companies [8]. This initiative started in 2001, and it has been an
ongoing process of continuous improvement, with both the complexity of the
registration and data requirements increasing over the years. In 2017, a protocol to
guarantee data quality was included in the PCD (Praktijkcode Drinkwater
1 ) no. 9:
‘Uniform failure registration’. In 2018, this PCD was released [8].
Within these protocols, one of the core ways of improving data quality is by
performing data validation. Data validation or, in other words, fault detection and
isolation (FDI) refers to the identification and handling of anomalies and outliers in
1 https://www.praktijkcodesdrinkwater.nl/
66
M. Castro-Gama et al.
decision-making and supports efficient systems operation has become evident, with
a growing number of companies viewing data as a key organizational aspects that
has to be properly managed, instead of an operational side-product. At the same
time, drinking water systems increase in complexity and feature smart sensors,
which in turn leads to data-richer operation environments for the water services.
Given this challenging context, the often-overlooked factor of ensuring high data
quality and preventing errors in data streams becomes increasingly important. In this
chapter the current data validation techniques, challenges and best practices of the
Dutch drinking water companies is presented.
Keywords Anomaly detection, Best practices, Data quality control,
Hydroinformatics, Netherlands
1 Introduction
1.1 Background
Despite the emerging need for holistic, efficient data management policies,
implementing a proper Data Quality Control (DQC) strategy is generally a
non-trivial task, as the protocols and techniques used are process- and contextdependent. For the water sector, protocols to standardize data acquisition and
analysis are being developed for different parts of the water cycle, targeting the
data streams of specific processes. For example, management frameworks in the
context of urban hydrology and sewer systems have been developed [1, 2], as well as
initiatives for European ocean and sea data management [3].
In the Netherlands, since 2012 a protocol for data quality is being developed by
KWR Water Research Institute and the Netherlands Organisation for Applied
Scientific Research (TNO) for registration of groundwater levels and hydraulic
heads [4–6]. A consequence of such protocol is the development of a validation
tool named Menyantes1 [7] for which a new version of this software, named
HydroMonitor, is currently under development. In the drinking water sector, a
uniform failure registration database (USTORE) has been used by eight Dutch
drinking water companies [8]. This initiative started in 2001, and it has been an
ongoing process of continuous improvement, with both the complexity of the
registration and data requirements increasing over the years. In 2017, a protocol to
guarantee data quality was included in the PCD (Praktijkcode Drinkwater
1 ) no. 9:
‘Uniform failure registration’. In 2018, this PCD was released [8].
Within these protocols, one of the core ways of improving data quality is by
performing data validation. Data validation or, in other words, fault detection and
isolation (FDI) refers to the identification and handling of anomalies and outliers in
1 https://www.praktijkcodesdrinkwater.nl/
66
M. Castro-Gama et al.
