• Checking the duration between sensor maintenance and anticipating operational
downtime periods.
• Maintaining and consulting a repository with data from past network failure
events, such as a Pipe Failure Data (PFD) repository, which is of use to update
physically based models and clarify certain anomalies which otherwise would be
identified as faulty data [87]. Standardization and integration of PFD repositories
among water companies helps streamline this task [88, 89].
• Combining expert-based judgment with ad hoc validation tools built for a specific
part and function of the water system [7].
4 Application of Data Validation
4.1 Overview of the Data and Selection of the Techniques
In order to apply and test diverse data validation techniques, two problems in the
field of drinking water distribution were identified:
• Anomaly detection in volume flow rates and energy use
• Anomaly detection in datasets of temperature, turbidity and pH.
In Table 2, the overview of the cases and the data validation techniques used
during the 2018 survey are presented where only the proposed data validation
applied to Company A’s data is discussed in this chapter.
The data provided by the water utilities contain several differences. In terms of
variables, water quality data tends to be more homogeneous. Data resolution was
also variable and depends on the type of registration for each utility. It can vary in
minutes, quarters (15 min), events (when significant change occurs) and pulses.
Timestamp registration is also very heterogeneous across utilities, dates can
contain summer and winter time as number or other indicator (+1.00 or +0.00),
and some information is consistently absent on the same timestamps, most likely due
to data communication. In this regard, at 00:00, Company A contains no data,
indicating that the data transfer is most likely to occur at this time at night. Data
from one utility was provided as a Last In, First Out (LIFO) format (reverse dates)
for some variables, while the same utility provided a more common First In, First
Out (FIFO) format.
Length of time series was also very variable as it was not possible to obtain more
than a few months of data for Company B. Some sensors have recently started to
send live data. Due to the high variability in data types and content, it was not
possible to perform a one size fits all analysis of data validation and more specifically
of faulty data detection, so the focus given to specific datasets is variable to present a
broader set of analyses within the funding research instrument BTO.
9 In each
9 BTO: Bedrijftaak Onderzoek. It corresponds to research within the consortium of the ten drinking
water companies from the Netherlands and one Belgian drinking water company.
84
M. Castro-Gama et al.
Précédent

- 102/357

Suivant