was no possible explanation identified by the utility for such atypical pattern
variation.
Another possible use of this analysis for data validation is to perform a correlation
analysis of flows with respect to the energy use obtained from the utility. In fact, as
presented in Fig. 16b, the energy use pattern follows obviously a similar trend than
that of the flows. As a matter of fact, this is not distant from the current operation of
the system as a single District Metered Area (DMA), or fully interconnected WDN.
Similar to what was done in the case of flows of the five pumping stations. The
obtained histogram of the bivariate probability of energy use in City A is presented
in Fig. 16b. As it was the case of flows, there is a double peak in the energy
consumption (this is expected) and a higher probability during the MNF.
From a data validation perspective, it is of particular interest that in this case a
larger number of timestamps occur with low probability for energy consumptions
under the average pattern (grey values), particularly during the hours of
06:00–10:00. However, given that there is no indication of status or flag data for
any of these time series, it is not possible to conclude whether this is due to data
anomalies or to a regular operation of the system. As a hypothesis, most of these low
values below the trend of energy pattern can be attributed in part to the fact that there
is a huge number of pump switches for all PS’ as it is presented in Fig. 15. Of notice
is also that such pump switches identified in the energy consumption are not
represented all the time in the flow data, mainly because pump switches occur at a
1 min resolution, while energy is a cumulative variable stored every 15 min.
From a feedback session with personnel from the utility, it was possible to
determine that such behaviour of low energy registrations is possible. Sometimes
during the mornings energy is self-produced by the utility from renewables, and the
third party (energy company) is not aware of such energy influx. For that reason,
there is a deviation between water and energy used for pumping which is not
registered by the energy company. In this case, expert knowledge of the daily
operations and workings of the utility became far more relevant; otherwise all
such data would be flagged as faulty by a data validation system.
4.5 Best Practices and Issues in Data Validation Identified
in the Case Studies
The following results are obtained from interviews (questionnaires, personal communications and feedback sessions). These are resented based on the analysis of four
companies contributing data and not only based on the presented results of the
previous sections. A summary of best practices and issues is identified (Table 6)
and the issues identified during the pilot of this bird’s-eye view (Table 7).
A general issue that arises from the comparison of the four companies is the lack
of standards. There are different types of registration protocols and standards used by
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
99
variation.
Another possible use of this analysis for data validation is to perform a correlation
analysis of flows with respect to the energy use obtained from the utility. In fact, as
presented in Fig. 16b, the energy use pattern follows obviously a similar trend than
that of the flows. As a matter of fact, this is not distant from the current operation of
the system as a single District Metered Area (DMA), or fully interconnected WDN.
Similar to what was done in the case of flows of the five pumping stations. The
obtained histogram of the bivariate probability of energy use in City A is presented
in Fig. 16b. As it was the case of flows, there is a double peak in the energy
consumption (this is expected) and a higher probability during the MNF.
From a data validation perspective, it is of particular interest that in this case a
larger number of timestamps occur with low probability for energy consumptions
under the average pattern (grey values), particularly during the hours of
06:00–10:00. However, given that there is no indication of status or flag data for
any of these time series, it is not possible to conclude whether this is due to data
anomalies or to a regular operation of the system. As a hypothesis, most of these low
values below the trend of energy pattern can be attributed in part to the fact that there
is a huge number of pump switches for all PS’ as it is presented in Fig. 15. Of notice
is also that such pump switches identified in the energy consumption are not
represented all the time in the flow data, mainly because pump switches occur at a
1 min resolution, while energy is a cumulative variable stored every 15 min.
From a feedback session with personnel from the utility, it was possible to
determine that such behaviour of low energy registrations is possible. Sometimes
during the mornings energy is self-produced by the utility from renewables, and the
third party (energy company) is not aware of such energy influx. For that reason,
there is a deviation between water and energy used for pumping which is not
registered by the energy company. In this case, expert knowledge of the daily
operations and workings of the utility became far more relevant; otherwise all
such data would be flagged as faulty by a data validation system.
4.5 Best Practices and Issues in Data Validation Identified
in the Case Studies
The following results are obtained from interviews (questionnaires, personal communications and feedback sessions). These are resented based on the analysis of four
companies contributing data and not only based on the presented results of the
previous sections. A summary of best practices and issues is identified (Table 6)
and the issues identified during the pilot of this bird’s-eye view (Table 7).
A general issue that arises from the comparison of the four companies is the lack
of standards. There are different types of registration protocols and standards used by
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
99
