a perfect agreement among the two analyses a Yes-Yes cell must contain all faulty
data flags for both DQC schemes.
• In the case that the current Company A’s DQC is unable to identify a faulty data
compared to the one proposed, a No-Yes coincidence is identified.
• In the case that the proposed method is able to identify faulty data, while the
current validation of Company A is not able to do it, a Yes-No coincidence is
identified.
It is also of interest that for all variables the number of faulty data identified with
the proposed validation is larger for this analysis than the number of flags obtained
with the current one of Company A. There are two possible reasons for this. Either
the proposed data validation is more compact and sensitive to faulty data or there is a
need to improve the current data validation of the utility. In the first case, this is a
disadvantage for the operatives, as this will translate to a large number of verifications required. For example, temperature data in WWTP I confirms more than 1,200
flags for a verification during a 2-year period, or almost two flags per day. As it is
Table 4 Confusion matrices of water quality data for company A’S and proposed DQC
pH
KWR
Temperature
KWR
Yes
No
Total
Yes
No Total
Company A
Yes
25
9
34
Company A
Yes
29
4
33
No
254
-
No
1250
-
Total
279
Total
1279
Turbidity II
KWR
Turbidity I
KWR
Yes
No
Total
Yes
No Total
Company A
Yes
61
6
67
Company A
Yes
20
10
30
No
187
-
No
39
-
Total
248
Total
59
The significance of the colors is linked between this table and Figs. 7, 8, 9 and 10. The blue color
and yellow color are represented both in this table and in Figs. 7, 8, 9 and 10. This means that the
cases represented in the figures correspond to specific cases of Yes–No and No–Yes which are
presented in this table
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
87
data flags for both DQC schemes.
• In the case that the current Company A’s DQC is unable to identify a faulty data
compared to the one proposed, a No-Yes coincidence is identified.
• In the case that the proposed method is able to identify faulty data, while the
current validation of Company A is not able to do it, a Yes-No coincidence is
identified.
It is also of interest that for all variables the number of faulty data identified with
the proposed validation is larger for this analysis than the number of flags obtained
with the current one of Company A. There are two possible reasons for this. Either
the proposed data validation is more compact and sensitive to faulty data or there is a
need to improve the current data validation of the utility. In the first case, this is a
disadvantage for the operatives, as this will translate to a large number of verifications required. For example, temperature data in WWTP I confirms more than 1,200
flags for a verification during a 2-year period, or almost two flags per day. As it is
Table 4 Confusion matrices of water quality data for company A’S and proposed DQC
pH
KWR
Temperature
KWR
Yes
No
Total
Yes
No Total
Company A
Yes
25
9
34
Company A
Yes
29
4
33
No
254
-
No
1250
-
Total
279
Total
1279
Turbidity II
KWR
Turbidity I
KWR
Yes
No
Total
Yes
No Total
Company A
Yes
61
6
67
Company A
Yes
20
10
30
No
187
-
No
39
-
Total
248
Total
59
The significance of the colors is linked between this table and Figs. 7, 8, 9 and 10. The blue color
and yellow color are represented both in this table and in Figs. 7, 8, 9 and 10. This means that the
cases represented in the figures correspond to specific cases of Yes–No and No–Yes which are
presented in this table
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
87
