values. Once a distribution is fitted, the quartiles of the data can be obtained and a
hypothesis test can be performed to identify values which fall outside the statistical
boundaries. If the test has statistical significance, the value can be registered as an
outlier, implying that is considered as faulty data. This approach can supplement the
CFPD method described above.
Regression Analysis, Correlation Checks and Principal Component Analysis
(PCA)
Statistical regression encompasses a large number of techniques that aim at fitting a
line or curve to a data and using it as a basis to determine outliers and – potentially –
faulty data points, as well as to predict/infer missing data, based on a subset of ‘good’
data points [46–49]. Such a regression model in principle includes two (2) components: (a) predictors or explanatory variables, which form the basis of the prediction
and (b) response variables, which can be predicted based on the explanatory variables.
5 Multivariate regression models can be also employed in case there are
multiple explanatory variables for the water production. Techniques when data are
sampled in unevenly distributed intervals also exist [50] and can be of use to water
utilities, where some variables are stored irregularly at specific events and subsequently interpolated prior to storage in databases/data warehouses. On the assessment side, multiple metrics
6 to assess the successful adjustment of the regression line
or curve exist [51], quantifying whether the model represents the response variable
as a function of the explanatory variables accurately. Once the regression model is
fitted, any large error between an actual data point and the estimation can be flagged
as a faulty data point.
The main issue of relevance to regression models for data validation is to properly
select the explanatory variables for a certain response variable. This issue is related
with the concept of data diet (Fig. 5), in the sense of looking for a correlation
between different variables by (1) deciding which correlations are possible by using
physical and domain knowledge (expertise of drinking water) and (2) employing
statistical correlation techniques used by data scientists.
7 The techniques used for the
selection of explanatory variables are known as input variable selection (IVS) and
5 For example, one may be interested in the relation between the monthly water production [m
3
] of a
utility and the total energy consumption [kWh] used for treatment, transmission and distribution. If
there is a missing or suspect faulty data in water production, then this value can be estimated based
on the total energy consumption of the utility (explanatory variable).
6 Examples include the Root Mean Square Error (RMSE), Coefficient of Determination (R
2
), NashSutcliffe Efficiency (NSE), Kling Gupta Efficiency (KGE). For a successful regression model,
RMSE should have low values (close to 0.0), while for efficiency metrics a value near 1.0 is
optimal.
7 Domain knowledge is always relevant in the water sector, as a SCADA, water data warehouse or
database can exceed a thousand variables (n v ~
1000Þ and thus a combination of response variables
that exceeds n dd ¼
nv nvÀ1
ð
Þ
2
~
500:000 values.
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
81
Précédent

- 99/357

Suivant