have broader applications in engineering and environmental sciences [48, 52,
53]. Such techniques apply correlation and stepwise selection of explanatory variables to perform the identification of significant dependencies and are based on two
main techniques:
(a) Correlation analysis, which explores auto- and cross-correlation structures
based on time series analysis [54]. A multitude of techniques exists in both
stochastic and deterministic time series analysis [55] and analytical models exist
that include long- scale dependency [56, 57], persistence [58–60] and more
sophisticated models such as ARIMA/ARIMAX [61–63]. In water quality
analysis, correlation has been extensively applied for open and pressurized
flow when data validation is required [64, 65].
(b) Principal component analysis (PCA), as its name states, has been applied mainly
for determining variables which may project the data into its principal components. The data is transformed using an orthogonal transformation and then
converted into a set of variables uncorrelated among themselves, which in turn
denominates the principal components. Particular applications for PCA in the
water sector are to support the identification of consumption patterns and leakage
detection [66].
3.2.3 Data-Driven Models
Another approach that features methods heavily relying on the dataset itself is called
data-driven modelling (DDM) [67, 68]. These methods originate from the computer
science fields of computational- or artificial-intelligence (CoAI) and machine learning (ML) and are used for data exploration, data mining [69] and also classification.
The latter case aims at building classifier models
8 based on a large number of
independent (predictor) variables and is of use to data validation. Data-driven
techniques, among others, include:
• Decision trees (DecT), which can be considered the simplest technique to perform
classification of large datasets [70], seeing use in the prediction of leakages and
breaks for WDN models [71]
• Support vector machines (SVM), which are non-linear regression algorithms over
multi-dimensional input spaces that have been used for leakages and demand
pattern identification [72–74]
• Artificial neural networks (ANN), which are machine learning algorithms
inspired by biological neural networks and used for classification. In the water
sector, ANN have been applied for the identification of losses and leakages in
8 Classifier models (or classifiers) output a categorical variable (e.g. with values ‘1’ or ‘0’, that can
be flags for a data validation problem), based on a (large) number of input variables, that can be for
instance geophysical or water system time series.
82
M. Castro-Gama et al.
53]. Such techniques apply correlation and stepwise selection of explanatory variables to perform the identification of significant dependencies and are based on two
main techniques:
(a) Correlation analysis, which explores auto- and cross-correlation structures
based on time series analysis [54]. A multitude of techniques exists in both
stochastic and deterministic time series analysis [55] and analytical models exist
that include long- scale dependency [56, 57], persistence [58–60] and more
sophisticated models such as ARIMA/ARIMAX [61–63]. In water quality
analysis, correlation has been extensively applied for open and pressurized
flow when data validation is required [64, 65].
(b) Principal component analysis (PCA), as its name states, has been applied mainly
for determining variables which may project the data into its principal components. The data is transformed using an orthogonal transformation and then
converted into a set of variables uncorrelated among themselves, which in turn
denominates the principal components. Particular applications for PCA in the
water sector are to support the identification of consumption patterns and leakage
detection [66].
3.2.3 Data-Driven Models
Another approach that features methods heavily relying on the dataset itself is called
data-driven modelling (DDM) [67, 68]. These methods originate from the computer
science fields of computational- or artificial-intelligence (CoAI) and machine learning (ML) and are used for data exploration, data mining [69] and also classification.
The latter case aims at building classifier models
8 based on a large number of
independent (predictor) variables and is of use to data validation. Data-driven
techniques, among others, include:
• Decision trees (DecT), which can be considered the simplest technique to perform
classification of large datasets [70], seeing use in the prediction of leakages and
breaks for WDN models [71]
• Support vector machines (SVM), which are non-linear regression algorithms over
multi-dimensional input spaces that have been used for leakages and demand
pattern identification [72–74]
• Artificial neural networks (ANN), which are machine learning algorithms
inspired by biological neural networks and used for classification. In the water
sector, ANN have been applied for the identification of losses and leakages in
8 Classifier models (or classifiers) output a categorical variable (e.g. with values ‘1’ or ‘0’, that can
be flags for a data validation problem), based on a (large) number of input variables, that can be for
instance geophysical or water system time series.
82
M. Castro-Gama et al.
