distribution-probability functions, the developed algorithms can be classified into
three categories: (i) Bayes-based algorithms (e.g., naïve Bayes and Bayes net),
(ii) functional algorithms (e.g., logistic regression, artificial neural network, support
vector machine), and (iii) decision-tree algorithms (e.g., random forest). The diversity of these algorithms is referred to as their applicability in classification problems.
Spatial prediction of environmental issues is related to two issues, which are
(i) model selection and (ii) factor/predictor selection and conditioning. In model
selection, users peruse existing models in search of one that is fitted to a training
dataset. There are neither standards or guidelines to clarify the optimal model for a
given task (Svetnik et al. 2003). This means that users may need to find the
applicable model through a trial-and-error approach, tuning parameters and running
algorithms on the training dataset and then comparing results to select the best fit.
What variables are appropriate to consider as predictors may vary according to data
availability, and differences in selected variables may create uncertainties during the
modeling process (Shirzadi et al. 2019). For example, a given LR model may return
different results in one case study than it would in another. This difference depends
on the fact that the variables defined in each case study are different, causing the
results to differ as well (Bui et al. 2018; Tien Bui et al. 2019).
Moreover, for a given case study, two models may return unlike results because
of differences in their probability-distribution functions. An LR model, for example,
might be fitted to the training dataset while an RF model might not. In other words,
the variables in a single studied area will be constant for all applied models, but in
this situation, the models would return different results. For this reason, a model’s
capacity to accurately predict is simply referred to as its ability to explain causal
processes, and there is enough room within the discipline for both standards of
evaluation (Shmueli 2010).
As per the above discussion and literature review, although some algorithms have
been developed for the classification task, few have been proposed or used for
groundwater contamination, including that of arsenic, fluoride, and nitrate. As we
have mentioned, the LR model has been more heavily relied upon for mapping
arsenic and fluoride contamination of groundwater. Although LR was employed as a
statistical machine learning technique in the 1990s to assess groundwater nitrate
vulnerability (Eckhardt and Stackelberg 1995; Tesoriero and Voss 1997), it has been
more recently used to map arsenic (Twarakavi and Kaluarachchi 2005;
Venkataraman 2010; Winkel et al. 2008) and fluoride-contamination resources
(Barzegar et al. 2017; Podgorski et al. 2018; Singh et al. 2013). The applicability
and usability of the LR model that makes it a promising alternative solution to the
mapping of groundwater contamination is derived from several advantages, including the ability to mathematically compute empirical weights for each variable, to
statistically remove the variables that lack predictive ability (via multi collinearity
tests), and to consider which variables most significantly affecting the results (at a
95% confidence interval) (Focazio 2002). Additionally, the obtained weights, which
are computed based on observed data, lead to the production of a reasonable, reliable
vulnerability map that can be used as an appropriate tool by water-resource managers
(Mair and El-Kadi 2013). The LR model has been considered a benchmark soft
4 Application of Artificial Intelligence in Predicting Groundwater Contaminants
99
Précédent

- 110/336

Suivant