computing model (Chapi et al. 2017; Miraki et al. 2019; Tien Bui et al. 2019) whose
performance has been confirmed in studies of groundwater-potential mapping (Elío
et al. 2017; Mair and El-Kadi 2013; Ozdemir 2011; Rizeei et al. 2018).
The current study also concludes that for mapping nitrate contamination of
groundwater, among the applied machine learning algorithms of recent years, the
random forest approach has predominated. The RF is a tree-based algorithm that is
extremely flexible (Muchlinski et al. 2016). It exchanges a high degree of variance
between each tree for a low bias in predicting the outcome variable. If the assumptions of the modeling process, including linearity of variables, collinearity, and
homoscedasticity, are not remarkable, other methods that are not regression-based
may provide better estimations. However, if the training dataset of an event is rare
and crucial, the flexibility of algorithms such as RF can outperform LR models
(Muchlinski et al. 2016). A study carried out by Rodriguez-Galiano et al. (2014)
indicated that the RF model had superior goodness-of-fit and prediction accuracy,
when compared with the LR model, for nitrate vulnerability mapping of the Vega de
Granada aquifer in the southern part of Iberia (southeast of Granada City). The
authors stated that the application of RF models is simple and that their results are
interpretable. Eventually, they summarized the advantages of the RF model for
nitrate vulnerability mapping and groundwater studies as follows:
• It can learn the complicated, nonlinear relationship between the independent and
dependent variables.
• The different types of variables can be incorporated for analysis because the RF
model does not need any assumption, whereas the LR model does.
• A large number of training datasets can be efficiently handled by the RF model
without variable deletion.
• It can estimate the predictive power of variables.
• It can produce an internal unbiased estimate of the prediction (out-of-bag) error.
• It is partly robust and sturdily resistant to outliers and spurious datasets.
• The computation time of RF model is lower than that of other machine learning
methods, including ANN or SVM. In other words, it is a high-predictive-accuracy
algorithm.
• Concerning noise, it is most robust than other machine learning models.
According to the above, it can be claimed that AI techniques for classification
tasks are diverse, and that therefore, they should be tested on a given groundwatercontamination dataset and their results compared. The application of various AI
techniques in predicting the presence of groundwater contaminants is a promising
approach but a recent phenomenon. Groundwater contamination can be triggered by
a myriad of environmental as well as anthropogenic processes. Geographic,
hydrogeologic, geological, hydrochemical, land use, and socioeconomic factors at
play in areas with contaminated water have vital roles in inducing the release of
specific contaminants into the groundwater. The above review also shows that
contamination is a regional phenomenon, so global prediction models may not be
appropriate for most groundwater contaminants. An array of ML techniques, including but not limited to regression and classification algorithms has been used to
100
S. K. Singh et al.
performance has been confirmed in studies of groundwater-potential mapping (Elío
et al. 2017; Mair and El-Kadi 2013; Ozdemir 2011; Rizeei et al. 2018).
The current study also concludes that for mapping nitrate contamination of
groundwater, among the applied machine learning algorithms of recent years, the
random forest approach has predominated. The RF is a tree-based algorithm that is
extremely flexible (Muchlinski et al. 2016). It exchanges a high degree of variance
between each tree for a low bias in predicting the outcome variable. If the assumptions of the modeling process, including linearity of variables, collinearity, and
homoscedasticity, are not remarkable, other methods that are not regression-based
may provide better estimations. However, if the training dataset of an event is rare
and crucial, the flexibility of algorithms such as RF can outperform LR models
(Muchlinski et al. 2016). A study carried out by Rodriguez-Galiano et al. (2014)
indicated that the RF model had superior goodness-of-fit and prediction accuracy,
when compared with the LR model, for nitrate vulnerability mapping of the Vega de
Granada aquifer in the southern part of Iberia (southeast of Granada City). The
authors stated that the application of RF models is simple and that their results are
interpretable. Eventually, they summarized the advantages of the RF model for
nitrate vulnerability mapping and groundwater studies as follows:
• It can learn the complicated, nonlinear relationship between the independent and
dependent variables.
• The different types of variables can be incorporated for analysis because the RF
model does not need any assumption, whereas the LR model does.
• A large number of training datasets can be efficiently handled by the RF model
without variable deletion.
• It can estimate the predictive power of variables.
• It can produce an internal unbiased estimate of the prediction (out-of-bag) error.
• It is partly robust and sturdily resistant to outliers and spurious datasets.
• The computation time of RF model is lower than that of other machine learning
methods, including ANN or SVM. In other words, it is a high-predictive-accuracy
algorithm.
• Concerning noise, it is most robust than other machine learning models.
According to the above, it can be claimed that AI techniques for classification
tasks are diverse, and that therefore, they should be tested on a given groundwatercontamination dataset and their results compared. The application of various AI
techniques in predicting the presence of groundwater contaminants is a promising
approach but a recent phenomenon. Groundwater contamination can be triggered by
a myriad of environmental as well as anthropogenic processes. Geographic,
hydrogeologic, geological, hydrochemical, land use, and socioeconomic factors at
play in areas with contaminated water have vital roles in inducing the release of
specific contaminants into the groundwater. The above review also shows that
contamination is a regional phenomenon, so global prediction models may not be
appropriate for most groundwater contaminants. An array of ML techniques, including but not limited to regression and classification algorithms has been used to
100
S. K. Singh et al.
