schist (Birimian), volcanite (basalt, andesite, rhyolite (Birimian)), granite, distance
to mineral deposits, and distance to granitoid rocks were the most important arseniccontamination predictors in the studied region (Bretzler et al. 2017). Podgorski and
his colleagues have developed a logistic regression model using 12 geochemical
predictors and 1200 groundwater samples in Pakistan (Podgorski et al. 2017). They
found that in there, potential evapotranspiration, precipitation, aridity, irrigated area,
slope, soil organic carbon, soil pH, and Holocene fluvial sediments were the most
important predictors of arsenic presence (Podgorski et al. 2017). In a recent study,
the authors developed a boosted regression tree machine learning model using
74 features of physical or geochemical triggers of arsenic contamination from
3283 water samples in the north-central USA (Erickson et al. 2018). They found
that 33 of the 74 variables served as essential arsenic predictors in the model
(Erickson et al. 2018).
It is apparent from this collection of studies and models that hydrogeological and
topographic factors play significant roles in predicting groundwater arsenic contamination and that logistic regression has been a preferred method of developing
arsenic-prediction models.
4.5 Application of AI in Predicting Fluoride Contamination
of Groundwater
Amini et al. (2008a, b) developed the first global model of fluoride contamination of
groundwater. The model is based on 60,000 groundwater samples collected from
25 countries around the world. Thirty-one variables were used to develop probabilistic models, applying linear regression and adaptive neuro-fuzzy inference system
(ANFIS) techniques (Table 4.5) (Amini et al. 2008b).
Nadiri et al. (2013), using a minimal dataset of the hydro-chemical properties of
132 groundwater samples collected over 4 years (2004–2008), developed ML
models: a Sugenofuzzylogic, a Mamdanifuzzylogic, an ANN, a neuro-fuzzy, and a
committee machine with artificial intelligence (CMAI) model (Nadiri et al. 2013). In
2017, Barzegar et al. (2017) incorporated additional samples into the same dataset to
develop extreme learning machine (ELM), multilayer perceptron (MLP), and SVM
models. Recently, Podgorski et al. (2018) have developed multinomial logistic
regression and random forest models based on geological, climatic, and soil features
of 12,600 groundwater samples collected from all over India. They found 15 of
25 variables to be the most crucial fluoride predictors in the studied area (Podgorski
et al. 2018). In another recent study, the authors developed instance-based k-nearest
neighbors, locally weighted learning, M5P, and regression-by-discretization models
on a small dataset of 143 groundwater samples using hydro-chemical properties of
the water samples (Khosravi et al. 2019).
It is readily apparent that all of these models except for one are based on a
minimal dataset (Barzegar et al. 2017; Khosravi et al. 2019; Nadiri et al. 2013) and
4 Application of Artificial Intelligence in Predicting Groundwater Contaminants
97
to mineral deposits, and distance to granitoid rocks were the most important arseniccontamination predictors in the studied region (Bretzler et al. 2017). Podgorski and
his colleagues have developed a logistic regression model using 12 geochemical
predictors and 1200 groundwater samples in Pakistan (Podgorski et al. 2017). They
found that in there, potential evapotranspiration, precipitation, aridity, irrigated area,
slope, soil organic carbon, soil pH, and Holocene fluvial sediments were the most
important predictors of arsenic presence (Podgorski et al. 2017). In a recent study,
the authors developed a boosted regression tree machine learning model using
74 features of physical or geochemical triggers of arsenic contamination from
3283 water samples in the north-central USA (Erickson et al. 2018). They found
that 33 of the 74 variables served as essential arsenic predictors in the model
(Erickson et al. 2018).
It is apparent from this collection of studies and models that hydrogeological and
topographic factors play significant roles in predicting groundwater arsenic contamination and that logistic regression has been a preferred method of developing
arsenic-prediction models.
4.5 Application of AI in Predicting Fluoride Contamination
of Groundwater
Amini et al. (2008a, b) developed the first global model of fluoride contamination of
groundwater. The model is based on 60,000 groundwater samples collected from
25 countries around the world. Thirty-one variables were used to develop probabilistic models, applying linear regression and adaptive neuro-fuzzy inference system
(ANFIS) techniques (Table 4.5) (Amini et al. 2008b).
Nadiri et al. (2013), using a minimal dataset of the hydro-chemical properties of
132 groundwater samples collected over 4 years (2004–2008), developed ML
models: a Sugenofuzzylogic, a Mamdanifuzzylogic, an ANN, a neuro-fuzzy, and a
committee machine with artificial intelligence (CMAI) model (Nadiri et al. 2013). In
2017, Barzegar et al. (2017) incorporated additional samples into the same dataset to
develop extreme learning machine (ELM), multilayer perceptron (MLP), and SVM
models. Recently, Podgorski et al. (2018) have developed multinomial logistic
regression and random forest models based on geological, climatic, and soil features
of 12,600 groundwater samples collected from all over India. They found 15 of
25 variables to be the most crucial fluoride predictors in the studied area (Podgorski
et al. 2018). In another recent study, the authors developed instance-based k-nearest
neighbors, locally weighted learning, M5P, and regression-by-discretization models
on a small dataset of 143 groundwater samples using hydro-chemical properties of
the water samples (Khosravi et al. 2019).
It is readily apparent that all of these models except for one are based on a
minimal dataset (Barzegar et al. 2017; Khosravi et al. 2019; Nadiri et al. 2013) and
4 Application of Artificial Intelligence in Predicting Groundwater Contaminants
97
