16
P. Muñoz et al.
we split the length of the data for training and test. For the Tomebamba catchment, training run from Jan/2015 to May/2017 and test from Feb/2017 to Jan/2019.
Whereas, for the Yanuncay catchment, training run from Jan/2016 to July/2017 and
test from Jan/2015 to Jan/2016 to capture more extreme events in the calibration
period.
In both catchments, flash-floods events might occur when consecutive precipitation events saturate the upper parts, and a further event (not necessarily extreme)
triggers a rapid response. It might also occur that a very intense event (above 100 mm
hour
−1 ) saturates and produces a flash-flood event regardless of antecedent soil saturation conditions in the catchment. On the other hand, hydrological droughts cause
water shortage cases due to the dependence on the Tomebamba and Yanuncay rivers.
2.3 Methodology
2.3.1 Random Forest
Random Forest (RF) is a supervised ML algorithm that ensembles a multitude of
decorrelated decision trees (DTs) voting for the most popular class (classification)
or the mean prediction of the individual trees (regression). In practice, a DT (or
particular model) is a hierarchical analysis based on a set of conditions consecutively
applied to a dataset. To assure decorrelation, a bagging technique is employed by
the RF algorithm for growing DTs from different randomly resampled training sets
obtained from the original dataset. Specifically, for regression problems, each DT
provides an independent numerical output of the phenomenon of interest (i.e., runoff),
contrary to class labels for classification applications.
In short, starting from the parent node of a DT, the RF algorithm splits each node
of a tree into two self-similar child nodes according to simple rules related to the
data and until a stopping criterion is reached. A node is split by randomly selecting
a number of features rather than using all of them. For this, a random component is
used to resample and to determine the optimal successive features (directions) for
splitting the data in order to obtain purer nodes than its parent one. This is aimed to
homogenize the outcomes of a single DT. At the end, every terminal node represents
a regression model that applies in that very node only. A complete description of the
RF functioning can be found in (Breiman 2017, 2001).
Some of the advantages of employing the RF algorithm are, for instance, the
reduced amount of parameters to be tuned when compared with other ML techniques and the use of a bagging process which hardens robustness and therefore
improves model accuracy. The RF algorithm can deal with both small and large size
samples, high dimensionality, and complex data structures (Biau and Scornet 2016).
Breiman’s original RF algorithm has already been implemented in programming
languages such as R® and Python®; we selected the Python language through the
P. Muñoz et al.
we split the length of the data for training and test. For the Tomebamba catchment, training run from Jan/2015 to May/2017 and test from Feb/2017 to Jan/2019.
Whereas, for the Yanuncay catchment, training run from Jan/2016 to July/2017 and
test from Jan/2015 to Jan/2016 to capture more extreme events in the calibration
period.
In both catchments, flash-floods events might occur when consecutive precipitation events saturate the upper parts, and a further event (not necessarily extreme)
triggers a rapid response. It might also occur that a very intense event (above 100 mm
hour
−1 ) saturates and produces a flash-flood event regardless of antecedent soil saturation conditions in the catchment. On the other hand, hydrological droughts cause
water shortage cases due to the dependence on the Tomebamba and Yanuncay rivers.
2.3 Methodology
2.3.1 Random Forest
Random Forest (RF) is a supervised ML algorithm that ensembles a multitude of
decorrelated decision trees (DTs) voting for the most popular class (classification)
or the mean prediction of the individual trees (regression). In practice, a DT (or
particular model) is a hierarchical analysis based on a set of conditions consecutively
applied to a dataset. To assure decorrelation, a bagging technique is employed by
the RF algorithm for growing DTs from different randomly resampled training sets
obtained from the original dataset. Specifically, for regression problems, each DT
provides an independent numerical output of the phenomenon of interest (i.e., runoff),
contrary to class labels for classification applications.
In short, starting from the parent node of a DT, the RF algorithm splits each node
of a tree into two self-similar child nodes according to simple rules related to the
data and until a stopping criterion is reached. A node is split by randomly selecting
a number of features rather than using all of them. For this, a random component is
used to resample and to determine the optimal successive features (directions) for
splitting the data in order to obtain purer nodes than its parent one. This is aimed to
homogenize the outcomes of a single DT. At the end, every terminal node represents
a regression model that applies in that very node only. A complete description of the
RF functioning can be found in (Breiman 2017, 2001).
Some of the advantages of employing the RF algorithm are, for instance, the
reduced amount of parameters to be tuned when compared with other ML techniques and the use of a bagging process which hardens robustness and therefore
improves model accuracy. The RF algorithm can deal with both small and large size
samples, high dimensionality, and complex data structures (Biau and Scornet 2016).
Breiman’s original RF algorithm has already been implemented in programming
languages such as R® and Python®; we selected the Python language through the
