2 Application of a Machine Learning Technique for Developing Short-Term Flood …
17
scikit-learn package for ML (Pedregosa et al. 2011). Details of the package can be
found online in https://scikit-learn.org.
2.3.2 Algorithm
The RF algorithm for regression can be summarized as follows:
1. Grow every DT based on a random selection of a number of bootstrap samples
(n_estimators parameter) drawn with (or without) replacement from the training
dataset. To compose each bootstrap, a different subset (roughly two-thirds) of
the dataset is selected in a process known as out-of-bag (OOB).
2. Split each parent node into two descendant ones by using the best split criteria.
The best split decision is taken by considering only a number of features
(max_features parameter) of the total number of predictor variables of the dataset
(n_features).
3. Construct n_trees as much as possible (largest extent) by repeating steps (1) and
(2). This must be done until reaching a defined number of nodes in the tree. The
depth of each tree is controlled by the max_depth and the min_samples_leaf
parameters; where min_ samples_ leaf is the minimum number of samples
required to be at a leaf node.
4. Determine the outcome of the RF model as the mean response from all DTs.
Some additional indications when using the RF algorithm:
– The OBB technique is intended to achieve unbiased estimates of the regression and
to estimate the importance of the features used for the tree construction process
(Probst et al. 2018).
– Conditioning max_ features to be lower than n_ features ensures the nonexistence
of duplicated DTs in the forest. It aims to mitigate overfitting. For regression
problems, Breiman (2001) recommends to set max_ features equal to the root
square of n_ features.
– The depth of each tree is controlled to reduce the structural complexity of the trees
(models). This is known as pruning criteria (Rodriguez-Galiano et al. 2014).
– Determination of the best splits are chosen based on the mean squared error (MSE)
for regression problems. The minimum number of samples required to split a node
is controlled by the min_ samples_ split parameter.
– The optimal number of trees is reached when the OOB error stops decreasing
significantly (depends on the research objective).
2.3.3 RF Hyper-Parameterization
Model hyper-parameters determines the structure of the forest and its level of randomness (Probst et al. 2018). Although the algorithm can be run with default parameters,
17
scikit-learn package for ML (Pedregosa et al. 2011). Details of the package can be
found online in https://scikit-learn.org.
2.3.2 Algorithm
The RF algorithm for regression can be summarized as follows:
1. Grow every DT based on a random selection of a number of bootstrap samples
(n_estimators parameter) drawn with (or without) replacement from the training
dataset. To compose each bootstrap, a different subset (roughly two-thirds) of
the dataset is selected in a process known as out-of-bag (OOB).
2. Split each parent node into two descendant ones by using the best split criteria.
The best split decision is taken by considering only a number of features
(max_features parameter) of the total number of predictor variables of the dataset
(n_features).
3. Construct n_trees as much as possible (largest extent) by repeating steps (1) and
(2). This must be done until reaching a defined number of nodes in the tree. The
depth of each tree is controlled by the max_depth and the min_samples_leaf
parameters; where min_ samples_ leaf is the minimum number of samples
required to be at a leaf node.
4. Determine the outcome of the RF model as the mean response from all DTs.
Some additional indications when using the RF algorithm:
– The OBB technique is intended to achieve unbiased estimates of the regression and
to estimate the importance of the features used for the tree construction process
(Probst et al. 2018).
– Conditioning max_ features to be lower than n_ features ensures the nonexistence
of duplicated DTs in the forest. It aims to mitigate overfitting. For regression
problems, Breiman (2001) recommends to set max_ features equal to the root
square of n_ features.
– The depth of each tree is controlled to reduce the structural complexity of the trees
(models). This is known as pruning criteria (Rodriguez-Galiano et al. 2014).
– Determination of the best splits are chosen based on the mean squared error (MSE)
for regression problems. The minimum number of samples required to split a node
is controlled by the min_ samples_ split parameter.
– The optimal number of trees is reached when the OOB error stops decreasing
significantly (depends on the research objective).
2.3.3 RF Hyper-Parameterization
Model hyper-parameters determines the structure of the forest and its level of randomness (Probst et al. 2018). Although the algorithm can be run with default parameters,
