2 Application of a Machine Learning Technique for Developing Short-Term Flood …
25
Fig. 2.7 Model results of the parsimonious 4-h discharge forecasting model of the Tomebamba
catchment (test period)
importance of the model. For the 4-h forecasting model of the Tomebamba catchment,
we found that including 9 lags from each precipitation station and 8 discharge lags
would be enough to achieve 80.36% of the total relative importance. The percentage
of reduction of model features was 58%. We did the same for each case; the percentage
of reduction ranged from 40 to 70%, and from 40 to 60% for all forecasting models
of the Tomebamba and Yanuncay catchments, respectively.
Table 2.2 summarizes the input data composition and total number of features
utilized for the RF forecasting models for all cases. Whereas, Table 2.3 presents the
obtained model performances in terms of the NSE coefficient. Notice that the NSE
coefficients were calculated for the whole spectrum of flows. Table 2.3 also contrasts
the NSE coefficients of the full-input and their parsimonious model obtained through
a feature selection process. Results prove that NSE coefficients obtained from the
parsimonious models do not differ significantly from the correspondent full-input
version of the model (maximum differences in calibration and validation of 0.01 and
0.02 for the Tomebamba and Yanuncay catchments). In some cases, parsimonious
models even outperformed their correspondent full-input models.
Regarding RF model overfitting, we found maximum differences between the
NSE coefficients of the calibration versus the validation period of 0.20 and 0.64 for
the Tomebamba and Yanuncay catchments, respectively. Significant differences for
the Yanuncay RF models can be partly explained by the fact that a shorter calibration
25
Fig. 2.7 Model results of the parsimonious 4-h discharge forecasting model of the Tomebamba
catchment (test period)
importance of the model. For the 4-h forecasting model of the Tomebamba catchment,
we found that including 9 lags from each precipitation station and 8 discharge lags
would be enough to achieve 80.36% of the total relative importance. The percentage
of reduction of model features was 58%. We did the same for each case; the percentage
of reduction ranged from 40 to 70%, and from 40 to 60% for all forecasting models
of the Tomebamba and Yanuncay catchments, respectively.
Table 2.2 summarizes the input data composition and total number of features
utilized for the RF forecasting models for all cases. Whereas, Table 2.3 presents the
obtained model performances in terms of the NSE coefficient. Notice that the NSE
coefficients were calculated for the whole spectrum of flows. Table 2.3 also contrasts
the NSE coefficients of the full-input and their parsimonious model obtained through
a feature selection process. Results prove that NSE coefficients obtained from the
parsimonious models do not differ significantly from the correspondent full-input
version of the model (maximum differences in calibration and validation of 0.01 and
0.02 for the Tomebamba and Yanuncay catchments). In some cases, parsimonious
models even outperformed their correspondent full-input models.
Regarding RF model overfitting, we found maximum differences between the
NSE coefficients of the calibration versus the validation period of 0.20 and 0.64 for
the Tomebamba and Yanuncay catchments, respectively. Significant differences for
the Yanuncay RF models can be partly explained by the fact that a shorter calibration
