32
P. Muñoz et al.
In this study, we examined the feasibility to develop precipitation-runoff forecasting models and their ability to forecast extreme high (floods) and low (drought)
flows. The applicability of this study relies in the difficulty to obtain key-spatial
information explaining short-term flow processes in mountainous regions such as the
Tropical Andes in South America. The machinery for constructing such models is
based on a data-driven approach, the Random Forest algorithm. This novel technique
has gained popularity in water-related studies in the past decades.
We utilized two comparable catchments to validate a methodology aimed to
develop parsimonious flash-flood forecasting models based on the Random Forest
algorithm. We extended the analysis to evaluate extreme low flows forecasting. Additionally, we firstly performed a randomized grid search procedure to calibrate the
hyper-parameter of all models; however, we found that optimal sets of parameters for a given lead time can be transferred to a comparable catchment. In this
regard, we found a minimum reduction of model performance for RF models using
imported hyper-parameters in contrast with the high computational cost required for
recalibration activities.
We could also validate the effectiveness of a feature selection technique (based on
output’s variance) aimed to reduce the complexity and dimension of the input before
training a model. Retaining only the features accounting for 80% of the model’s
variance did not compromise forecasts but rather optimized computation times. Only
slight differences in NSE coefficients proved that the selection of the most important
features was successfully achieved.
Generally, it seems conclusive that is more difficult to forecast floods than
droughts. From a data-driven perspective, this is occasioned by imbalance data problems (i.e., the number of independent events for peak flows are much scarce than
the quantity of low flow events). Although the RF algorithm is capable to deal with
this issue, the reduced number of peak flow events was not enough for the models to
properly learn from data.
As expected, the ability of the RF models to forecast extreme high values decreased
as the lead time increased. For extreme low flows, the magnitude of the errors involved
are at a certain degree independent of the lead time selected (from 4 to 24 h).
Although we used only punctual precipitation and discharge records for building
up forecasting models, we motivate the use of spatial information and the inclusion
of other relevant variables involved in the flow generation process (e.g. soil moisture
and soil type maps). It is obvious how additional relevant information could significantly improve extreme high and low flows forecasting when using a ML technique
such as the RF algorithm. This information can be obtained by means of remote
sensing data. However, budget constraints in the Andean region and particularly in
Ecuador often limits its viability. As an alternative, we also suggest expanding the
rain gauge network to improve the representativeness of precipitation in mountain
catchments. This will improve, at a certain degree, model performances. However,
it must be taken into account that the major shortcoming of the use of rain gauges
in the Andean region is the occurrence of local rains due to complex topography.
Thus, an adequate representation of the spatial variability of precipitation is rarely
available for forecasting applications.
Précédent

- 38/353

Suivant