2 Application of a Machine Learning Technique for Developing Short-Term Flood …
21
2.4.1 Feature Reduction
In the pursuit of model parsimony, we reduced the input dimension of each model
through a process known as feature selection. It consists on the calculation of the
relative importance of each feature to the model´s output with the purpose to keep
only the most important features (defined criterion) for model construction. Apart
from shorten computation times, in some cases, feature selection even improves
model accuracy (Tang et al. 2007).
There is a number of techniques for feature selection, based on a variance sensitivity analysis, based on univariate statistical tests, recursive elimination, etc. We
employed the variance sensitivity analysis introduced by (Cortez 2010). This technique measures the output’s variance produced by a single feature alone. With this
approach, there is no consideration of the influence of features interaction. Thus, the
impact of each feature can be isolated and attributed to the feature itself. A relevant
feature to the model is to produce a higher output variance, therefore, the variance
(V k ) and its relative importance (R k ) are as follows:
V k =
L
j=1 (
y t−k ( j) −
y t−k ( j))
2
L − 1
R k =
V k
m
i=1 V i
x100
where, y
t−k is the model output obtained by holding all m input features at their
average values except y
t−k , which varies according to the sample or time step, along
the interval j ∈ {1, . . . , L}.
The criterion to obtain parsimonious models is to retain the most relevant features
until reaching, at least, 80% of the total relative importance. The remaining features
can be considered unimportant and therefore we trimmed them off from the model’s
input.
Figure 2.2 summarizes the methodology proposed for the construction and evaluation of flood and drought forecasting models. It is an extension of the methodology
proposed by Muñoz et al. (2018) for flood forecasting. We applied this scheme for
each lead time and for each catchment, separately. In short, we start by using an
autoregressive (discharge) model as the base forecasting model. The input is then
enriched with precipitation information to improve extreme flow forecasting. Finally,
we reduced the dimension of the input through a feature selection process to retain
only the most relevant information. Model hyper-parameterization is performed for
each model (input data scenario).
Each model forecasts runoff for a defined timestep, and at the end, sequential forecasts results in a forecasted runoff timeseries. The forecasted and observed timeseries
are ultimately used to perform a detailed model assessment specifically focused on
floods and droughts.
Précédent

- 27/353

Suivant