234
M. Gevrey . S. Lek· T.Oberdorff
12.2
Contribution of Environmental Variables
In the Multi-Linear Regression (MLR) model, the influence of each variable can
be roughly assessed by checking the final values of the regression coefficients.
However, it is more difficult with ANNs to find the contribution of the input
variables directly from the models, specific algorithms are necessary to use.
Most authors have used the principle of step by step elimination of the variables
to determine this influence (Balls et al. 1996; Maier and Dandy 1996) but others
have tried different methods using connection weights (Garson 1991; Goh 1995),
the perturbation of input variables (Scardi 1996), the partial derivatives of the
output according to the input variables (Dimopoulos et al. 1995; 1999) or the
successive study of the variables by a variation of one of them while the others are
fixed to a determined value (Lek et al. 1996a,b), etc. The two methods retained for
this study are the Profile method (Lek et al. 1996a,b) and the PaD method
(Dimopoulos et al. 1999).
The "Profile algorithm"
This method was suggested by Lek et al (1996a, b). The general idea is to study
each input variable successively, to do so the others are blocked during the
utilisation of the model. The principle of this algorithm is to construct a fictitious
matrix considering the range of all input variables. In greater detail, the values of
each variable are divided into 12 values at equal intervals between their minimum
and maximum values. For all variables except one, the 12 values are set at their
minimum values, then successively their first quartile, median, third quartile and
maximum. For each studied variable, five values for each of the 12 points are
obtained. These five values are reduced to the median value. Then the profile of
the output variable can be plotted for 12 values levels of the variable considered.
The same calculations can then be repeated for each of the other variables. For
each variable a curve is then obtained, wh ich gives a set of profiles of the variation
of the dependent variable according to the increase of the input variables.
The "Pad algorithm"
This method gives two different results, the first one being a profile of the output
for each input variable and the second being a classification of the variable in
increasing order of importance.
"The derivatives profile"
This is a sensitivity analysis proposed by Dimopoulos et al. (1995; 1999). It is
based on the principle of partial derivatives of the ANN response with respect to
each descriptor. When the input x j is modified, the output Y j changes, yj=J( x). Then
the sensitivity of the network outputs according to small input perturbations can be
studied, which is represented by the Jacobian matrix dy / dx T = [ßy /tJx lxn.
For a network with n inputs, one hidden layer with ni nodes, and one output (i.e.
m=I), the gradient vector of Y j with respect to x j is d j = [d jl ,K ,d je ,K ,d jn r
(Dimopoulos et al. 1995), with:
M. Gevrey . S. Lek· T.Oberdorff
12.2
Contribution of Environmental Variables
In the Multi-Linear Regression (MLR) model, the influence of each variable can
be roughly assessed by checking the final values of the regression coefficients.
However, it is more difficult with ANNs to find the contribution of the input
variables directly from the models, specific algorithms are necessary to use.
Most authors have used the principle of step by step elimination of the variables
to determine this influence (Balls et al. 1996; Maier and Dandy 1996) but others
have tried different methods using connection weights (Garson 1991; Goh 1995),
the perturbation of input variables (Scardi 1996), the partial derivatives of the
output according to the input variables (Dimopoulos et al. 1995; 1999) or the
successive study of the variables by a variation of one of them while the others are
fixed to a determined value (Lek et al. 1996a,b), etc. The two methods retained for
this study are the Profile method (Lek et al. 1996a,b) and the PaD method
(Dimopoulos et al. 1999).
The "Profile algorithm"
This method was suggested by Lek et al (1996a, b). The general idea is to study
each input variable successively, to do so the others are blocked during the
utilisation of the model. The principle of this algorithm is to construct a fictitious
matrix considering the range of all input variables. In greater detail, the values of
each variable are divided into 12 values at equal intervals between their minimum
and maximum values. For all variables except one, the 12 values are set at their
minimum values, then successively their first quartile, median, third quartile and
maximum. For each studied variable, five values for each of the 12 points are
obtained. These five values are reduced to the median value. Then the profile of
the output variable can be plotted for 12 values levels of the variable considered.
The same calculations can then be repeated for each of the other variables. For
each variable a curve is then obtained, wh ich gives a set of profiles of the variation
of the dependent variable according to the increase of the input variables.
The "Pad algorithm"
This method gives two different results, the first one being a profile of the output
for each input variable and the second being a classification of the variable in
increasing order of importance.
"The derivatives profile"
This is a sensitivity analysis proposed by Dimopoulos et al. (1995; 1999). It is
based on the principle of partial derivatives of the ANN response with respect to
each descriptor. When the input x j is modified, the output Y j changes, yj=J( x). Then
the sensitivity of the network outputs according to small input perturbations can be
studied, which is represented by the Jacobian matrix dy / dx T = [ßy /tJx lxn.
For a network with n inputs, one hidden layer with ni nodes, and one output (i.e.
m=I), the gradient vector of Y j with respect to x j is d j = [d jl ,K ,d je ,K ,d jn r
(Dimopoulos et al. 1995), with:
