that, contrary to most constrained ordination methods described in Chap. 6, where
the selection of explanatory variables is made on the basis of explanatory power,
MRT focuses on prediction, making it a very interesting tool for practical applications, as in environmental management. The focus on prediction is embedded in the
method, as will become clear below.
MRT is a powerful and robust method, which can handle a wide variety of
situations, even those where some values are missing, and where the relationships
between the response and explanatory variables are nonlinear, or where high-order
interactions among explanatory variables are present.
4.12.2 Computation (Principle)
The computation of a MRT consists in two procedures running together:
(1) constrained partitioning of the data and (2) cross-validation of the results. Let
us first briefly explain the two procedures. After that we will see how they are
applied together to produce a model that has the form of a decision tree.
4.12.2.1 Constrained Partitioning of the Data
• For each explanatory variable, produce all possible partitions of the sites into two
groups. For a quantitative variable, this is done by sorting the sites according to
the ordered values of the variable, and repeatedly splitting the series after the first,
second... (n À 1)
th object. For a categorical variable, allocate the objects to two
groups, screening all possible combinations of levels. In all cases, compute the
resulting sum of within-group sums of squared distances to the group means
(within-group SS) for the response data. Retain the solution minimizing this
quantity, along with the identity and value of the explanatory variable or the
level of the categorical variable producing the partition retained.
• Repeat the same procedure within each of the two subgroups retained above; in
each group, retain the best partition along with the corresponding explanatory
variable and its threshold value.
• Continue within all partitions until all objects form their own group or until a
preselected smallest number of objects per group is reached. At that point, select
the tree with size (number of groups) appropriate to the aims of the study. For
studies with a predictive objective, cross-validation, which is a procedure to
identify the best predictive tree, is developed below.
• Apart from the number and composition of the leaves, an important characteristic
of a tree is its relative error (RE), i.e., the sum of the within-group SS over all
leaves divided by the overall SS of the data. In other words, this is the fraction of
variance not explained by the tree. Without cross-validation, among the successive partitioning levels, one would retain the solution minimizing RE; this would
4.12 Multivariate Regression Trees (MRT): Constrained Clustering
127
Précédent

- 140/444

Suivant