par(mfrow = c(1, 2))
spe.ch.mvpart data.matrix(spe.norm) ~ .,
env,
margin = 0.08,
cp = 0,
xv = "pick",
xval = nrow(spe),
xvmult = 100
)
summary(spe.ch.mvpart)
printcp(spe.ch.mvpart)
Hint Argument xv = "pick" allows one to interactively pick a tree among those
proposed. If one prefers that the tree with the minimum CVRE be automatically
chosen, then use xv = "min".
If argument xv ¼ "pick" has been used, which we recommend, one left-clicks
on the point representing the desired number of groups. A tree is then drawn. Here
we decided to pick the 4-group solution. While not the absolute best, it still ranks
among the good ones and avoids producing too many small groups (Fig. 4.27, right).
The tree produced by this analysis is rich in information. Apart from the general
statistics appearing at the bottom of the plot (residual error, i.e. the one-complement
of the R
2 of the model; cross-validated error; standard error), the following features
are important:
• Each node is characterized by a threshold value of an explanatory variable. For
instance, the first node splits the data into two groups of 16 and 13 sites on the
basis of elevation. The critical value (here 341 m) is often not found among the
data; it is the mean of the two values delimiting the split. If two or more
explanatory variables lead to equal results, an arbitrary choice is made among
them. In this example, for instance, variable dfs with value 204.8 km would
yield the same split.
• Each leaf (terminal group) is characterized by its number of sites and its RE as
well as by a small barplot representing the abundances of the species (in the same
order as is the response data matrix). Although difficult to read if there are many
species, these plots show that the different groups are indeed characterized by
different species. A more formal statistical approach is to search for characteristic
or indicator species (Sect. 4.11). See below for an example.
• The tree can be used to allocate a new observation to one of the groups on the
basis of the values of the relevant environmental variables. “Relevant” means
here that the variables needed to allocate an object may differ depending on the
4.12 Multivariate Regression Trees (MRT): Constrained Clustering
131
Précédent

- 144/444

Suivant