correlation) highlight the same strongest associations between species and selected
groups. However, it may be interesting to display all the phi values to identify
possible sets of environmental conditions (represented by the groups, if these have
been defined on the basis of environmental variables) that are avoided by some
species (with the caveat mentioned above):
round(iva.phi$str, 3)
For instance, the bleak (Alal) shows a very strong negative association
(ϕ ¼ À0.92) with the group of sites “1 + 2”, which mirrors its strongest positive
association (ϕ ¼ 0.92) with the group of sites “3 + 4”. Indeed, the bleak is absent
from all but 3 sites of groups 1 + 2, and is present in all sites of groups 3 + 4.
Obviously, the bleak prefers the lower regions of the river and avoids the higher
ones. The ecological characteristics of the various sections of the Doubs River are
described and analysed in various sections of this book.
It is also possible to compute bootstrap confidence intervals around the phi
values, using the same strassoc() function as was used to compute the IndVal
coefficient:
iva.phi.boot <- strassoc(spe, grps, func = "r.g", nboot = 1000)
Since the phi values can be negative or positive, you can use this opportunity to
obtain a bootstrap test of significance of the phi values: all values where both
confidence limits have the same sign (i.e., CI not encompassing 0) are deemed
significant. This complements the permutational tests for the negative values. For
instance, the avoidance of the bullhead (Cogo) of the conditions prevailing in group
1 is significant in that sense, the CI interval being [À0.313, À0.209].
4.12 Multivariate Regression Trees (MRT): Constrained
Clustering
4.12.1 Introduction
Multivariate regression trees (MRT; De’ath 2002) are an extension of univariate
regression trees, a method allowing the recursive partitioning of a quantitative
response variable under the control of a set of quantitative or categorical explanatory
variables (Breiman et al. 1984). Such a procedure is sometimes called constrained or
supervised clustering. The result is a tree whose “leaves” (terminal groups of sites)
are composed of subsets of sites chosen to minimize the within-group sums of
squares (as in a k-means clustering), but where each successive partition is defined
by a threshold value or a state of one of the explanatory variables. Among the
numerous potential solutions in terms of group composition and number of leaves,
one usually retains the one that has the best predictive power. This stresses the fact
126
4 Cluster Analysis
Précédent

- 139/444

Suivant