minimal CVRE value plus one standard error of the CVRE values. This is called
“the 1 SE rule”.
• To obtain an estimate of the error of this process, run it a large number of times
(100 or 500 times) with other random assignments of the objects to k groups.
• The final tree retained is the one showing most of the smallest CVRE values over
all permutations, or the one respecting most often the 1 SE rule.
• In mvpart, the smallest possible number of objects in one group is equal to
ceiling(log2(n)) when argument minauto ¼ TRUE. Setting
minauto ¼ FALSE allows the partitioning to proceed until the data set is
fully split into individual observations.
In MRT analysis, the computations of sums-of-squares (SS) are done in Euclidean space. To account for special characteristics of the data, they can be
pre-transformed prior to being submitted to the procedure. Pre-transformations for
species data prior to their analysis by Euclidean-based methods are presented in Sect.
2.2.4. For response data that are environmental descriptors with different physical
units, the variables should be standardized before they are used in MRT.
4.12.3 Application Using Packages mvpart
and MVPARTwrap
As of this writing, the only package implementing a complete and handy version of
MRT is mvpart. Unfortunately, this package is no longer supported by the R Core
Team, so that no updates are available for R versions posterior to R 3.0.3. Nevertheless, mvpart can still be installed on more recent versions of R by applying the
following code.
# On Windows machines, Rtools (3.4 and above) must be installed
# first. Go to: https://cran.r-project.org/bin/windows/Rtools/
# After that (for Windows and MacOS), type:
install.packages("devtools")
library(devtools)
install_github("cran/mvpart", force = TRUE)
install_github("cran/MVPARTwrap", force = TRUE)
Package mvpart has been written to compute MRT, using univariate regression
trees computed by a function called rpart(). Its use requires that the response data
belong to class ‘matrix’ and the explanatory variables to class ‘data frame’. The
relationship is written as a formula of the same type as those used in regression
functions (see ?lm). The example below shows the simplest implementation, where
one uses all the variables contained in the explanatory data frame.
4.12 Multivariate Regression Trees (MRT): Constrained Clustering
129
Précédent

- 142/444

Suivant