be equivalent to retaining the solution maximizing the R
2 . However, this would
be an explanatory rather than predictive approach. De’ath (2002) states that “RE
gives an over-optimistic estimate of how accurately a tree will predict for new
data, and predictive accuracy is better estimated from the cross-validated relative error (CVRE)”.
4.12.2.2 Cross-Validation of the Partitions and Pruning of the Tree
At which level should we prune a tree, i.e., cut each of its branches so as to retain the
most sensible partition? To answer this question in a prediction-oriented manner,
one uses a subset of the objects (training set) to construct the tree, and the remaining
objects (test set) to validate the result by allocating them to the constructed
The measure of predictive error is the cross-validated relative error (CVRE). The
function is:
CVRE ¼
P ν
k¼1
P n
i¼1
P p
j¼1
y ij k
ð Þ À b y j k
ð Þ
2
P n
i¼1
P p
j¼1
À
y ij À
y j
Á 2
ð4:1Þ
where y ij(k) is one observation of the test set k, b y j k
ð Þ is the predicted value of one
observation in one leaf (centroid of the sites of that leaf), and the denominator
represents the overall dispersion (sum of squares) of the response data.
The CVRE can thus be defined as the ratio between the dispersion unexplained by
the tree (summed over the k test sets) divided by the overall dispersion of the
response data. Of course, the numerator changes after every partitioning event.
CVRE is 0 for perfect predictors and close to 1 for a poor set of predictors.
4.12.2.3 MRT Procedure
Now that we have both components of the methods, let us put them together to
explain the sequence of events of a cross-validated MRT run:
• Randomly split the data into k groups; by default k ¼ 10.
• Leave one of the k groups out and build a tree by constrained partitioning, the
decisions being made on the basis of the minimal within-group SS.
• Rerun the step above k À 1 times, leaving out each of the test groups in turn.
• In each of the k solutions above, and for each possible partition size (number of
groups) within these solutions, reallocate the test set. Compute the CVRE for all
partition sizes of the k solutions (one CVRE value per size). Eq. 4.1 encompasses
the computation for one level of partitioning and all k solutions.
• Pruning of the tree: retain the partition size for which the CVRE is smallest. An
alternative solution is to retain a smallest size for which the CVRE value is the
128
4 Cluster Analysis
Précédent

- 141/444

Suivant