Sites are ordered to match the two dendrogramsas well as possible. Colours
highlight common clusters, whereas sites printed in black have different positions
in the two trees. Can you recognize particularly “robust” clusters?
4.7.3.3 Multiscale Bootstrap Resampling
Unconstrained cluster analysis is a family of heuristic methods that are not aimed at
testing a priori hypotheses. However, natural variation leads to sampling variability,
and the results of a classification are likely to reflect it. It is therefore legitimate to
assess the uncertainty (or its counterpart the robustness) of a classification. This has
been done abundantly in phylogenetic analysis.
The choice approach for such validation procedures is bootstrap resampling
(e.g. Efron 1979; Felsenstein 1985), which consists in randomly sampling subsets
of the data and computing the clustering on these subsets. After having done this a
large number of times, one counts the proportion of the replicate clustering results
where a given cluster appears. This proportion is called the bootstrap probability
(BP) of the cluster. Multiscale bootstrap resampling has been developed as an
enhancement to answer some criticisms about the classical bootstrap procedure
(Efron et al. 1996; Shimodaira 2002, 2004). In this method bootstrap samples of
several different sizes are used to estimate the p-value of each cluster. This improvement produces “approximately unbiased” (AU) p-values. Readers are referred to the
original publications for more details.
The pvclust package (Suzuki and Shimodaira 2006) provides functions to plot
a dendrogram with bootstrap p-values associated to each cluster. AU p-values are
printed in red. The less accurate BP values are printed in green. Clusters with high
AU values (e.g. 0.95 or more) can be considered as strongly supported by the data.
Let us apply this analysis to the dendrogram obtained with the Ward method
(Fig. 4.10). Beware: in function pvclust() the data object must be transposed
with respect to our usual layout (rows are variables).
# Compute p-values for all clusters (edges) of the dendrogram
spech.pv
method.hclust = "ward.D2",
method.dist = "euc",
parallel = TRUE)
# Plot dendrogram with p-values
plot(spech.pv)
# Highlight clusters with high AU p-values
pvrect(spech.pv, alpha = 0.95, pv = "au")
lines(spech.pv)
pvrect(spech.pv, alpha = 0.91, border = 4)
4.7 Interpreting and Comparing Hierarchical Clustering Results
79
highlight common clusters, whereas sites printed in black have different positions
in the two trees. Can you recognize particularly “robust” clusters?
4.7.3.3 Multiscale Bootstrap Resampling
Unconstrained cluster analysis is a family of heuristic methods that are not aimed at
testing a priori hypotheses. However, natural variation leads to sampling variability,
and the results of a classification are likely to reflect it. It is therefore legitimate to
assess the uncertainty (or its counterpart the robustness) of a classification. This has
been done abundantly in phylogenetic analysis.
The choice approach for such validation procedures is bootstrap resampling
(e.g. Efron 1979; Felsenstein 1985), which consists in randomly sampling subsets
of the data and computing the clustering on these subsets. After having done this a
large number of times, one counts the proportion of the replicate clustering results
where a given cluster appears. This proportion is called the bootstrap probability
(BP) of the cluster. Multiscale bootstrap resampling has been developed as an
enhancement to answer some criticisms about the classical bootstrap procedure
(Efron et al. 1996; Shimodaira 2002, 2004). In this method bootstrap samples of
several different sizes are used to estimate the p-value of each cluster. This improvement produces “approximately unbiased” (AU) p-values. Readers are referred to the
original publications for more details.
The pvclust package (Suzuki and Shimodaira 2006) provides functions to plot
a dendrogram with bootstrap p-values associated to each cluster. AU p-values are
printed in red. The less accurate BP values are printed in green. Clusters with high
AU values (e.g. 0.95 or more) can be considered as strongly supported by the data.
Let us apply this analysis to the dendrogram obtained with the Ward method
(Fig. 4.10). Beware: in function pvclust() the data object must be transposed
with respect to our usual layout (rows are variables).
# Compute p-values for all clusters (edges) of the dendrogram
spech.pv
method.dist = "euc",
parallel = TRUE)
# Plot dendrogram with p-values
plot(spech.pv)
# Highlight clusters with high AU p-values
pvrect(spech.pv, alpha = 0.95, pv = "au")
lines(spech.pv)
pvrect(spech.pv, alpha = 0.91, border = 4)
4.7 Interpreting and Comparing Hierarchical Clustering Results
79
