Which method produces cophenetic distances (ordinate) that are well linearly
related to the original distances (abscissa)?
Another possible statistic for the comparison of clustering results is the Gower
(1983) distance
1 , computed as the sum of squared differences between the original
dissimilarities and cophenetic distances. The clustering method that produces the
smallest Gower distance may be seen as the one that provides the best clustering
model of the dissimilarity matrix. The cophenetic correlation and Gower distance
criteria do not always designate the same clustering result as the best.
# Gower (1983) distance
(gow.dist.single <- sum((spe.ch - spe.ch.single.coph) ^ 2))
(gow.dist.comp <- sum((spe.ch - spe.ch.comp.coph) ^ 2))
(gow.dist.UPGMA <- sum((spe.ch - spe.ch.UPGMA.coph) ^ 2))
(gow.dist.ward <- sum((spe.ch - spe.ch.ward.coph) ^ 2))
Hint Enclosing in brackets a command line producing an object induces the immediate
screen display of the object.
4.7.3 Looking for Interpretable Clusters
To interpret and compare clustering results, users generally look for interpretable
clusters. This means that a decision must be made: at what level should the
dendrogram be cut? Although it is not mandatory to select a single cutting level
for a whole dendrogram (some parts of the dendrogram may be interpretable at finer
levels than others), it is often practical to find one or a few levels where interpretations are made. These levels can be defined subjectively by visual examination of the
dendrogram, or they can be chosen to fulfil some criteria, such as a predetermined
number of groups for instance. In any case, adding information on the dendrograms
or plotting additional information about the clustering results can be very useful.
4.7.3.1 Graph of the Fusion Level Values
The fusion level values of a dendrogram are the dissimilarity values where a fusion
between two branches of a dendrogram occurs. Plotting the fusion level values may
1 This Gower distance is not to be confused with the Gower dissimilarity presented in Chap. 3.
74
4 Cluster Analysis
Précédent

- 87/444

Suivant