# Compute complete-linkage agglomerative clustering
spe.ch.complete <- hclust(spe.ch, method = "complete")
plot(spe.ch.complete,
labels = rownames(spe),
main = "Chord - Complete linkage")
Given that these are sites along a river (with the numbers following the stream),
does this result tend to place sites that are neighbours along the river in the same
groups?
How can two perfectly valid clustering methods produce such different results
when applied to the same data?
The comparison between the two dendrograms (Figs. 4.1 and 4.2) shows the
difference in the philosophy and the results of the two methods: single linkage
allows an object to agglomerate easily to a group, since a link to a single object of
the group suffices to induce fusion. This is a “closest friend” procedure, so to say.
The resulting dendrogram does not always show clearly separated groups, but can be
used to identify gradients in the data. At the opposite, complete linkage clustering is
much more contrasting. A group admits a new member only at a dissimilarity
corresponding to the furthest object of the group: one could say that the admission
requires unanimity of the members of the group. It follows that the larger a group is,
the more difficult it is to agglomerate with it. Complete linkage, therefore, tends to
produce many small separate groups, which tend to be rather spherical in multivariate space and agglomerate at large distances. Therefore, this method is interesting to
search for and identify discontinuities in data.
4.4 Average Agglomerative Clustering
This family comprises four methods that are based on average dissimilarities among
objects or on centroids of clusters. The differences among them are in the way of
computing the positions of the groups (arithmetic average versus centroids) and in
the weighting or non-weighting of the groups according to the number of objects
they contain when computing fusion levels. Table 4.1 summarizes their names and
properties.
The best-known method of this family, UPGMA, allows an object to join a group
at the mean of the dissimilarities between this object and all members of the group.
When two groups join, they do it at the mean of the dissimilarities between all
members of one group and all members of the other. Let us apply it to our data
(Fig. 4.3):
4.4 Average Agglomerative Clustering
65
Précédent

- 78/444

Suivant