Are the two results fairly similar? Which object(s) is (or are) classified
differently?
To compute k-means partitioning from dissimilarity indices that cannot be
obtained by a transformation of the raw data followed by calculation of the Euclidean distance, for example the percentage difference (aka Bray-Curtis) dissimilarity,
one has to compute first a rectangular data table with n rows by principal coordinate
analysis (PCoA, Sect. 5.5) of the dissimilarity matrix, then use that rectangular table
as input to k-means partitioning. For the percentage difference dissimilarity, one has
to compute the PCoA on the square root of the percentage difference dissimilarities
to obtain a fully Euclidean solution, or use a PCoA function that provides a
correction for negative eigenvalues. These points are discussed in Chap. 5.
A partitioning yields a single partition with a predefined number of groups. If you
want to try several solutions with different k values, you must rerun the analysis. But
which solution is the best in terms of number of clusters? To answer this question,
one has to state what “best” means. Many criteria exist; some of them are available in
the function clustIndex() of the package cclust. Milligan and Cooper
(1985) recommend maximizing the Calinski-Harabasz index (F-statistic comparing
the among-group to the within-group sum of squares of the partition), although its
value tends to be lower for unequal-sized partitions. The maximum of ‘ssi’ (“Simple
Structure Index”, see the documentation file of clustIndex() for details) is
another good indicator of the best partition in the least-squares sense.
Fortunately, one can avoid running kmeans() many times by hand. vegan’s
function cascadeKM() is a wrapper for the kmeans() function, that is, a
function that uses a basic function, adding new properties to it. It creates several
partitions forming a cascade from small (argument inf.gr) to large values of
k (argument sup.gr). Let us apply this function to our dataset, asking for 2–10
groups and the simple structure index criterion for clustering quality, followed by a
plot of the results (Fig. 4.20).
# k-means partitioning, 2 to 10 groups
spe.KM.cascade spe.norm,
inf.gr = 2,
sup.gr = 10,
iter = 100,
criterion = "ssi"
)
summary(spe.KM.cascade)
plot(spe.KM.cascade, sortg = TRUE)
98
4 Cluster Analysis
Précédent

- 111/444

Suivant