# Silhouette profile for k = 4 groups, k-means and PAM
par(mfrow = c(1, 2))
k <- 4
sil <- silhouette(spe.kmeans.g, spe.ch)
rownames(sil) <- row.names(spe)
plot(sil,
main = "Silhouette plot - k-means",
cex.names = 0.8,
col = 2:(k + 1))
plot(
silhouette(spe.ch.pam),
main = "Silhouette plot - PAM",
cex.names = 0.8,
col = 2:(k + 1)
)
On the basis on this plot, you should be able to tell which solution (PAMor kmeans) has a better silhouette profile.
You could also compare this result with the silhouette plot of the optimized Ward
clustering produced earlier. Except cluster numbering, is there any difference
between k-means and optimized Ward classifications? You can also address this
question by examining a contingency table:
# Compare classifications from k-means and optimized Ward
# clustering
table(spe.kmeans.g, spech.ward.gk)
Hint The PAM method is presented as “robust” because it minimizes a sum of
dissimilarities instead of a sum of squared Euclidean distances. It is also robust in
that it tends to converge to the same solution with a wide array of starting
medoids for a given k value; this does not guarantee, however, that the solution is
the most appropriate for a given research purpose.
This example shows that even two methods that are devoted to the same goal and
belong to the same general class (here non-hierarchical clustering) may provide
diverging results. It is up to the user to choose the one that yields classifications that
are bringing out more pertinent information or are more closely interpretable using
environmental variables (next section).
106
4 Cluster Analysis
par(mfrow = c(1, 2))
k <- 4
sil <- silhouette(spe.kmeans.g, spe.ch)
rownames(sil) <- row.names(spe)
plot(sil,
main = "Silhouette plot - k-means",
cex.names = 0.8,
col = 2:(k + 1))
plot(
silhouette(spe.ch.pam),
main = "Silhouette plot - PAM",
cex.names = 0.8,
col = 2:(k + 1)
)
On the basis on this plot, you should be able to tell which solution (PAMor kmeans) has a better silhouette profile.
You could also compare this result with the silhouette plot of the optimized Ward
clustering produced earlier. Except cluster numbering, is there any difference
between k-means and optimized Ward classifications? You can also address this
question by examining a contingency table:
# Compare classifications from k-means and optimized Ward
# clustering
table(spe.kmeans.g, spech.ward.gk)
Hint The PAM method is presented as “robust” because it minimizes a sum of
dissimilarities instead of a sum of squared Euclidean distances. It is also robust in
that it tends to converge to the same solution with a wide array of starting
medoids for a given k value; this does not guarantee, however, that the solution is
the most appropriate for a given research purpose.
This example shows that even two methods that are devoted to the same goal and
belong to the same general class (here non-hierarchical clustering) may provide
diverging results. It is up to the user to choose the one that yields classifications that
are bringing out more pertinent information or are more closely interpretable using
environmental variables (next section).
106
4 Cluster Analysis
