# Heat map of the doubly ordered community table, with dendrogram
heatmap(
t(spe[rev(or$species)]),
Rowv = NA,
Colv = dend,
col = c("white", brewer.pal(5, "Greens")),
scale = "none",
margin = c(4, 4),
ylab = "Species (weighted averages of sites)",
xlab = "Sites"
)
Hint A similar result can be obtained using the tabasco() function of the vegan
package.
4.8 Non-hierarchical Clustering
Non-hierarchical partitioning consists in looking for a single partition of a set of
objects. The problem can be stated as follows: given n objects in a p-dimensional
space, determine a partition of the objects into k groups, or clusters, such that the
objects within each cluster are more similar to one another than to objects in the other
clusters. The user determines the number of groups, k. The partitioning algorithms
require an initial configuration, i.e. an initial attribution of the objects to the k groups,
which will be optimized in a recursive process. The initial configuration may be
provided by theory, but users often choose a random starting configuration. In that
case, the analysis is run a large number of times with different random initial
configurations in the hope to find the best solution.
Here we shall present two related methods, k-means partitioning and partitioning
around medoids (PAM). These algorithms operate in Euclidean space. An important
note is that if the variables in the data table are not dimensionally homogeneous, they
must be standardized prior to partitioning. Otherwise, the total variance of the data
has dimension equal to the sum of the squared dimensions of the individual variables, which is meaningless.
4.8.1 k-means Partitioning
The k-means method uses the local structure of the data to delineate clusters: groups
are formed by identifying high-density regions in the data. To achieve this, the
method iteratively minimizes an objective function called the total error sum of
squares (E
2
k or TESS or SSE), which is the sum of the within-group sums-ofsquares. This quantity is the sum, over the k groups, of the sums of the squared
(Euclidean) distances among the objects in the groups, each divided by the number
96
4 Cluster Analysis
heatmap(
t(spe[rev(or$species)]),
Rowv = NA,
Colv = dend,
col = c("white", brewer.pal(5, "Greens")),
scale = "none",
margin = c(4, 4),
ylab = "Species (weighted averages of sites)",
xlab = "Sites"
)
Hint A similar result can be obtained using the tabasco() function of the vegan
package.
4.8 Non-hierarchical Clustering
Non-hierarchical partitioning consists in looking for a single partition of a set of
objects. The problem can be stated as follows: given n objects in a p-dimensional
space, determine a partition of the objects into k groups, or clusters, such that the
objects within each cluster are more similar to one another than to objects in the other
clusters. The user determines the number of groups, k. The partitioning algorithms
require an initial configuration, i.e. an initial attribution of the objects to the k groups,
which will be optimized in a recursive process. The initial configuration may be
provided by theory, but users often choose a random starting configuration. In that
case, the analysis is run a large number of times with different random initial
configurations in the hope to find the best solution.
Here we shall present two related methods, k-means partitioning and partitioning
around medoids (PAM). These algorithms operate in Euclidean space. An important
note is that if the variables in the data table are not dimensionally homogeneous, they
must be standardized prior to partitioning. Otherwise, the total variance of the data
has dimension equal to the sum of the squared dimensions of the individual variables, which is meaningless.
4.8.1 k-means Partitioning
The k-means method uses the local structure of the data to delineate clusters: groups
are formed by identifying high-density regions in the data. To achieve this, the
method iteratively minimizes an objective function called the total error sum of
squares (E
2
k or TESS or SSE), which is the sum of the within-group sums-ofsquares. This quantity is the sum, over the k groups, of the sums of the squared
(Euclidean) distances among the objects in the groups, each divided by the number
96
4 Cluster Analysis
