simultaneously the member of two genera: membership is binary (0 or 1). Some
methods, less commonly used, consider fuzzy partitions, in which membership is
continuous (between 0 and 1). Depending on the clustering model, the result can be a
single partition or a series of hierarchically nested partitions. Except for its
constrained version, clustering is not a statistical method in the strict sense since it
does not test any a priori hypothesis. Post-hoc cluster robustness tests are available,
however (see Sect. 4.7.3.3). Clustering helps bring out some features hidden in the
data; it is the user who decides if these structures are interesting and worth
interpreting in ecological terms.
Note that most clustering methods are computed from association matrices, which
stresses the importance of the choice of an appropriate association coefficient.
Clustering methods differ by their algorithms, which are the effective methods for
solving a problem using a finite sequence of instructions. One can recognize the
following families of clustering methods:
1. Sequential or simultaneous algorithms. Most clustering algorithms are sequential
and consist in the repetition of a given procedure until all objects have found their
place. The less frequent simultaneous algorithms find the solution in a single step.
2. Agglomerative or divisive. Among the sequential algorithms, agglomerative procedures begin with the discontinuous collection of objects, which are successively grouped into larger and larger clusters until a single, all-encompassing
cluster is obtained. Divisive methods, on the contrary, start with the collection of
objects considered as one single group, and divide it into subgroups, and so on
until the objects are completely separated. In either case it is left to the user to
decide which of the intermediate partitions should be retained, given the problem
under study.
3. Monothetic versus polythetic. Divisive methods may be monothetic or polythetic.
Monothetic methods use a single descriptor (the one that is considered the best for
that level) at each step for partitioning, whereas polythetic methods use all
descriptors; in most polythetic cases, the descriptors are combined into an
association matrix.
4. Hierarchical versus non-hierarchical methods. In hierarchical methods, the
members of inferior-ranking clusters become members of larger, higher-ranking
clusters. Most of the time, hierarchical methods produce non-overlapping clusters. Non-hierarchical methods (e.g. k-means partitioning) produce a single partition, without any hierarchy among the groups.
5. Probabilistic versus non-probabilistic methods. Probabilistic methods define
groups in such a way that the within-group association matrices have a given
probability of being homogeneous. Probabilistic methods are sometimes used to
define species associations.
6. Unconstrained or constrained methods. Unconstrained clustering relies upon the
information of a single data set, whereas constrained clustering uses two matrices:
the one being clustered, and a second one containing a set of explanatory variables which provides a constraint (or guidance) as to where to group or divide the
data of the first matrix.
60
4 Cluster Analysis
methods, less commonly used, consider fuzzy partitions, in which membership is
continuous (between 0 and 1). Depending on the clustering model, the result can be a
single partition or a series of hierarchically nested partitions. Except for its
constrained version, clustering is not a statistical method in the strict sense since it
does not test any a priori hypothesis. Post-hoc cluster robustness tests are available,
however (see Sect. 4.7.3.3). Clustering helps bring out some features hidden in the
data; it is the user who decides if these structures are interesting and worth
interpreting in ecological terms.
Note that most clustering methods are computed from association matrices, which
stresses the importance of the choice of an appropriate association coefficient.
Clustering methods differ by their algorithms, which are the effective methods for
solving a problem using a finite sequence of instructions. One can recognize the
following families of clustering methods:
1. Sequential or simultaneous algorithms. Most clustering algorithms are sequential
and consist in the repetition of a given procedure until all objects have found their
place. The less frequent simultaneous algorithms find the solution in a single step.
2. Agglomerative or divisive. Among the sequential algorithms, agglomerative procedures begin with the discontinuous collection of objects, which are successively grouped into larger and larger clusters until a single, all-encompassing
cluster is obtained. Divisive methods, on the contrary, start with the collection of
objects considered as one single group, and divide it into subgroups, and so on
until the objects are completely separated. In either case it is left to the user to
decide which of the intermediate partitions should be retained, given the problem
under study.
3. Monothetic versus polythetic. Divisive methods may be monothetic or polythetic.
Monothetic methods use a single descriptor (the one that is considered the best for
that level) at each step for partitioning, whereas polythetic methods use all
descriptors; in most polythetic cases, the descriptors are combined into an
association matrix.
4. Hierarchical versus non-hierarchical methods. In hierarchical methods, the
members of inferior-ranking clusters become members of larger, higher-ranking
clusters. Most of the time, hierarchical methods produce non-overlapping clusters. Non-hierarchical methods (e.g. k-means partitioning) produce a single partition, without any hierarchy among the groups.
5. Probabilistic versus non-probabilistic methods. Probabilistic methods define
groups in such a way that the within-group association matrices have a given
probability of being homogeneous. Probabilistic methods are sometimes used to
define species associations.
6. Unconstrained or constrained methods. Unconstrained clustering relies upon the
information of a single data set, whereas constrained clustering uses two matrices:
the one being clustered, and a second one containing a set of explanatory variables which provides a constraint (or guidance) as to where to group or divide the
data of the first matrix.
60
4 Cluster Analysis
