4.8.2 Partitioning Around Medoids (PAM)
Partitioning around medoids (Chap. 2 in Kaufman and Rousseeuw 2005) “searches
for k representative objects or medoids among the observations of the dataset. These
observations should represent the structure of the data. After finding a set of k
medoids, k clusters are constructed by assigning each observation to the nearest
medoid. The goal is to find k representative objects which minimize the sum of the
dissimilarities of the observations to their closest representative object” (excerpt
from the pam() documentation file). By comparison, k-means minimizes the sum of
the squared Euclidean distances within the groups. k-means is thus a traditional
least-squares method, while PAM is not
2 . As implemented in R, pam()(package
cluster) accepts raw data or dissimilarity matrices (an advantage over
0
50
100
150
20
40
60
80
100
Four optimized Ward clusters along the Doubs River
x coordinate (km)
y coordinate (km)
Upstream
Downstream
1
2
3
4
5
6
7
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
Cluster 1
Cluster 2
Cluster 3
Cluster 4
Fig. 4.22 The four optimized Ward clusters on a map of the Doubs River
2 Many dissimilarity functions used in community ecology are non-Euclidean (for example the
Jaccard, Sørensen and % difference indices). However, the square-root of these D functions is
Euclidean. PAM minimizes the sum of the dissimilarities of the observations to their closest
medoid. Since D is the square of the square-rooted (Euclidean) dissimilarities, PAM is a leastsquares method when applied to these non-Euclidean D functions. See Legendre and De Cáceres
(2013), Appendix S2.
4.8 Non-hierarchical Clustering
103
Précédent

- 116/444

Suivant