133
Clustering and Classification
In the next step, for each example in the feature space, we need to determine
the cluster to which the example belongs to. This choice is simply made based
on the proximity of the examples to the cluster centers. The results of this cluster
assignment step are shown in Figure 7.2c.
Next, we need to average the members in each of the resulting cluster to find the
new cluster centers and then reassign the samples to the clusters, this time using
the new cluster centers. The new cluster centers are shown in Figure 7.2d. As can
be seen, the centers of the clusters are now closer to the expected centers of the
actual clusters. At the next iteration, we use the new cluster centers to assign every
example to one of the three clusters. The results of this assignment are shown in
Figure 7.2e. Finally, we calculate the new centers based on the clusters obtained
in Figure 7.2e, which gives the cluster centers depicted in Figure 7.2f. At t
  his point,
the cluster centers are indeed in the center of actual clusters and repeating the
process for extra iteration will not change the clusters. At this point, the algorithm
has converged to the actual clusters, and, therefore, the iterations are terminated.
This example shows how K-means can effectively cluster unlabeled data.
In MATLAB ® , the command “k-means” is used to perform K-means clustering.
We explore using MATLAB for K-means clustering in the following example.
Example 7.3
In this example, we first generate 40 random samples. Twenty samples in this pool
are generated using a 2-D normal distribution centered at (1, 1), and the remaining 20 samples are from another normal distribution centered at (−1, −1). Then,
a K-means algorithm is applied to cluster these samples into two clusters. The
MATLAB codes for this example are as follows:
X=[randn(20,2) + 2.8 * ones(20,2);

randn(20,2)−2.8*ones(20,2)]

[cidx, ctrs] = kmeans(X, 2, ‘dist’,‘city’, ‘rep’,5, ‘disp’,
‘final’)
plot(X(cidx==1,1),X(cidx==1,2),‘r.’, …
X(cidx==2,1),X(cidx==2,2), ‘b.’, ctrs(:,1),ctrs(:,2), ‘kx’);
In the preceding code, first, we use the command “randn” to generate 40 normally
distributed random samples out of which 20 samples are around (1, 1) and the rest
are centered at (−1, −1). Then, we use K-means algorithm to perform clustering of
these samples and divide them into two clusters. Number “2” in k-means command reflects our desire to form two clusters for the data. The options “dist” and
“sqEuclidean” specify the Euclidean distance as the distance measure for clustering. The rest of the code deals with the labeling of the samples in each cluster.
Figure 7.3 shows the result of clustering. Samples of one cluster are shown with
red, and the samples belonging to the other class are graphed in blue color. We
have also marked the center of each cluster by X’s. In this example, many points
generated by the first random generator, centered at (1, 1), are correctly assigned
to the “blue” cluster. However, some points closer to the origin, while generated
by the random generator at (1, 1), may be assigned to the red group, which is
mainly formed by the sample of the other normal distribution. This shows that
even though K-means is a simple and rather fast algorithm, it has some limitations.
Précédent

- 160/412

Suivant