454
M. N. Bojnordi and P. Behnam
9.5.1 Data Clustering
Data clustering refers to partitioning a set of objects into meaningful groups
(a.k.a. clusters) without using predefined labels [103]. Data clustering tasks are
computationally difficult (NP-hard) problems. The entities of each cluster are more
similar to each other than to those in other clusters. The class of k-means algorithms
are the most prominent clustering techniques that have been successfully employed
in numerous fields of science and engineering [104]. Algorithm 9.1 shows the basic
steps of k-means clustering, where k centroids are defined to represent the clusters.
A centroid is either a representative member of the cluster, such as the median of
the cluster, or an additional data point computed based on the similarities among
all of the cluster members (e.g., the arithmetic mean). The former has been proven
to find better clusters than the latter due to its resistance against outlier members
[103, 104]. Prior to partitioning the data, the k centroids are randomly selected for
the clusters. The clustering task is carried out through two algorithmic steps (lines 3
and 4 in Algorithm 9.1) that are repeated after the initial step. Firstly, the clusters are
formed by assigning data points to their closest centroids. Secondly, new centroids
are computed for all of the clusters. These two steps are repeated for a constant
number of iterations defined by the application or until convergence is reached and
none of the members switches their clusters during the first step.
Algorithm 9.1 Basic k-Means Clustering
1: select k initial centroids randomly
2: repeat
3:
from k clusters: assign data points to their closest centroids
4:
recompute the centroid of each new cluster
5: until convergence is reached
9.5.2 Applications of Data Clustering
We can find numerous applications of k-means clustering in the literature. This
section reviews only two representative examples for gene expression analysis
(GEA) and text data mining.
9.5.2.1 Gene Expression Analysis
Recently, clustering has seen wide use in the field of medical research, such
as cancer diagnosis and drug discovery. An accurate clustering algorithm can
significantly improve the correctness of these applications. Lu and Han [105] have
shown that data clustering techniques may be employed to classify cancerous cells
Précédent

- 458/647

Suivant