188
Ö. Akay et al.
E =
k
l=1
n
i=1
y il d(X i , Q l )
(18.1)
Here, Q l = [q l1 , q l2 , …, q lm ] is the representative vector or prototype for cluster l,
and y il is an element of a partition matrix Y nxl . d(X i , Q l ) is the dissimilarity measure
defined as follows:
d(X i , Q l ) =
p
j=1
T
t=1
X
r
i j (t) − q
r
l j
2 + μ l
m
j= p+1
δ
X
c
i j (t) − q
c
l j
(18.2)
where δ(p, q) = 0 for p = q, and δ(p, q) = 1 for p = q; X
r
i j (t)
X
c
i j (t)
is the value of
the jth numeric (categorical) attribute at the time t for the data object i; q
r
l j
q
c
l j
is the
prototype of the jth numeric (categorical) attribute in the cluster l;μ l is a weight for
categorical attributes in the cluster l (Ji et al. 2012). The process of the k-prototype
algorithm is described as follows:
The process of the K-prototype algorithm is defined as follows:
Step 1. Randomly select k data objects from the dataset X as the initial prototype
of the sets.
Step 2. For each data object in X, assign it to the cluster whose prototype is closest
to that data object in terms of Eq. (18.2). After each assignment, update the prototype
of the cluster.
Step 3. After all data objects have been assigned to a cluster, recalculate the
similarity of the data objects with the existing prototypes. If a data object is found to
belong to another cluster rather than the closest prototype, reassign that data object
to that cluster and update the prototypes of both clusters.
Step 4. After the full circle test of X, terminate the algorithm if no data object has
changed the sets, or else repeat step 3 (Ji et al. 2013).
Different clustering algorithms often lead to different clusters of data, even for
the same algorithm, the choice of different parameters or the order of presentation
of data objects can greatly affect the final clusters. Therefore, effective assessment
standards and criteria are critical to reassuring users of cluster results. For all that,
these evaluations provide meaningful information on how many clusters are hidden
in the data. Actually, the user is faced with the dilemma of selecting the number of
clusters or partitions in the underlying data. Therefore, numerous indices have been
proposed to determine the number of clusters in a data set (Charrad et al. 2012).
Some clustering validity indices are used to select the optimal number of clusters.
These indices are The C-Index, Dunn index, Gamma index, Gplus index, McClain
index, Ptbiserial index, Silhouette index and Tau index. The minimum values of
the C-Index, Gplus and McClain index are used to indicate the optimal number of
clusters. The maximum values of the Dunn, Gamma, Ptbiserial, Silhouette and Tau
index are used to indicate the optimal number of clusters.
Précédent

- 186/206

Suivant