Assessing Inhomogeneous Indicator-Related Typologies Through the Reverse. . .
35
Fig. 3 An image of categorization, for which the reverse clustering would reveal divergences
from a “natural” partition. Rounded shapes indicate objects, belonging to (five) clusters “along
the primary indicator axis”, while the rectangular ones – belonging to (two) “divergent clusters”
3 On the Reverse Clustering
As outlined in the Introduction, the reverse clustering technique aims at identifying
the partition P B of a certain set of n entities into clusters (A q , q = 1, . . . ,p), which, for
the given set X of (descriptions of) these entities, X = {x i = {x i1 , . . . ,x im }} i = 1, . . . ,n ,
and a given a priori partition P A of these entities, is the closest to P A . In the search
for the best P B we apply a natural measure of similarity, or distance, between
the partitions, namely the Rand index (Rand 1971) or some of its variations (see,
e.g., Hubert and Arabie 1985). The Rand index in its original form counts, for two
partitions, the following four numbers: of pairs of objects which fall into the same
cluster in both partitions (a), those that are in different clusters in both partitions (b),
and that are in the same cluster in one partition and in different clusters in the other
(c and d). The original Rand index is simply the quotient: (a + b)/(a + b + c + d),
with a + b + c + d = ½n(n − 1), of course. The search is performed with
the evolutionary algorithm of own design (Sta´ nczak 2003), featuring two-level
selection: of individuals and of the operations.
Although it could be argued that the very problem of the reverse clustering is
equivalent to some kind of “supervised classifier choice”, it is, in general, not.
Namely, (1) the aim is not to provide the basis for classifying individual incoming
new entities, one after another, but rather to map entire new samples (of cardinality
perhaps even much bigger than n) into the obtained partition P B ; (2) the new
partition P B needs not be composed of the same number of clusters as P A ; (3) in
particular, the approach may lead to the identification of outliers, not entering any
of the essential clusters, forming the partition.
In addition, let us emphasise that the generic problem statement encompasses, in
fact, a whole variety of the potential situations, having quite different interpretations,
35
Fig. 3 An image of categorization, for which the reverse clustering would reveal divergences
from a “natural” partition. Rounded shapes indicate objects, belonging to (five) clusters “along
the primary indicator axis”, while the rectangular ones – belonging to (two) “divergent clusters”
3 On the Reverse Clustering
As outlined in the Introduction, the reverse clustering technique aims at identifying
the partition P B of a certain set of n entities into clusters (A q , q = 1, . . . ,p), which, for
the given set X of (descriptions of) these entities, X = {x i = {x i1 , . . . ,x im }} i = 1, . . . ,n ,
and a given a priori partition P A of these entities, is the closest to P A . In the search
for the best P B we apply a natural measure of similarity, or distance, between
the partitions, namely the Rand index (Rand 1971) or some of its variations (see,
e.g., Hubert and Arabie 1985). The Rand index in its original form counts, for two
partitions, the following four numbers: of pairs of objects which fall into the same
cluster in both partitions (a), those that are in different clusters in both partitions (b),
and that are in the same cluster in one partition and in different clusters in the other
(c and d). The original Rand index is simply the quotient: (a + b)/(a + b + c + d),
with a + b + c + d = ½n(n − 1), of course. The search is performed with
the evolutionary algorithm of own design (Sta´ nczak 2003), featuring two-level
selection: of individuals and of the operations.
Although it could be argued that the very problem of the reverse clustering is
equivalent to some kind of “supervised classifier choice”, it is, in general, not.
Namely, (1) the aim is not to provide the basis for classifying individual incoming
new entities, one after another, but rather to map entire new samples (of cardinality
perhaps even much bigger than n) into the obtained partition P B ; (2) the new
partition P B needs not be composed of the same number of clusters as P A ; (3) in
particular, the approach may lead to the identification of outliers, not entering any
of the essential clusters, forming the partition.
In addition, let us emphasise that the generic problem statement encompasses, in
fact, a whole variety of the potential situations, having quite different interpretations,
