32
J. W. Owsi´ nski et al.
The data set
analysed X
The prior parƟƟon
of X, i.e. P A
The clustering
algorithms and the
data processing
parameters: Z
The search (optimisation) procedure:
maximising the Q(P A ,P B )
The criterion
Q(P A ,P B ):
similarity of
the two
partitions
The obtained
partition of X:
P B
Fig. 1 Schematic view of the reverse clustering procedure
The optimisation procedure is applied to the vector Z, describing the selected
clustering algorithm that yields the partition P B . This vector is composed of: (i) the
choice of the clustering algorithm, (ii) the choice of its essential parameter(s) – e.g.
the number of clusters, or some threshold distance etc., (iii) the weights, or choice,
of variables, (iv) the distance definitions used (e.g. as expressed through Minkowski
exponent).
The working of the entire procedure of reverse clustering is schematically shown
in Fig. 1.
The preliminary results from the concrete study, considered here, were reported
in Owsi´ nski et al. (2018). Now, besides presenting an ampler view of the results,
we focus on the broader implications for the use of a similar approach in other
settings, where the “composite indicator” context applies, while either continuous
variables are used along with the “strongly” discrete ones to categorise objects, or
the procedure applied involves branchings, so that, in effect, a single-axis-indicator
might not render appropriately the resulting categories.
2 The Study with Its Narrow and Broader Motivations
We present here a study, in which the data on all of Polish municipalities are
analysed in the presence of a definite typology of these municipalities, elaborated
for a concrete (spatial planning) purpose by the specialists from the Institute of
Geography and Spatial Organization of the Polish Academy of Sciences ( ´
Sleszy´ nski
and Komornicki 2016). We apply in this context the reverse clustering approach,
meaning that we try to find the parameters of the broadly conceived clustering
procedure, which, when applied to the data on the municipalities, yield a possibly
J. W. Owsi´ nski et al.
The data set
analysed X
The prior parƟƟon
of X, i.e. P A
The clustering
algorithms and the
data processing
parameters: Z
The search (optimisation) procedure:
maximising the Q(P A ,P B )
The criterion
Q(P A ,P B ):
similarity of
the two
partitions
The obtained
partition of X:
P B
Fig. 1 Schematic view of the reverse clustering procedure
The optimisation procedure is applied to the vector Z, describing the selected
clustering algorithm that yields the partition P B . This vector is composed of: (i) the
choice of the clustering algorithm, (ii) the choice of its essential parameter(s) – e.g.
the number of clusters, or some threshold distance etc., (iii) the weights, or choice,
of variables, (iv) the distance definitions used (e.g. as expressed through Minkowski
exponent).
The working of the entire procedure of reverse clustering is schematically shown
in Fig. 1.
The preliminary results from the concrete study, considered here, were reported
in Owsi´ nski et al. (2018). Now, besides presenting an ampler view of the results,
we focus on the broader implications for the use of a similar approach in other
settings, where the “composite indicator” context applies, while either continuous
variables are used along with the “strongly” discrete ones to categorise objects, or
the procedure applied involves branchings, so that, in effect, a single-axis-indicator
might not render appropriately the resulting categories.
2 The Study with Its Narrow and Broader Motivations
We present here a study, in which the data on all of Polish municipalities are
analysed in the presence of a definite typology of these municipalities, elaborated
for a concrete (spatial planning) purpose by the specialists from the Institute of
Geography and Spatial Organization of the Polish Academy of Sciences ( ´
Sleszy´ nski
and Komornicki 2016). We apply in this context the reverse clustering approach,
meaning that we try to find the parameters of the broadly conceived clustering
procedure, which, when applied to the data on the municipalities, yield a possibly
