36
J. W. Owsi´ nski et al.
and therefore also the statuses of the obtained solutions. A broader account on
this is provided in Owsi´ nski et al. (2018, 2021), and we shall only sketch here the
respective image. We can speak, namely, of two “axes” in this image:
(I) the degree of “certainty” of P A , meaning that this partition may be either a
“solid” one, e.g. in our case, the classification of the municipalities according
to provinces (a province constituting a cluster), or just a hypothesis (an expert
opinion);
(II) the degree of association of the partition P A with the data set X (the known
or assumed foundation of P A on the actual data set X) – in the above case
of provinces and municipalities there is no such association, while an expert
would most probably base her/his opinion on the data from X, when, for
instance, designing a functional typology, like in the present case, although
not necessarily in an exact manner.
In this perspective, Table 1 presents the relevant examples of situations potentially encountered.
As already indicated, the vector Z of the clustering procedure parameters sought,
is composed of the choice of the algorithm itself, the crucial parameter(s) of
the algorithm, the weights of variables, and the scaling of the distance measure
(Minkowski exponent). The algorithms accounted for are k-means and similar
(Steinhaus 1956; Lloyd 1957), general hierarchical aggregation (parameterised with
the Lance-Williams formula, see Lance and Williams 1966, 1967), and DBSCAN
(Ester et al. 1996), as a representative of the local density-based algorithms. Thus,
quite a broad range of algorithms is covered, with, indeed, a very significant scope
of search, regarding the variables composing Z.
Given the composition of the optimised Z one could include in it, for instance,
an explicit feature selection procedure or another operation, oriented at the shaping
of the description space, as, say a preprocessing stage. We preferred, though, to
encapsulate the entire procedure in one optimisation task, aiming integrally at
getting possibly close to P A , including all the parameters we thought would be
important.
4 The Case Studied
The particular case here studied (described also preliminarily in Owsi´ nski et al.
2018) concerns the typology of close to 2500 Polish municipalities, elaborated for
definite planning purposes by a team from the Institute of Geography and Spatial
Organization of the Polish Academy of Sciences (see ´
Sleszy´ nski and Komornicki
2016). As already indicated, the typological procedure was quite complex, with the
use of a variety of variables and criteria, and including branching decisions. For our
purposes here suffice to quote the “headings” of the typology elaborated, as given
in Table 2.
J. W. Owsi´ nski et al.
and therefore also the statuses of the obtained solutions. A broader account on
this is provided in Owsi´ nski et al. (2018, 2021), and we shall only sketch here the
respective image. We can speak, namely, of two “axes” in this image:
(I) the degree of “certainty” of P A , meaning that this partition may be either a
“solid” one, e.g. in our case, the classification of the municipalities according
to provinces (a province constituting a cluster), or just a hypothesis (an expert
opinion);
(II) the degree of association of the partition P A with the data set X (the known
or assumed foundation of P A on the actual data set X) – in the above case
of provinces and municipalities there is no such association, while an expert
would most probably base her/his opinion on the data from X, when, for
instance, designing a functional typology, like in the present case, although
not necessarily in an exact manner.
In this perspective, Table 1 presents the relevant examples of situations potentially encountered.
As already indicated, the vector Z of the clustering procedure parameters sought,
is composed of the choice of the algorithm itself, the crucial parameter(s) of
the algorithm, the weights of variables, and the scaling of the distance measure
(Minkowski exponent). The algorithms accounted for are k-means and similar
(Steinhaus 1956; Lloyd 1957), general hierarchical aggregation (parameterised with
the Lance-Williams formula, see Lance and Williams 1966, 1967), and DBSCAN
(Ester et al. 1996), as a representative of the local density-based algorithms. Thus,
quite a broad range of algorithms is covered, with, indeed, a very significant scope
of search, regarding the variables composing Z.
Given the composition of the optimised Z one could include in it, for instance,
an explicit feature selection procedure or another operation, oriented at the shaping
of the description space, as, say a preprocessing stage. We preferred, though, to
encapsulate the entire procedure in one optimisation task, aiming integrally at
getting possibly close to P A , including all the parameters we thought would be
important.
4 The Case Studied
The particular case here studied (described also preliminarily in Owsi´ nski et al.
2018) concerns the typology of close to 2500 Polish municipalities, elaborated for
definite planning purposes by a team from the Institute of Geography and Spatial
Organization of the Polish Academy of Sciences (see ´
Sleszy´ nski and Komornicki
2016). As already indicated, the typological procedure was quite complex, with the
use of a variety of variables and criteria, and including branching decisions. For our
purposes here suffice to quote the “headings” of the typology elaborated, as given
in Table 2.
