Assessing Inhomogeneous Indicator-Related Typologies Through the Reverse. . .
33
similar typology to the one we are given at the outset, in this case – provided by the
geographers.
The study had, therefore, three essential motivations:
1. Yet another check on the capacities and effectiveness of the reverse clustering
approach (Owsi´ nski et al. 2017a, b; 2021), for yet another set of data and for a
different kind of substantive prerequisites;
2. Analysis of a typology, given by a complex and “branching” process, so as to
derive conclusions on the “deviations” of this typology from the one obtained on
the basis of a coherent, unified data set on the same subject, this analysis leading,
hopefully, to some broader conclusions; and
3. The substantive analysis of the given data set, for comparison with the original
typology, and, perhaps, some tangible substantive conclusions.
The broader meaning of this exercise (implied in the motivation 2 above) results
from the following image of the situation:
There is a set of data, concerning a collection of entities, describing certain
features of these entities. We wish to categorise these entities into a relatively small
number of categories (much smaller than the number of entities). (We abstract here
from the question whether we deal with an entire “population” or a “sample”. In the
latter case we assume the “sample” is “representative”.) In general, it would often be
convenient, if the categories formed a linear order (for the reason of categorization
in many cases refers somehow to the “composite indicator” context), although it
may happen that they do not. Namely, in many contexts, even if we are aware of
the essential multidimensionality of the subject matter, we deal with some sort of
“general axis”, corresponding to the potential or hypothetical “composite indicator”,
this axis representing the “magnitude / intensity / graveness of the phenomenon”,
to which the indicator is supposed to refer. In this particular case we deal with
the “urban-rural-peripheral” axis, and the potential divergences are associated with
some special phenomena and corresponding groups of units, e.g. urban areas
featuring different patterns of development (or, indeed, decay), or peripheral rural
areas, where the share of settling urbanites plays an important role (see also Fig. 3
further on).
We assume that we do dispose of a certain categorization of the entities in
question, this categorization coming out of a special procedure, which involves, say,
categorical variables, branchings, various data sets in various branches of the procedure, etc., like in the examples of Fig. 2, showing two typical procedures, related to
social care / unemployment benefit registration and relevant data production, which
can hardly be translated into a unified data set and easily processed as such.
Having the data set and its partition, we wish to recreate the partition for this data
set as faithfully as possible, using clustering. We shall not be using the “decision
variables” of the procedure (or, if used, they will be treated like other variables), and
the data will be the same for all objects. In the here analysed case we used a different
set of variables, as our assumption was to use only the publicly available data.
The substantive sense of the data remains, though, except for one or two original
variables, not accessible to us, very much the same.
Précédent

- 52/324

Suivant