34
J. W. Owsi´ nski et al.
Fig. 2 Two examples of the procedures, leading to the kind of prior categorization of interest here
From the analysis we wish to get an image more “naturally” related to the
data and to compare it with the original categorization, in terms of: (a) existence
of potential “twists” or “artifacts” in the original partition, resulting from the
application of “decision variables”, “thresholds”, etc.; (b) the number and the nature
of categories (clusters) obtained in a more flexible environment; and (c) detecting
the potential outliers, which might, again, be overlooked in the original partition.
In the case there is a definite need of the categories to form an ordering (along
the hypothetical axis of a “composite indicator”), the exercise might also serve to
(d) confirm or put to doubt the possibility of actual formation of such an order in a
more “natural” manner. This kind of situation is schematically depicted in Fig. 3. We
would like to indicate, in the context of this illustration, that two quite typical issues
arise, in connection with a potential partition of a data set that in its general shape is
distributed along the already mentioned “main indicator axis”, namely: 1. Frequent
arbitrary manner of cutting into pieces the “cloud” of data points, stretching along
this main axis; 2. In the cases of “branchings” out of this main axis the question of
their relation to the main axis (which ones lie along the main axis, and which one
diverge from it?) and the existence of an actual separation from the main axis.
An example of such a situation may be also provided by the case of poverty
measurement, and the categories, established in this context. In the measurement
the leading variable seems to be income per capita in the household, social and
other benefits included, while other variables, such as number of persons in the
household, ages of household members, their education, health conditions, housing
situation, etc., being usually either highly or at least significantly correlated. There
may, however, be yet another variable, or group of variables, that are less correlated,
but for some definite reason (e.g. crosschecking) included in the measurement. Say:
driving license? obesity? political attitude? arms possession?
Précédent

- 53/324

Suivant