CHAPTER 21 . (hemometries for Sampling and Analysis: Theory and Environmental Applications
variables have the same importance in regard to information, because they have the
same dispersion. Then, the autoscaled variables can be divided by the standard deviation of the relative noise, SN / Sx (weighting). So, the importance of a variable becomes
proportional to the signal/noise ratio. Simple autoscaling must not be used with homogeneous data with almost constant noise. For example, consider absorbances at two
wavelengths, with the same noise: the first has high absorption, the second has no
absorption. Simple auto scaling gives the same importance to informative and the noisy
absorbance. Weighting restores the original data. In this case an often-used transform
is the simple subtraction of the column mean (centring). Therefore, it is evident that
only the knowledge of data can suggest which pretreatment we have to use. A very
simple mathematical operation as autoscaling can have pernicious effects when improperly used.
Frequently, in the case of homogeneous data, row profiles (percentages) are computed:
Xoriginal
Xprofile =
L Xoriginal
Variables
This other simple mathematical operation cancels a part of the original information, so proportional objects become equal. Only the knowledge of the problem allows
the decision about the adequacy of this pretreatment, since the knowledge of the problem is fundamental to select the suitable treatment in the case of homogeneous ordered variables, as those of spectra, chromatography, hyphenated techniques. Smoothing, first and second derivative, Fourier transforms, wavelet transforms, base line subtraction and row autoscaling (also known as Standard Normal Variate) are examples
of important pretreatment tools.
21.2.2
Similarity and Clustering
Also similarity is a problem-dependent concept. Usually, two samples are considered
similar when the difference in composition is small if compared to the variability.
Fig. 21.1 shows nine samples described by two variables.
A current definition of the similarity between two objects i and j is:
d·
= 1- _'J_
sij
d max
where the distance between the two objects (dependent on the pretreatment, obviously)
is divided by the maximum inter-object distance of the data set, so that the farthest
objects have similarity o.
However, there is a second concept of similarity, based on the membership to a structure. Fig. 21.2 shows that object 2 is more similar to object 1 than to object 3, because
both 1 and 2 belong to the same structure.
As for objects, the similarity between variables can be defined. A usual definition
is that similarity is measured by the (1 + r) / 2, where r is the linear correlation coefficient. So, the similarity is zero when the two variables are opposite, linked by a linear
relationship with negative slope. A second definition is that the similarity is measured
Précédent

- 395/447

Suivant