CHAPTER 21 • (hemometries for Sampling and Analysis: Theory and Environmental Applications
391
highly correlated. Because of this correlation, in the V-dimensional information space
the objects are not uniformly distributed, but they are grouped into a restricted area
of the space; they form a structure.
Let us consider a very simple example, where the objects are solutions of the same
compounds at different concentrations, described by the absorbance at two wavelengths:
tl;l
kl Ci + eil
tl;2
k2 Ci + ei2
The two absorbances are proportional to the concentration (with a noise component for each object i for the two absorbances). The space of information is two dimensional, but the two absorbances are closely correlated, because
k2
a·2 = - c· + e·
I
kl I
I
so that (Fig. 21.3) the objects are on a straight line. The useful information has an unidimensional structure; the distance of an object from the structure is a consequence
of the noise, i.e. the space outside the structure is that of useless information.
There are a lot of possible structures (linear, non-linear, clusters), but in real problems the V-dimensional space of the information can be always divided into an inner
subspace of structured, useful information, and in an outer space of noise. Very frequently the dimension of the inner space is small, also when the dimension of the original information is very large.
peA aims to detect the structure and to evaluate the number of dimensions of the
inner space.
Principal components are the eigenvectors of centred data, often of auto scaled data.
In a bidimensional example (Fig. 21.4) the first eigenvector E1 is the direction of maximum variance, with origin in the centroid. Its loadings (the direction cosines of the
two axes) can be obtained as in the usual univariate regression, as the direction for
Fig. 21.3. Linear structure
02
caused by correlated variables
°1
391
highly correlated. Because of this correlation, in the V-dimensional information space
the objects are not uniformly distributed, but they are grouped into a restricted area
of the space; they form a structure.
Let us consider a very simple example, where the objects are solutions of the same
compounds at different concentrations, described by the absorbance at two wavelengths:
tl;l
kl Ci + eil
tl;2
k2 Ci + ei2
The two absorbances are proportional to the concentration (with a noise component for each object i for the two absorbances). The space of information is two dimensional, but the two absorbances are closely correlated, because
k2
a·2 = - c· + e·
I
kl I
I
so that (Fig. 21.3) the objects are on a straight line. The useful information has an unidimensional structure; the distance of an object from the structure is a consequence
of the noise, i.e. the space outside the structure is that of useless information.
There are a lot of possible structures (linear, non-linear, clusters), but in real problems the V-dimensional space of the information can be always divided into an inner
subspace of structured, useful information, and in an outer space of noise. Very frequently the dimension of the inner space is small, also when the dimension of the original information is very large.
peA aims to detect the structure and to evaluate the number of dimensions of the
inner space.
Principal components are the eigenvectors of centred data, often of auto scaled data.
In a bidimensional example (Fig. 21.4) the first eigenvector E1 is the direction of maximum variance, with origin in the centroid. Its loadings (the direction cosines of the
two axes) can be obtained as in the usual univariate regression, as the direction for
Fig. 21.3. Linear structure
02
caused by correlated variables
°1
