7 NIR Data Exploration and Regression by Chemometrics—A Primer
155
the inspection of the obtained data alone, and that in a very short time by using
interactive graphical displays with easily interpretable symbols and colors according
to the metadata of the dataset.
The term metadata is used for any information associated with the data that is
not a part of the PCA itself. Typically, such data are categorical and do only exist as
discrete levels, like the name of the person who measured a specific sample, or if the
measurement is first, second or third of a set of replicate measurements. Metadata
can also be numerical, but not relevant to include in the PCA itself—e.g., “seconds
elapsed since a reference measurement was made on the equipment,” “the content
of an analyte in the sample” or “the relative humidity in the laboratory on that day”.
The most important source of metadata is the experimental design parameters such
as harvest year, variety and field. Metadata are very important to gain insight into
a dataset and can be exploited in the PCA, in the validation of multivariate models,
in PLS-DA models and in ASCA models. Some chemometric programs allow for
coloring the samples according to a quantitative response parameter, which is going
to be target for regression analysis. As an example, see Fig. 7.17 where the score
plot is colored according to the sugar contents. This allows for a quick, graphical
investigation on how much of the total variance is related to the quantitative response
parameter, and how systematic it is distributed over the sample set.
Using PCA makes it easy to evaluate the validity of a dataset simply by observing
the location of the replicates in a score scatter plot. An experiment can contain two
types of sample replicates: experimental replicates (i.e., “mixing” or “chemical”)
and measurement or analytical replicates. Concerning experimental replicates, the
samples are to be considered experimentally alike, but have different origins (for
instance, the same type of beer, but brewed on three different days). When each of
the experimental replicates is measured several times, they become measurement
or analytical replicates. This is illustrated in Fig. 7.18, where a score plot for two
components resulting from a PCA displays three experimental replicates which each
has been measured three times (in a random order).
Based on the location of the colored groups, an inspection reveals that the samples
originating from the red group are significantly different (distant), than the green
and blue groups, which are very similar (close). In addition, the green and blue
measurement collections appear more similar (closer), than the red group, which
spreads out more indicating a higher inter-group variance. The next step would be to
inspect the loadings of the two components to investigate why the difference in red
and green/blue is so significant, or to look back into the experimental logbook to see
if there is anything known about red group that can hint at this separation.
It should be noted here that PCA on real spectral data always is able to find
and illuminate such replicate variances and groupings, no matter how small. It is
thus important to compare the replicate (intra-group) variance to the sample (intergroup) variance. When conducting this exercise, it is important to have always the
explained variance of the investigated PCs in mind as they carry the information of
the magnitude of the explained variance (i.e., importance) in the two directions.
PCA can be advantageous in analyzing performance of an analytical technique
or sample preparation over time. One such diagnostic feature is the pool sample,
155
the inspection of the obtained data alone, and that in a very short time by using
interactive graphical displays with easily interpretable symbols and colors according
to the metadata of the dataset.
The term metadata is used for any information associated with the data that is
not a part of the PCA itself. Typically, such data are categorical and do only exist as
discrete levels, like the name of the person who measured a specific sample, or if the
measurement is first, second or third of a set of replicate measurements. Metadata
can also be numerical, but not relevant to include in the PCA itself—e.g., “seconds
elapsed since a reference measurement was made on the equipment,” “the content
of an analyte in the sample” or “the relative humidity in the laboratory on that day”.
The most important source of metadata is the experimental design parameters such
as harvest year, variety and field. Metadata are very important to gain insight into
a dataset and can be exploited in the PCA, in the validation of multivariate models,
in PLS-DA models and in ASCA models. Some chemometric programs allow for
coloring the samples according to a quantitative response parameter, which is going
to be target for regression analysis. As an example, see Fig. 7.17 where the score
plot is colored according to the sugar contents. This allows for a quick, graphical
investigation on how much of the total variance is related to the quantitative response
parameter, and how systematic it is distributed over the sample set.
Using PCA makes it easy to evaluate the validity of a dataset simply by observing
the location of the replicates in a score scatter plot. An experiment can contain two
types of sample replicates: experimental replicates (i.e., “mixing” or “chemical”)
and measurement or analytical replicates. Concerning experimental replicates, the
samples are to be considered experimentally alike, but have different origins (for
instance, the same type of beer, but brewed on three different days). When each of
the experimental replicates is measured several times, they become measurement
or analytical replicates. This is illustrated in Fig. 7.18, where a score plot for two
components resulting from a PCA displays three experimental replicates which each
has been measured three times (in a random order).
Based on the location of the colored groups, an inspection reveals that the samples
originating from the red group are significantly different (distant), than the green
and blue groups, which are very similar (close). In addition, the green and blue
measurement collections appear more similar (closer), than the red group, which
spreads out more indicating a higher inter-group variance. The next step would be to
inspect the loadings of the two components to investigate why the difference in red
and green/blue is so significant, or to look back into the experimental logbook to see
if there is anything known about red group that can hint at this separation.
It should be noted here that PCA on real spectral data always is able to find
and illuminate such replicate variances and groupings, no matter how small. It is
thus important to compare the replicate (intra-group) variance to the sample (intergroup) variance. When conducting this exercise, it is important to have always the
explained variance of the investigated PCs in mind as they carry the information of
the magnitude of the explained variance (i.e., importance) in the two directions.
PCA can be advantageous in analyzing performance of an analytical technique
or sample preparation over time. One such diagnostic feature is the pool sample,
