182
K. M. Sørensen et al.
the random variation. If formalized as a PLS-DA problem, this is known as multi-level
PLS-DA [55].
A limitation of the application of ASCA is that it is most optimally applied to
balanced design structures. On the other hand, the statistical significance of the
strength of each design factor and their interactions can be evaluated using a permutation test [56], thus providing a link to classical statistics. Needless to say that
appropriate pre-processing and centering of the X matrix will be essential to the
subsequent ASCA.
7.8.1 Application of ASCA to NIR Spectra
Since the ASCA description above can seem a little theoretical, it will be illustrative
to demonstrate the potential of ASCA from an example. Here, it is used to partition
(split) the variance due to different sampling geometries in Dataset 4. The ASCA
model thus reflects the experimental factors shown in Eq. 7.32 and Fig. 7.33. Only
three factors are included in the model, and the “individual” factor is left out.
It will make no sense to include “individual” as a design factor in the ASCA
model, as the different individuals are not the same for the different varieties. They
are physically different kernels and can thus not be compared across varieties. This
leaves us to focus on the partitioned effects of the three design factors:
X = X variety + X position + X orientation + E
(7.32)
The ASCA model will be computed on Dataset 4 with the spectra X, pre-processed
by a 2nd order, derivative Savitzky-Golay filter of window width 7, followed by a
MSC, and finally a mean centering. The effect is evaluated with 1000 permutations.
For simplicity, interactions between the three factors are not included. The full rank
of each factor is also calculated.
The primary output of the ASCA model is the effect table (Table 7.2). From the
table, it is clear that largest effect is found in the residual (81.1%). In the current case,
this is primarily caused by the individual seed variation, which we choose to disregard
here. The second highest effect stems from variety, followed by the two sample
presentation parameters of position and orientation. Based on the permutation test,
we can determine the statistical significance of each of the factors. The orientation
factor is found not to be significant due to its high p-value. However, it seems that
variety and position both are highly important and are contributing significantly to
the variation in the dataset.
It can be concluded that the position is an important experimental factor and can
be directly compared to the more important variety effect. As previously mentioned,
the orientation factor describes the orientation of the single seed in the instrument—
which is thus a factor that should be controlled in the future use of the instrument or
in, e.g., a high throughput scenario with single-seed sorting!
K. M. Sørensen et al.
the random variation. If formalized as a PLS-DA problem, this is known as multi-level
PLS-DA [55].
A limitation of the application of ASCA is that it is most optimally applied to
balanced design structures. On the other hand, the statistical significance of the
strength of each design factor and their interactions can be evaluated using a permutation test [56], thus providing a link to classical statistics. Needless to say that
appropriate pre-processing and centering of the X matrix will be essential to the
subsequent ASCA.
7.8.1 Application of ASCA to NIR Spectra
Since the ASCA description above can seem a little theoretical, it will be illustrative
to demonstrate the potential of ASCA from an example. Here, it is used to partition
(split) the variance due to different sampling geometries in Dataset 4. The ASCA
model thus reflects the experimental factors shown in Eq. 7.32 and Fig. 7.33. Only
three factors are included in the model, and the “individual” factor is left out.
It will make no sense to include “individual” as a design factor in the ASCA
model, as the different individuals are not the same for the different varieties. They
are physically different kernels and can thus not be compared across varieties. This
leaves us to focus on the partitioned effects of the three design factors:
X = X variety + X position + X orientation + E
(7.32)
The ASCA model will be computed on Dataset 4 with the spectra X, pre-processed
by a 2nd order, derivative Savitzky-Golay filter of window width 7, followed by a
MSC, and finally a mean centering. The effect is evaluated with 1000 permutations.
For simplicity, interactions between the three factors are not included. The full rank
of each factor is also calculated.
The primary output of the ASCA model is the effect table (Table 7.2). From the
table, it is clear that largest effect is found in the residual (81.1%). In the current case,
this is primarily caused by the individual seed variation, which we choose to disregard
here. The second highest effect stems from variety, followed by the two sample
presentation parameters of position and orientation. Based on the permutation test,
we can determine the statistical significance of each of the factors. The orientation
factor is found not to be significant due to its high p-value. However, it seems that
variety and position both are highly important and are contributing significantly to
the variation in the dataset.
It can be concluded that the position is an important experimental factor and can
be directly compared to the more important variety effect. As previously mentioned,
the orientation factor describes the orientation of the single seed in the instrument—
which is thus a factor that should be controlled in the future use of the instrument or
in, e.g., a high throughput scenario with single-seed sorting!
