286
B. Igne et al.
NIR) [7]. The ASTM E55.01 committee has also issued several documents on the
topic [10, 11]. Proper testing of any model requires independence of the calibration
and representative samples used for performance evaluation.
12.3.5.2 Sample Variability or Origin
Samples for building a NIR spectroscopic model may come from natural sources
or manufacturing sources. In cases where it is difficult to design specific samples
for model development, care should be taken to ensure that the calibration set is
well balanced and not over representing a particular source of variability that may
artificially influence the predictions.
The modeling team should ensure that all relevant variability is included in the
calibration set. Algorithms such as the Kennard and Stone [13] sample selection
approach can help reduce redundancy in spectra presenting the same variability for
large datasets but the team must understand the data included and excluded, and
rationalize the sample selection.
When feasible, particularly for the chemical and pharmaceutical industries, artificial samples can be produced at small or pilot scales. Designs of experiments are
often used to derive a suitable spectral space by varying the chemical and physical
properties in the samples to represent normal operating ranges of the process and that
the model should be expected to suitably handle. Figure 12.4 provides an example of
such a design (a 3-factor 2-level full factorial design with a center point) where the
Fig. 12.4 Example of full factorial design commonly used for designing synthetic samples
B. Igne et al.
NIR) [7]. The ASTM E55.01 committee has also issued several documents on the
topic [10, 11]. Proper testing of any model requires independence of the calibration
and representative samples used for performance evaluation.
12.3.5.2 Sample Variability or Origin
Samples for building a NIR spectroscopic model may come from natural sources
or manufacturing sources. In cases where it is difficult to design specific samples
for model development, care should be taken to ensure that the calibration set is
well balanced and not over representing a particular source of variability that may
artificially influence the predictions.
The modeling team should ensure that all relevant variability is included in the
calibration set. Algorithms such as the Kennard and Stone [13] sample selection
approach can help reduce redundancy in spectra presenting the same variability for
large datasets but the team must understand the data included and excluded, and
rationalize the sample selection.
When feasible, particularly for the chemical and pharmaceutical industries, artificial samples can be produced at small or pilot scales. Designs of experiments are
often used to derive a suitable spectral space by varying the chemical and physical
properties in the samples to represent normal operating ranges of the process and that
the model should be expected to suitably handle. Figure 12.4 provides an example of
such a design (a 3-factor 2-level full factorial design with a center point) where the
Fig. 12.4 Example of full factorial design commonly used for designing synthetic samples
