4.3 Chemometrics
Chemometrics is the chemical discipline that uses mathematics and statistics to
design or select optimal experimental procedures and to obtain knowledge about a
chemical system, such as a water matrix. In spectroscopy the primary uses for data
analysis algorithms are grouping and classification and modelling relationships
between different analytical data. Examples include the classification of samples,
such as chemical compounds or materials, based on spectra, and the building of
calibration models for calculation of concentrations of chemical constituents in a
mixture, e.g. a water sample. The superposition of numerous single substance
signals in real-world samples causes cross-sensitivities; when the concentration of
an analyte is directly deduced from the signal at one individual wavelength, it will
often respond to other, non-related, variations in the matrix. Chemometric models
are required to extract information on the specific parameters from the spectra.
The most frequently used method to develop calibration models is indirect
modelling using multivariate analysis [10]. In indirect modelling, a calibration
model is built from a dataset containing spectral data and the concentrations of the
parameters of interest. These concentration values are acquired through separate
analytical methods, such as the standard methods for water quality analysis [11]. The
multivariate approach has the advantage that interactions between analytes, or
between analytes and the matrix, can be accounted for in the calibration model.
Also, indirect modelling can deal with any correlations between target analytes. In
water applications such correlations are often quite prevalent, such as, for example,
the strong correlation between chemical oxygen demand (COD) and total suspended
solids (TSS) in wastewater. The first step in the indirect modelling approach is the
grouping of analytical data into clusters. An example of a projection method often
used in spectroscopy is principal component analysis (PCA) [9]. PCA reduces the
dimensionality of a set of variables while conserving the variability within the data
as much as possible. In other words, PCA tries to explain the variance-covariance
structure of the data using a new coordinate system that is lesser in dimension than
the number of original variables; spectra typically consist of 200+ wavelength
measurements (variables). The deduced principal components (PC) are new,
uncorrelated, orthogonal variables that describe a maximum of variance in the
dataset. Another method widely used in spectroscopy is partial least squares (PLS)
regression [8]. Working in a similar way as PCA, PLS reduces a complex
multidimensional dataset into a smaller number of components accounting for as
much variation as possible while also modelling the Y-variables (the reference
values). Because PLS is suited for cases with insufficient data to construct a model
to predict all variability, it is especially popular in industrial applications where
sufficiently complete datasets are often impossible to obtain. Both PCA and PLS are
combined with cross-validation procedures and outlier tests to reach both high
correlation quality and robustness of the model [8]. The result of the calibration
procedure is a function describing how to combine a selection of wavelengths to
calculate the target variables. The goodness of fit is described in the recovery
Spectroscopic Methods for Online Water Quality Monitoring
291
Chemometrics is the chemical discipline that uses mathematics and statistics to
design or select optimal experimental procedures and to obtain knowledge about a
chemical system, such as a water matrix. In spectroscopy the primary uses for data
analysis algorithms are grouping and classification and modelling relationships
between different analytical data. Examples include the classification of samples,
such as chemical compounds or materials, based on spectra, and the building of
calibration models for calculation of concentrations of chemical constituents in a
mixture, e.g. a water sample. The superposition of numerous single substance
signals in real-world samples causes cross-sensitivities; when the concentration of
an analyte is directly deduced from the signal at one individual wavelength, it will
often respond to other, non-related, variations in the matrix. Chemometric models
are required to extract information on the specific parameters from the spectra.
The most frequently used method to develop calibration models is indirect
modelling using multivariate analysis [10]. In indirect modelling, a calibration
model is built from a dataset containing spectral data and the concentrations of the
parameters of interest. These concentration values are acquired through separate
analytical methods, such as the standard methods for water quality analysis [11]. The
multivariate approach has the advantage that interactions between analytes, or
between analytes and the matrix, can be accounted for in the calibration model.
Also, indirect modelling can deal with any correlations between target analytes. In
water applications such correlations are often quite prevalent, such as, for example,
the strong correlation between chemical oxygen demand (COD) and total suspended
solids (TSS) in wastewater. The first step in the indirect modelling approach is the
grouping of analytical data into clusters. An example of a projection method often
used in spectroscopy is principal component analysis (PCA) [9]. PCA reduces the
dimensionality of a set of variables while conserving the variability within the data
as much as possible. In other words, PCA tries to explain the variance-covariance
structure of the data using a new coordinate system that is lesser in dimension than
the number of original variables; spectra typically consist of 200+ wavelength
measurements (variables). The deduced principal components (PC) are new,
uncorrelated, orthogonal variables that describe a maximum of variance in the
dataset. Another method widely used in spectroscopy is partial least squares (PLS)
regression [8]. Working in a similar way as PCA, PLS reduces a complex
multidimensional dataset into a smaller number of components accounting for as
much variation as possible while also modelling the Y-variables (the reference
values). Because PLS is suited for cases with insufficient data to construct a model
to predict all variability, it is especially popular in industrial applications where
sufficiently complete datasets are often impossible to obtain. Both PCA and PLS are
combined with cross-validation procedures and outlier tests to reach both high
correlation quality and robustness of the model [8]. The result of the calibration
procedure is a function describing how to combine a selection of wavelengths to
calculate the target variables. The goodness of fit is described in the recovery
Spectroscopic Methods for Online Water Quality Monitoring
291
