Chapter 21
(hemometrics for Sampling and Analysis:
Theory and Environmental Applications
M. Forina . S. Lanteri . R. Todeschini
21.1
Introduction
Chemometrics is the chemical discipline that uses mathematics, statistics, artificial
intelligence and formal logic
• to design or select optimal experiments,
• to extract maximum relevant information from chemical measured or computed data,
and
• to obtain knowledge about chemical complex systems, processes and products.
Chemometrics is characterized by:
• the data base,
• the tools,
• the models and
• the objective.
Data for chemometrics are complex, multivariate, in form of a matrix of many rows
(samples, molecules), described by many columns (measured physical or chemical
quantities, computed quantities) and frequently tridimensional (e.g. sampling sites,
time, measured quantities). The below reported examples will generally use only two
variables for representation purposes. Hundreds or thousands of variables often describe real objects: e.g. a spectrophotometer can give the absorbance at about one thousand wavelengths. GRID, a technique used to describe molecules by interaction potentials, can produce 10 000 or more descriptors for each studied molecule.
The variability of data between objects depends on many known or unknown factors and on measurement error (afterward simply "noise"). The information content
of the data is closely related to the factor-dependent variability (afterward simply "variability"). When the noise is higher than the variability, data cannot contain information. The situation is that of "quality" - all objects are equal within the noise:
• Variability = Information
• Non-variability = Quality
The tools of chemometrics (Massart et al.I998; Meloun et al. 1992; Frank and
Todeschini 1994) range from those of the base theory of probability (with particular
attention to the special case of experimental data, characterized by the discontinuity
Précédent

- 393/447

Suivant