147
Artificial Neural Networks (ANN) ANN are defined as structures comprising
densely interconnected adaptive simple processing elements, called artificial neurons (or nodes), which can perform massively parallel computations for data processing and knowledge representation.
Classification and Regression Trees (CARTs) CARTs is a powerful and flexible
classification tool. It handles both ordered and categorical predictor variables. The
final classification rule has a simple form, which is easy to interpret and to use for
future classification. CARTs takes into account the fact that different relationships
may hold between variables in different parts of the data. It does automatic stepwise
variable selection and calculates the importance rank of the variables. CARTs calculates misclassification error estimated by both re-substitution and
cross-validation.
Linear Discriminant Analysis (LDA) LDA is probably the most frequently used
supervised pattern recognition method and the most studied one. LDA is based on
the determination of linear discriminant functions, which maximize the ratio of
between-class variance and minimize the ratio of within-class variance. In LDA,
classes are supposed to follow a multivariate normal distribution and be linearly
separated. LDA can be considered, as PCA, a feature-reduction method in the sense
that both, LDA and PCA, determine a smaller dimension hyperplane on which the
points will be projected from the higher dimension.
Cluster Analysis (CA) clustering is the task of assigning a set of objects into groups
(called clusters) so that the objects in the same cluster are more similar (in some
sense or another) to each other than to those in other clusters. In this way, we categorize the samples in order to succeed with discrimination.
Principal Component Analysis (PCA) PCA is a technique that, by the reduction of
the data dimensionality, allows their visualization while retaining as much as possible the information present in the original data. So, PCA transforms the original
measured variables into new uncorrelated variables called principal components
(PCs). Each PC is a linear combination of the original measured variables. This
technique affords a group of orthogonal axes that represent the directions of greatest
variance in the data. The first PC (PC1) accounts for the maximum of the total
variance, the second (PC2) is uncorrelated with the first and accounts for the maximum of the residual variance, and so on, until the total variance is accounted for.
Analysis of Variance (ANOVA) ANOVA is a collection of statistical models, and
their associated procedures, in which the observed variance in a particular variable
is partitioned into components attributable to different sources of variation. In its
simplest form, ANOVA provides a statistical test of whether or not the means of
several groups are all equal and therefore generalizes t-test to more than two groups.
Performing multiple two-sample t-tests would result in an increased chance of committing a type I error. For this reason, ANOVA is useful in comparing two, three, or
more means and has been used to compare elemental profiles of foods of different
origin.
5 Proposing Chemometric Tool for Efficacy Surface Dust Deposition Tracking…
Artificial Neural Networks (ANN) ANN are defined as structures comprising
densely interconnected adaptive simple processing elements, called artificial neurons (or nodes), which can perform massively parallel computations for data processing and knowledge representation.
Classification and Regression Trees (CARTs) CARTs is a powerful and flexible
classification tool. It handles both ordered and categorical predictor variables. The
final classification rule has a simple form, which is easy to interpret and to use for
future classification. CARTs takes into account the fact that different relationships
may hold between variables in different parts of the data. It does automatic stepwise
variable selection and calculates the importance rank of the variables. CARTs calculates misclassification error estimated by both re-substitution and
cross-validation.
Linear Discriminant Analysis (LDA) LDA is probably the most frequently used
supervised pattern recognition method and the most studied one. LDA is based on
the determination of linear discriminant functions, which maximize the ratio of
between-class variance and minimize the ratio of within-class variance. In LDA,
classes are supposed to follow a multivariate normal distribution and be linearly
separated. LDA can be considered, as PCA, a feature-reduction method in the sense
that both, LDA and PCA, determine a smaller dimension hyperplane on which the
points will be projected from the higher dimension.
Cluster Analysis (CA) clustering is the task of assigning a set of objects into groups
(called clusters) so that the objects in the same cluster are more similar (in some
sense or another) to each other than to those in other clusters. In this way, we categorize the samples in order to succeed with discrimination.
Principal Component Analysis (PCA) PCA is a technique that, by the reduction of
the data dimensionality, allows their visualization while retaining as much as possible the information present in the original data. So, PCA transforms the original
measured variables into new uncorrelated variables called principal components
(PCs). Each PC is a linear combination of the original measured variables. This
technique affords a group of orthogonal axes that represent the directions of greatest
variance in the data. The first PC (PC1) accounts for the maximum of the total
variance, the second (PC2) is uncorrelated with the first and accounts for the maximum of the residual variance, and so on, until the total variance is accounted for.
Analysis of Variance (ANOVA) ANOVA is a collection of statistical models, and
their associated procedures, in which the observed variance in a particular variable
is partitioned into components attributable to different sources of variation. In its
simplest form, ANOVA provides a statistical test of whether or not the means of
several groups are all equal and therefore generalizes t-test to more than two groups.
Performing multiple two-sample t-tests would result in an increased chance of committing a type I error. For this reason, ANOVA is useful in comparing two, three, or
more means and has been used to compare elemental profiles of foods of different
origin.
5 Proposing Chemometric Tool for Efficacy Surface Dust Deposition Tracking…
