internal nodes are called decision nodes. Each decision node is labeled by a test which
can be applied to any description of an individual in the population.
2.2.2 Support Vector Machine (SVM)
Support Vector Machines is a phenomenon f (possibly non-deterministic) which, from
a certain set of inputs x, produces an output y ¼ f ðxÞ. This approach, often translated
by the name of Support Vector Machine (SVM), is a class of learning algorithms
initially defined for discrimination and prediction of a binary qualitative variable. The
main objective is to find f from the only observation of a certain number of input-output
pairs fðxi; yiÞ: i ¼ 1; . . .; ng. Among its advantages, SVM overcomes various common problems related to the recognition of shapes.
2.2.3 K-Nearest Neighbors (KNN)
The principle of this model consists in choosing the k data closest to the point studied
in order to predict its value. The objective is to make a classification without making a
hypothesis on the function y ¼ f x1; x2; . . .; xn
ð
Þ which links the dependent variable y
to the independent variables x1; x2; . . .; xn. Otherwise, the idea of the KNN algorithm
is for a new observation (u1; u2; . . .; up) to predict the k observations that are most
similar to it in the training data [1].
2.2.4 Multi-nominal Naïve Bayes (MNNB)
The Multi-Nominal Naïve Bayes classifier is derived from Bayesian decision theory. It
is a fundamental statistical approach in pattern recognition. Bayesian decision theory
chooses the best decision among the possible decisions based on these laws and the
costs associated with each decision. The objective consists in finding a decision rule
which minimizes an average cost and in defining which decision (action) to take
according to the observed entity.
2.2.5 Random Forest (RF)
The algorithm of “random forests” was proposed by Leo Breiman and Adèle Cutler in
2001 [3]. It performs parallel learning on multiple decision trees randomly constructed
and trained on subsets of data different. The ideal number of trees, which can go up to
several hundred or more, is an important parameter: it is very variable and depends on
the problem.
2.2.6 Gradient Boosting (GB)
This boosting technique is mainly used with decision trees (it is then called Gradient
Tree Boosting). Again, the main idea is to aggregate several classifiers together but to
create them iteratively. These “mini-classifiers” are generally simple and parameterized
functions, most often decision trees, each parameter of which is the split criterion of the
branches.
350
M. Khadhraoui et al.
can be applied to any description of an individual in the population.
2.2.2 Support Vector Machine (SVM)
Support Vector Machines is a phenomenon f (possibly non-deterministic) which, from
a certain set of inputs x, produces an output y ¼ f ðxÞ. This approach, often translated
by the name of Support Vector Machine (SVM), is a class of learning algorithms
initially defined for discrimination and prediction of a binary qualitative variable. The
main objective is to find f from the only observation of a certain number of input-output
pairs fðxi; yiÞ: i ¼ 1; . . .; ng. Among its advantages, SVM overcomes various common problems related to the recognition of shapes.
2.2.3 K-Nearest Neighbors (KNN)
The principle of this model consists in choosing the k data closest to the point studied
in order to predict its value. The objective is to make a classification without making a
hypothesis on the function y ¼ f x1; x2; . . .; xn
ð
Þ which links the dependent variable y
to the independent variables x1; x2; . . .; xn. Otherwise, the idea of the KNN algorithm
is for a new observation (u1; u2; . . .; up) to predict the k observations that are most
similar to it in the training data [1].
2.2.4 Multi-nominal Naïve Bayes (MNNB)
The Multi-Nominal Naïve Bayes classifier is derived from Bayesian decision theory. It
is a fundamental statistical approach in pattern recognition. Bayesian decision theory
chooses the best decision among the possible decisions based on these laws and the
costs associated with each decision. The objective consists in finding a decision rule
which minimizes an average cost and in defining which decision (action) to take
according to the observed entity.
2.2.5 Random Forest (RF)
The algorithm of “random forests” was proposed by Leo Breiman and Adèle Cutler in
2001 [3]. It performs parallel learning on multiple decision trees randomly constructed
and trained on subsets of data different. The ideal number of trees, which can go up to
several hundred or more, is an important parameter: it is very variable and depends on
the problem.
2.2.6 Gradient Boosting (GB)
This boosting technique is mainly used with decision trees (it is then called Gradient
Tree Boosting). Again, the main idea is to aggregate several classifiers together but to
create them iteratively. These “mini-classifiers” are generally simple and parameterized
functions, most often decision trees, each parameter of which is the split criterion of the
branches.
350
M. Khadhraoui et al.
