21
Identification from Wearable Device Brain Signals
The stopping criteria for the recursive decision tree algorithm is commonly a threshold
for the minimum number of remaining instances in a node’s partition or a threshold for
the minimum difference in impurity calculated between the parent and child nodes [10].
Overfitting occurs when a classification model performs well when tested against training
data but poorly when applied to unseen data. Tree pruning techniques [10,11] and statistical stopping criteria [10] have been used to mitigate overfitting.
2.3.2 Random Forest
The random forest technique, proposed by Breiman (2001), leverages multiple decision
trees to predict an outcome [12, 34]. Its output is determined by the prediction that appears
the most often in each of the individual decision trees [10,12]. Multiple trees, or an ensemble of trees, can be used to mitigate the instability of a single decision tree [10]. An ensemble of trees is created with random samples picked from the input training data [10,11].
The  instances excluded with each random sample can be considered “out-of-bag” and
used as test samples for measuring out-of-bag prediction accuracy.
Tree ensembles minimize overfitting with a set of diverse trees that tend to converge
when the set is sufficiently large [12]. By randomly restricting the attributes used to generate the trees, attributes that would otherwise not have been chosen in a single decision
tree can result in the discovery of cross-attribute correlations and patterns that otherwise
would have been missed [10]. This has the potential to improve global prediction and
accuracy.
2.3.3 Support Vector Machine
Support vector machine is used for binary classification [3,13–15]. The attributes of input
training data are referred to as features. SVM works by first mapping the input data into
a higher dimensional feature space. The SVM model then works to produce an optimal
hyperplane in the new high dimensional feature space. The hyperplane separates the data
into two groups, representing the two classes of the input data. In a two-dimensional space,
we can separate instances into two groups with a line. In a higher dimensional space, we
use hyperplanes. An optimal hyperplane maximizes the margin, or separation, between
the two groups. The dataset used in this chapter has more than two classes. There are five
classes of activities and four classes representing the persons. Since SVM are explicitly
designed to classify into two groups, a specialized approach is required to handle multiple
classes in the classification dataset, referred to as multiclass classification.
The implementation of SVM used in this chapter used the “one-against one” or “oneversus-one” approach for multiclass classification [16]. For k classes, there are
k k −
( 1)
2
binary classifiers trained, and then a voting scheme decides the appropriate single class
predicted [13,16]. Each binary classifier is given a newly constructed training dataset, such
that one class is considered the positive class and another class is considered the negative
class [13]. Finally, each of the binary classifiers can vote on the single class that they predict
is correct, and the class with the most votes is considered the combined prediction.
2.3.4 Neural Network
Artificial neural networks, used for information processing, are inspired by the interactions within the biological nervous system of the human brain [17]. Neurons are connected
Précédent

- 46/358

Suivant