30
Internet of Things (IoT)
2.7.2 Classification
The dataset used to develop the classification models therefore consists of 178 records.
Each record contains 160 attributes. We used the following four well-known classifiers:
1. Decision tree
2. Random forest
3. Support vector machines
4. Neural networks
The classifiers were trained using default parameters and options. Each classifier was
applied to three predictions:
1. Predicting a person—Five classes corresponding to five persons
2. Predicting an activity—Five classes corresponding to five activities
• Doing nothing
• Playing a game
• Listening to music
• Watching a video
• Reading
3. Predicting a person as well as the activity—Twenty-five classes corresponding to
the cross-product of the set of persons and the set of activities.
To test how well our knowledge representation and classifiers describe the activities,
we first applied the classifiers to the entire dataset. This allowed us to study the importance of different attributes in the classification process.
In Tables 2.4(a), 2.4(b), and 2.5, you can see the most significant independent variables
obtained by the random forest classifier when classifying by person, activity, or person and
activity, respectively.
Using a complete dataset for training can lead to overtraining of the classifiers. The classifier may not work for new datasets. In order to see if the models were general enough to
predict new recordings, we applied ten fold cross-validation.
In ten fold cross-validation [28], 10% of randomly selected data from the dataset is set
aside for testing the model. The remaining 90% of the dataset is used for training the model.
This process is repeated 10 times and the results are summarized.
From our tests we learned that our knowledge representation paired with SVM, random
forest, or neural networks could predict a class for person or activity with fairly high
precision. However, using a decision tree classifier proved to be ineffective.
The primary objective of the study was to explore the effectiveness of a number of wellknown classification techniques combined with a data representation that reduces variable
length signals to more manageable and uniformly fixed-length distributions. Using this representation and the classification techniques, the hypothesis was that the EEG brain signals
will be identifiable for both persons and activities with a reasonable degree of precision.
2.7.3 Semi-Supervised Evolutionary Learning
We used the two-point crossover technique for the genetic algorithm. The two-point
crossover technique uses two points to divide both of the parents’ genomes into three
Internet of Things (IoT)
2.7.2 Classification
The dataset used to develop the classification models therefore consists of 178 records.
Each record contains 160 attributes. We used the following four well-known classifiers:
1. Decision tree
2. Random forest
3. Support vector machines
4. Neural networks
The classifiers were trained using default parameters and options. Each classifier was
applied to three predictions:
1. Predicting a person—Five classes corresponding to five persons
2. Predicting an activity—Five classes corresponding to five activities
• Doing nothing
• Playing a game
• Listening to music
• Watching a video
• Reading
3. Predicting a person as well as the activity—Twenty-five classes corresponding to
the cross-product of the set of persons and the set of activities.
To test how well our knowledge representation and classifiers describe the activities,
we first applied the classifiers to the entire dataset. This allowed us to study the importance of different attributes in the classification process.
In Tables 2.4(a), 2.4(b), and 2.5, you can see the most significant independent variables
obtained by the random forest classifier when classifying by person, activity, or person and
activity, respectively.
Using a complete dataset for training can lead to overtraining of the classifiers. The classifier may not work for new datasets. In order to see if the models were general enough to
predict new recordings, we applied ten fold cross-validation.
In ten fold cross-validation [28], 10% of randomly selected data from the dataset is set
aside for testing the model. The remaining 90% of the dataset is used for training the model.
This process is repeated 10 times and the results are summarized.
From our tests we learned that our knowledge representation paired with SVM, random
forest, or neural networks could predict a class for person or activity with fairly high
precision. However, using a decision tree classifier proved to be ineffective.
The primary objective of the study was to explore the effectiveness of a number of wellknown classification techniques combined with a data representation that reduces variable
length signals to more manageable and uniformly fixed-length distributions. Using this representation and the classification techniques, the hypothesis was that the EEG brain signals
will be identifiable for both persons and activities with a reasonable degree of precision.
2.7.3 Semi-Supervised Evolutionary Learning
We used the two-point crossover technique for the genetic algorithm. The two-point
crossover technique uses two points to divide both of the parents’ genomes into three
