38
Internet of Things (IoT)
divide by the cardinality of the known classification. The inclusive nature of upper bound
provides us with higher recall values. On the contrary, the exclusivity of lower bounds
leads to lower recall values.
2.10 Conclusion
Wearable technology provides exciting opportunities for data mining. This study demonstrates the viability of supervised, unsupervised, and semi-supervised learning techniques
for identifying individuals and activities based on EEG data collected from a commercially
available wearable headband. In a real-world scenario, the signals will have even more variation in length than what has been evaluated here. Most data mining techniques are based
on fixed-length object representation. This chapter shows that by using histograms of EEG
brain signals, we can apply a wide variety of data mining techniques. Histograms are very
advantageous as they reduce the raw variable length data to fixed-length representations.
No matter the length of the data recorded, the representation will be the same.
All four supervised learning techniques explored were successful in identifying persons
based on their collected EEG brain signals. The ten fold cross-validation results varied
from 95% to 93% for SVM, neural networks, and random forest. Decision tree provided
the least precision with 65%. SVM’s precision of 78% for activities and 85% for person and
activity showed that the proposed representation is quite credible for predicting activities from brain signals as well. Optimizing and tuning the classifier parameters may yield
increased accuracy and precision; however, that was not explored in this chapter.
The traditional clustering methods had precision values ranging from 27% to 82%
for K-means, and improved values of 68% to 95% for K-medoids. The proposed rough
K-medoids evolutionary clustering method provided varying results. The precision of the
lower bounds had a very large range, from 0% to 100%. Recall values were lower overall
for K-means, with a range of 40% to 98%, as compared to K-medoids, with a range of 65%
to 96%. For the proposed rough K-medoids algorithm, the recall for the upper bound was
overall higher, but with a wide range from 13% to 100%.
The proposed rough K-medoids algorithm provided widely varying results depending
on the person. This is demonstrated by the 0% recall and precision values for person 3 and
the 100% precision value for person 0 in the lower bound. Correspondingly, in the upper
bound, the variation is demonstrated by the 6% precision and 13% recall for person 0 and
the 100% recall for person 1. With rough clustering, the upper bounds are inclusive in
nature. If a pattern has reasonable chance of belonging to a cluster, it goes into its upper
bound. This leads to generally higher recall values. The lower bounds, on the contrary,
are exclusive. A pattern goes into the lower bound of a cluster only if there is a very high
chance that the pattern belongs to the cluster. This leads to lower recall and higher precision values for the lower bounds of clustering. An analyst can use the upper bounds of a
cluster when higher recall is desired. Lower bounds of the clusters will be useful when the
precision is an overwhelming criterion.
All clustering methods were reasonably successful with matching clusters to persons.
K-medoids provided the most reliable results, while the proposed rough K-medoids algorithm had both very good and very poor results. This suggests that further optimization
and tuning of the algorithm may be necessary.
Internet of Things (IoT)
divide by the cardinality of the known classification. The inclusive nature of upper bound
provides us with higher recall values. On the contrary, the exclusivity of lower bounds
leads to lower recall values.
2.10 Conclusion
Wearable technology provides exciting opportunities for data mining. This study demonstrates the viability of supervised, unsupervised, and semi-supervised learning techniques
for identifying individuals and activities based on EEG data collected from a commercially
available wearable headband. In a real-world scenario, the signals will have even more variation in length than what has been evaluated here. Most data mining techniques are based
on fixed-length object representation. This chapter shows that by using histograms of EEG
brain signals, we can apply a wide variety of data mining techniques. Histograms are very
advantageous as they reduce the raw variable length data to fixed-length representations.
No matter the length of the data recorded, the representation will be the same.
All four supervised learning techniques explored were successful in identifying persons
based on their collected EEG brain signals. The ten fold cross-validation results varied
from 95% to 93% for SVM, neural networks, and random forest. Decision tree provided
the least precision with 65%. SVM’s precision of 78% for activities and 85% for person and
activity showed that the proposed representation is quite credible for predicting activities from brain signals as well. Optimizing and tuning the classifier parameters may yield
increased accuracy and precision; however, that was not explored in this chapter.
The traditional clustering methods had precision values ranging from 27% to 82%
for K-means, and improved values of 68% to 95% for K-medoids. The proposed rough
K-medoids evolutionary clustering method provided varying results. The precision of the
lower bounds had a very large range, from 0% to 100%. Recall values were lower overall
for K-means, with a range of 40% to 98%, as compared to K-medoids, with a range of 65%
to 96%. For the proposed rough K-medoids algorithm, the recall for the upper bound was
overall higher, but with a wide range from 13% to 100%.
The proposed rough K-medoids algorithm provided widely varying results depending
on the person. This is demonstrated by the 0% recall and precision values for person 3 and
the 100% precision value for person 0 in the lower bound. Correspondingly, in the upper
bound, the variation is demonstrated by the 6% precision and 13% recall for person 0 and
the 100% recall for person 1. With rough clustering, the upper bounds are inclusive in
nature. If a pattern has reasonable chance of belonging to a cluster, it goes into its upper
bound. This leads to generally higher recall values. The lower bounds, on the contrary,
are exclusive. A pattern goes into the lower bound of a cluster only if there is a very high
chance that the pattern belongs to the cluster. This leads to lower recall and higher precision values for the lower bounds of clustering. An analyst can use the upper bounds of a
cluster when higher recall is desired. Lower bounds of the clusters will be useful when the
precision is an overwhelming criterion.
All clustering methods were reasonably successful with matching clusters to persons.
K-medoids provided the most reliable results, while the proposed rough K-medoids algorithm had both very good and very poor results. This suggests that further optimization
and tuning of the algorithm may be necessary.
