4 The Analysis of Event-Related Potentials
77
received a strong impulsion thanks to development of ERP-based brain computer
interfaces (BCI: [101]). In fact, a popular family of such interfaces is based on the
recognition of the P300 ERP. The most famous example is the P300 Speller [34], a
system allowing the user to spell text without moving, but just by focusing attention
on symbols (e.g., letters) that are flashed on a virtual keyboard.
The fundamental criterion for choosing a classification method is the achieved
accuracy for the data at hand. However, other criteria may be relevant. In BCI systems,
the training of the classifier starts with a calibration session carried out just before the
actual session. Such calibration phase makes the usage of BCI system impractical
and annoying. To avoid this, there are at least two other desirable characteristics
that a classification method should possess [60]: its ability to generalize and its
ability to adapt. Generalization allows the so-called transfer learning, thanks to
which data from other sessions and/or other subjects can be used to initialize a BCI
system so as to avoid the calibration phase. Transfer learning may involve using data
from previous sessions of the same subject (“cross-session”) and/or data from other
subjects (“cross-subject”). The continuous (on-line) adaptation of the classifier [47,
48] ensures that optimal performance is achieved once the initialization is obtained by
transfer learning [18]. Taken together, generalization and on-line adaptation ensure
also the stability of the system in adverse situations, that is, when the SNR of the
incoming data is low and when there are sudden environmental, instrumental or
biological changes during the session. This is very important for effective use of a
BCI outside the controlled environment of research laboratories.
Classification methods differ from each other in the way they define the set of features and in the discriminant function they employ. Traditionally, the classification
approaches for ERPs have given emphasis to the optimization of either one or the
other aspect in order to increase accuracy. The approaches emphasizing the definition
of the set of features try to increase the SNR of single-sweeps by using multivariate
filtering, as those we have encountered in the section on time domain analysis, but
specifically designed to increase the separation of the classes in a reduced feature
space where the filter projects the data [83, 104]. For data filtered in this way, the
choice of the discriminant function is not critical, in the sense that similar accuracy is
obtained using several types of discriminant functions. In general, these approaches
perform well even if the training set is small, but generalize poorly across sessions
and across subjects because the spatial filters are optimal only for the session and
subject on whom they are estimated. Instead, the approaches emphasizing the discriminant function use sharp machine learning algorithms on raw data or on data that
has underwent little-preprocessing. Many machine learning algorithms have been
tried in the BCI literature for this purpose [59, 60]. The three traditional approaches
that have been found effective in P300 single-sweep classification are the supportvector machine, the stepwise linear discriminant analysis and the Bayesian linear
discriminant analysis. In general, those require large training sets and have high
computational complexity, but generalize fairly well across sessions and across subjects. The use of a random forest classifier is currently gaining popularity in the BCI
community, incited by good accuracy properties [33]. However, its generalization
and adaptation capability have not been established yet. The deep neural networks
77
received a strong impulsion thanks to development of ERP-based brain computer
interfaces (BCI: [101]). In fact, a popular family of such interfaces is based on the
recognition of the P300 ERP. The most famous example is the P300 Speller [34], a
system allowing the user to spell text without moving, but just by focusing attention
on symbols (e.g., letters) that are flashed on a virtual keyboard.
The fundamental criterion for choosing a classification method is the achieved
accuracy for the data at hand. However, other criteria may be relevant. In BCI systems,
the training of the classifier starts with a calibration session carried out just before the
actual session. Such calibration phase makes the usage of BCI system impractical
and annoying. To avoid this, there are at least two other desirable characteristics
that a classification method should possess [60]: its ability to generalize and its
ability to adapt. Generalization allows the so-called transfer learning, thanks to
which data from other sessions and/or other subjects can be used to initialize a BCI
system so as to avoid the calibration phase. Transfer learning may involve using data
from previous sessions of the same subject (“cross-session”) and/or data from other
subjects (“cross-subject”). The continuous (on-line) adaptation of the classifier [47,
48] ensures that optimal performance is achieved once the initialization is obtained by
transfer learning [18]. Taken together, generalization and on-line adaptation ensure
also the stability of the system in adverse situations, that is, when the SNR of the
incoming data is low and when there are sudden environmental, instrumental or
biological changes during the session. This is very important for effective use of a
BCI outside the controlled environment of research laboratories.
Classification methods differ from each other in the way they define the set of features and in the discriminant function they employ. Traditionally, the classification
approaches for ERPs have given emphasis to the optimization of either one or the
other aspect in order to increase accuracy. The approaches emphasizing the definition
of the set of features try to increase the SNR of single-sweeps by using multivariate
filtering, as those we have encountered in the section on time domain analysis, but
specifically designed to increase the separation of the classes in a reduced feature
space where the filter projects the data [83, 104]. For data filtered in this way, the
choice of the discriminant function is not critical, in the sense that similar accuracy is
obtained using several types of discriminant functions. In general, these approaches
perform well even if the training set is small, but generalize poorly across sessions
and across subjects because the spatial filters are optimal only for the session and
subject on whom they are estimated. Instead, the approaches emphasizing the discriminant function use sharp machine learning algorithms on raw data or on data that
has underwent little-preprocessing. Many machine learning algorithms have been
tried in the BCI literature for this purpose [59, 60]. The three traditional approaches
that have been found effective in P300 single-sweep classification are the supportvector machine, the stepwise linear discriminant analysis and the Bayesian linear
discriminant analysis. In general, those require large training sets and have high
computational complexity, but generalize fairly well across sessions and across subjects. The use of a random forest classifier is currently gaining popularity in the BCI
community, incited by good accuracy properties [33]. However, its generalization
and adaptation capability have not been established yet. The deep neural networks
