76
M. Congedo
Another kind of FWER-controlling permutation test in ERP analysis is the suprathreshold cluster size test [43]. A variant of this test has been implemented in the
EEG toolbox Fieldtrip [71], following the review of Maris and Oostenveld [63]. This
procedure assesses the probability to observe a concentration of the effect simultaneously along one or more studied dimensions. For example, in testing the mean
amplitude difference of a P300 ERP, one expects the effect to be concentrated both
along time, around 300–500 ms, and along space, at midline central and adjacent
parietal locations. This leads to a typical correlation structure of hypothesis in ERP
data; under the null hypothesis the effect would instead be scattered all over both time
and spatial dimensions. An example of the supra-threshold cluster size test applied
in the time-space ERP domain is shown in Fig. 4.8.
Another family of testing procedures controls the false discovery rate (FDR). The
FDR is the expected proportion of falsely rejected hypotheses [3]. Indicating by R
the number of rejected hypotheses and by F the number of those that have been
falsely rejected, the FDR controls the expectation of the ratio F/R. This is clearly a
less stringent criterion as compared to the FWER, since, as the number of discoveries
increases, we allow proportionally more errors. The original FDR procedure of Benjamini and Hochberg [3] assumes that all hypotheses are independent, which is clearly
not the case in general for ERP data. A later work has extended the FDR procedure
to the case of arbitrary dependence structure among variables [4], however, contrary
to what one would expect, the resulting procedure is more conservative, yielding low
power in practice. The FDR procedure and its version for dependent hypotheses have
been the subject of several improvements (e.g., [38, 87]). Recent research on FDRcontrolling procedures attempts to increase their power by sorting the hypotheses
based on a priori information [32]. Such sorting may be guided by previous findings in similar experiments, by the total variance of the variables when using central
location tests, or by any criterion that is independent to the test-statistics. Another
trend in this direction involves arranging the hypotheses in hierarchical trees prior to
testing [102] and in analyzing experimental replicability [40]. The FDR procedures
tend to be unduly conservative when the number of hypotheses is very large, although
much less so than Bonferroni-like procedures. In contrast to FWER-controlling procedures, FDR-controlling procedures are much simpler and faster to compute. They
offer, however, a much looser guarantee against the actual type I error rates and, like
Bonferroni-like procedures, do not take explicitly into consideration the correlation
structure of ERP data.
4.7 Single-Sweep Classification
The goal of a classification method is to automatically estimate the class to which a
single-sweep belongs. The task is challenging because of the very low amplitude of
ERPs as compared to the background EEG. Large artifacts, the non-stationary nature
of EEG and inter-sweep variability exacerbate the difficulty of the task. Although
single-sweep classification has been investigated since a long time [29], it has recently
M. Congedo
Another kind of FWER-controlling permutation test in ERP analysis is the suprathreshold cluster size test [43]. A variant of this test has been implemented in the
EEG toolbox Fieldtrip [71], following the review of Maris and Oostenveld [63]. This
procedure assesses the probability to observe a concentration of the effect simultaneously along one or more studied dimensions. For example, in testing the mean
amplitude difference of a P300 ERP, one expects the effect to be concentrated both
along time, around 300–500 ms, and along space, at midline central and adjacent
parietal locations. This leads to a typical correlation structure of hypothesis in ERP
data; under the null hypothesis the effect would instead be scattered all over both time
and spatial dimensions. An example of the supra-threshold cluster size test applied
in the time-space ERP domain is shown in Fig. 4.8.
Another family of testing procedures controls the false discovery rate (FDR). The
FDR is the expected proportion of falsely rejected hypotheses [3]. Indicating by R
the number of rejected hypotheses and by F the number of those that have been
falsely rejected, the FDR controls the expectation of the ratio F/R. This is clearly a
less stringent criterion as compared to the FWER, since, as the number of discoveries
increases, we allow proportionally more errors. The original FDR procedure of Benjamini and Hochberg [3] assumes that all hypotheses are independent, which is clearly
not the case in general for ERP data. A later work has extended the FDR procedure
to the case of arbitrary dependence structure among variables [4], however, contrary
to what one would expect, the resulting procedure is more conservative, yielding low
power in practice. The FDR procedure and its version for dependent hypotheses have
been the subject of several improvements (e.g., [38, 87]). Recent research on FDRcontrolling procedures attempts to increase their power by sorting the hypotheses
based on a priori information [32]. Such sorting may be guided by previous findings in similar experiments, by the total variance of the variables when using central
location tests, or by any criterion that is independent to the test-statistics. Another
trend in this direction involves arranging the hypotheses in hierarchical trees prior to
testing [102] and in analyzing experimental replicability [40]. The FDR procedures
tend to be unduly conservative when the number of hypotheses is very large, although
much less so than Bonferroni-like procedures. In contrast to FWER-controlling procedures, FDR-controlling procedures are much simpler and faster to compute. They
offer, however, a much looser guarantee against the actual type I error rates and, like
Bonferroni-like procedures, do not take explicitly into consideration the correlation
structure of ERP data.
4.7 Single-Sweep Classification
The goal of a classification method is to automatically estimate the class to which a
single-sweep belongs. The task is challenging because of the very low amplitude of
ERPs as compared to the background EEG. Large artifacts, the non-stationary nature
of EEG and inter-sweep variability exacerbate the difficulty of the task. Although
single-sweep classification has been investigated since a long time [29], it has recently
