8 It Started with Templates: The Future of Profiling in Side-Channel Analysis
137
augmentation techniques they call Shifting and Add-Remove [122]. They use
convolutional neural networks (CNN) and find data augmentation to significantly
improve the performance of CNN. Pu et al. use a data augmentation technique
where they randomly shift each measurement in order to increase the number
of measurements available in the profiling phase [489]. They report that even
such simple augmentation can effectively improve the performance of profiling
SCA. Picek et al. experiment with several data augmentation and class balancing
techniques in order to decrease the influence of highly unbalanced datasets that
occur when considering HW/HD models [478]. They show that by using a wellknown machine learning technique called SMOTE, it is possible to reduce the
number of measurements needed for a successful attack by up to 10 times. Kim et
al. investigate how the addition of artificial noise to the input signal can be beneficial
to the performance of the neural network [329].
8.2.3 Feature Engineering
When discussing the feature engineering tasks, we can recognize a few directions
that researchers follow in the context of SCA:
• feature selection. Here, the most important subsets of features are selected. We
can distinguish between filter, wrapper, and hybrid techniques.
• dimensionality reduction. The original features are transformed into new features. A common example of such a technique is Principal Component Analysis
(PCA) [25].
When discussing feature engineering, it is important to mention the curse of
dimensionality. This describes the effects of an exponential increase in volume
associated with the increase in the dimensions [71]. As a consequence, as the
dimensionality of the problem increases, the classifier’s performance increases until
the optimal feature subset is reached. Further increasing the dimensionality without
increasing the number of training samples results in a decrease in the classifier
performance.
In the SCA community, there are several standard techniques to conduct feature
selection:
• Pearson Correlation Coefficient. The Pearson correlation coefficient measures the
linear dependence between two variables, x and y, in the range [−1, 1], where
1 is a total positive linear correlation, 0 is no linear correlation, and −1 means
a total negative linear correlation. The Pearson correlation for a sample of the
entire population is defined by [301]:
P earson(x, y) =
N
i=1 ((x i − ¯
x)(y i − ¯
y))
N
i=1 (x i − ¯
x) 2
N
i=1 (y i − ¯
y) 2
,
(8.3)
137
augmentation techniques they call Shifting and Add-Remove [122]. They use
convolutional neural networks (CNN) and find data augmentation to significantly
improve the performance of CNN. Pu et al. use a data augmentation technique
where they randomly shift each measurement in order to increase the number
of measurements available in the profiling phase [489]. They report that even
such simple augmentation can effectively improve the performance of profiling
SCA. Picek et al. experiment with several data augmentation and class balancing
techniques in order to decrease the influence of highly unbalanced datasets that
occur when considering HW/HD models [478]. They show that by using a wellknown machine learning technique called SMOTE, it is possible to reduce the
number of measurements needed for a successful attack by up to 10 times. Kim et
al. investigate how the addition of artificial noise to the input signal can be beneficial
to the performance of the neural network [329].
8.2.3 Feature Engineering
When discussing the feature engineering tasks, we can recognize a few directions
that researchers follow in the context of SCA:
• feature selection. Here, the most important subsets of features are selected. We
can distinguish between filter, wrapper, and hybrid techniques.
• dimensionality reduction. The original features are transformed into new features. A common example of such a technique is Principal Component Analysis
(PCA) [25].
When discussing feature engineering, it is important to mention the curse of
dimensionality. This describes the effects of an exponential increase in volume
associated with the increase in the dimensions [71]. As a consequence, as the
dimensionality of the problem increases, the classifier’s performance increases until
the optimal feature subset is reached. Further increasing the dimensionality without
increasing the number of training samples results in a decrease in the classifier
performance.
In the SCA community, there are several standard techniques to conduct feature
selection:
• Pearson Correlation Coefficient. The Pearson correlation coefficient measures the
linear dependence between two variables, x and y, in the range [−1, 1], where
1 is a total positive linear correlation, 0 is no linear correlation, and −1 means
a total negative linear correlation. The Pearson correlation for a sample of the
entire population is defined by [301]:
P earson(x, y) =
N
i=1 ((x i − ¯
x)(y i − ¯
y))
N
i=1 (x i − ¯
x) 2
N
i=1 (y i − ¯
y) 2
,
(8.3)
