8 It Started with Templates: The Future of Profiling in Side-Channel Analysis
143
discuss those. If not all data from datasets are used, it is necessary to state how the
samples are chosen and how many are used in the experiments. One needs to define
the level of noise appearing in the data in a clearly reproducible way, e.g., using the
signal-to-noise ratio (SNR). Finally, if some feature engineering procedure is used,
it needs to be clearly stated in order to know what features are used in the end.
Algorithms
When discussing the choice of algorithms, first it is necessary either to specify which
framework and algorithms are used or provide pseudo-code (for example, when
custom algorithms are used). As a rule of thumb, more than one algorithm should
always be used: the algorithms should ideally belong to different machine learning
approaches (e.g., a decision tree method like Random Forest and a kernel method
like Support Vector Machine (SVM)). Next, all parameters that uniquely define the
algorithm need to be enumerated.
Experiments
Regarding the experiments, it is first necessary to discuss how the data are divided
into training and testing sets. Then, for the training phase, one needs to define the
test options (e.g., whether to use the whole dataset or cross-validation, etc.) After
that, for each algorithm, one needs to define a set of parameter values to conduct the
tuning phase. There are different options for tuning, but we consider starting with
the default parameters as a reasonable approach and continue varying them until
there is no more improvement. Naturally, this should be done in a reasonable way,
since the tuning phase is the most expensive from the computational perspective and
it is usually not practical to test all combinations of parameters.
Results
For the tuning phase, it is usually sufficient to report the accuracy. For the testing
results, one should report the accuracy but also some other metric like the area under
the ROC curve (AUC) or the F-measure. The area under the ROC curve is used
to measure the accuracy and is calculated via Mann-Whitney statistics [580]; the
ROC curve is the ratio between the true positive rate and the false positive rate.
An AUC close to 1 represents a good test, while a value close to 0.5 represents a
random guess. The F-measure is the harmonic mean of the precision and recall,
where precision is the ratio between true positive (TP, the number of examples
predicted positive that are actually positive) and predicted positive. The recall is
the ratio between true positives and actual positives [488]. Both the F-Measure and
the AUC can help in situations where accuracy can be misleading, i.e., where we
are also interested in the number of false positive and false negative values.
143
discuss those. If not all data from datasets are used, it is necessary to state how the
samples are chosen and how many are used in the experiments. One needs to define
the level of noise appearing in the data in a clearly reproducible way, e.g., using the
signal-to-noise ratio (SNR). Finally, if some feature engineering procedure is used,
it needs to be clearly stated in order to know what features are used in the end.
Algorithms
When discussing the choice of algorithms, first it is necessary either to specify which
framework and algorithms are used or provide pseudo-code (for example, when
custom algorithms are used). As a rule of thumb, more than one algorithm should
always be used: the algorithms should ideally belong to different machine learning
approaches (e.g., a decision tree method like Random Forest and a kernel method
like Support Vector Machine (SVM)). Next, all parameters that uniquely define the
algorithm need to be enumerated.
Experiments
Regarding the experiments, it is first necessary to discuss how the data are divided
into training and testing sets. Then, for the training phase, one needs to define the
test options (e.g., whether to use the whole dataset or cross-validation, etc.) After
that, for each algorithm, one needs to define a set of parameter values to conduct the
tuning phase. There are different options for tuning, but we consider starting with
the default parameters as a reasonable approach and continue varying them until
there is no more improvement. Naturally, this should be done in a reasonable way,
since the tuning phase is the most expensive from the computational perspective and
it is usually not practical to test all combinations of parameters.
Results
For the tuning phase, it is usually sufficient to report the accuracy. For the testing
results, one should report the accuracy but also some other metric like the area under
the ROC curve (AUC) or the F-measure. The area under the ROC curve is used
to measure the accuracy and is calculated via Mann-Whitney statistics [580]; the
ROC curve is the ratio between the true positive rate and the false positive rate.
An AUC close to 1 represents a good test, while a value close to 0.5 represents a
random guess. The F-measure is the harmonic mean of the precision and recall,
where precision is the ratio between true positive (TP, the number of examples
predicted positive that are actually positive) and predicted positive. The recall is
the ratio between true positives and actual positives [488]. Both the F-Measure and
the AUC can help in situations where accuracy can be misleading, i.e., where we
are also interested in the number of false positive and false negative values.
