382
W. Chang et al.
• Saliency maps: Based on the back propagation of results in the neural network,
saliency maps [21] highlight those portions of an image that have greatest
influence on classification results and can be used to provide a manual plausibility
check of results as well as to determine potential causes of failed tests.
• Explanations: Another line of research tries to generate natural language explanations referring in human understandable terms to the discriminating contents
of an input image to explain which features were relevant for the classification
[22].
Due to the inherent restrictions of the applicability of white-box approaches to
the verification of the trained function, a strong emphasis will remain on testing
as a means to estimate the achieved performance of the trained function. Standard
approaches to testing machine learning functions involve reserving a proportion of
the data collected for training purposes to performing validation tests. These tests
naturally suffer from the same inadequacies as described above for the training data.
Several additional test approaches are therefore being developed.
• Synthetic data generation and search-based testing: Based on advances in
computer graphics realism as well as the possibility to generate data with specific
properties, the use of synthetically generated data may also play a role [23]
in the assurance case. Synthetic data can be used to generate huge numbers
of test cases, in particular to cover critical or rare situations, otherwise not
adequately represented in naturally occurring data. The use of synthetic data also
allows test cases to be automatically generated together with the corresponding
ground truth. This allows for search-based optimization approaches to be applied
to automatically generate (physically feasible) images which produce incorrect
classifications. However, the use of synthetic data also implies the introduction
of the additional assumption in the assurance case that the synthetic data would
lead to test results that are indeed representative of the operational environment.
• White-box coverage tests: At present, there is no clear consensus on which
stopping criteria to apply when testing machine learning functions. Due to the
fact that deep neural networks operate in a highly dimensional feature space,
choosing test cases based on a set of domain-specific equivalence classes is less
likely to be effective, as there is a high chance that these do not match the
feature dimensions learnt by the neural network. White-box criteria have been
proposed based on the concept of neuron coverage to determine the completeness
and effectiveness of the test data. This involves calculating the ratio of activated
neurons (activation values above a given threshold) to the total number of neurons
for a given set of input data [24, 25]. These approaches have also been combined
with search-based testing techniques to create variations of test data that achieve
coverage. These techniques are only applicable in combination with functional
criteria, and it is as yet unclear how effective such white-box techniques are at
discovering performance issues in the neural networks.
The objective of applying techniques such as those described above should not
only be to demonstrate that a given performance requirement has been met but
W. Chang et al.
• Saliency maps: Based on the back propagation of results in the neural network,
saliency maps [21] highlight those portions of an image that have greatest
influence on classification results and can be used to provide a manual plausibility
check of results as well as to determine potential causes of failed tests.
• Explanations: Another line of research tries to generate natural language explanations referring in human understandable terms to the discriminating contents
of an input image to explain which features were relevant for the classification
[22].
Due to the inherent restrictions of the applicability of white-box approaches to
the verification of the trained function, a strong emphasis will remain on testing
as a means to estimate the achieved performance of the trained function. Standard
approaches to testing machine learning functions involve reserving a proportion of
the data collected for training purposes to performing validation tests. These tests
naturally suffer from the same inadequacies as described above for the training data.
Several additional test approaches are therefore being developed.
• Synthetic data generation and search-based testing: Based on advances in
computer graphics realism as well as the possibility to generate data with specific
properties, the use of synthetically generated data may also play a role [23]
in the assurance case. Synthetic data can be used to generate huge numbers
of test cases, in particular to cover critical or rare situations, otherwise not
adequately represented in naturally occurring data. The use of synthetic data also
allows test cases to be automatically generated together with the corresponding
ground truth. This allows for search-based optimization approaches to be applied
to automatically generate (physically feasible) images which produce incorrect
classifications. However, the use of synthetic data also implies the introduction
of the additional assumption in the assurance case that the synthetic data would
lead to test results that are indeed representative of the operational environment.
• White-box coverage tests: At present, there is no clear consensus on which
stopping criteria to apply when testing machine learning functions. Due to the
fact that deep neural networks operate in a highly dimensional feature space,
choosing test cases based on a set of domain-specific equivalence classes is less
likely to be effective, as there is a high chance that these do not match the
feature dimensions learnt by the neural network. White-box criteria have been
proposed based on the concept of neuron coverage to determine the completeness
and effectiveness of the test data. This involves calculating the ratio of activated
neurons (activation values above a given threshold) to the total number of neurons
for a given set of input data [24, 25]. These approaches have also been combined
with search-based testing techniques to create variations of test data that achieve
coverage. These techniques are only applicable in combination with functional
criteria, and it is as yet unclear how effective such white-box techniques are at
discovering performance issues in the neural networks.
The objective of applying techniques such as those described above should not
only be to demonstrate that a given performance requirement has been met but
