7 Intelligent and Connected Cyber-Physical Systems: A Perspective. . .
381
a wider system as described here (see Fig. 7.10), the input to the neural network
will have typically been processed by a number of elements already [15], such as
image filters and buffering mechanisms. These elements may vary between the
training and operation environments leading to the trained function becoming
dependent on hidden features of the training environment not relevant in the
target system. In addition, typical reliability issues in the target hardware (e.g.,
random hardware failures) may not manifest themselves directly in obviously
erroneous outputs due to the data driven approach, where deviations of individual
parameters or calculations may have subtle but relevant effects on the overall
decision made by the neural network.
7.3.1.4 Sources of Evidence and Structuring the Assurance Case
The development of methods for demonstrating the performance of machine
learning functions to the level of integrity required by safety-critical systems is
currently an emerging field of research. It is expected that, analogous to traditional
algorithmic-based software approaches, a diverse set of complementary evidence
based on constructive measures, formal analysis, and test methods will be required
to make a robust assurance case. In this section, we discuss different categories of
potential evidence that can be used to support such an assurance case.
The choice of training data has a direct impact on accuracy of a machine
learning function. Criteria are therefore required in order to determine whether or
not the training data has the potential to lead to a sufficient level of performance,
including:
• Training data volume: A sufficient amount of training data is used to provide a
statistically relevant distribution of scenarios and to ensure a stabilization of a
strong coverage of weightings in the neural network.
• Coverage of known, critical scenarios: Domain experience based on wellunderstood physical properties of the system and environment as well as previous
validation exercises ensures the identification of classes of scenarios that should
exhibit similar behavior in the function.
• Minimization of unknown, critical scenarios: Some critical attributes of the input
space may not be known during system design [20]. A combination of systematic
identification of equivalence classes in the training data and statistical coverage
during training and validation will therefore be essential to minimize the residual
risk of insufficiencies due to inadequate training data.
Key components of demonstrating the correctness of traditional safety-critical
software are introspective techniques that include manual code review, static
analysis, code coverage and formal verification. These techniques allow for an
argument to be formulated on the detailed algorithmic design and implementation
but cannot be easily transferred to the machine learning paradigms. Other arguments
must therefore be found that make use of knowledge of the internal behavior of the
neural networks.
381
a wider system as described here (see Fig. 7.10), the input to the neural network
will have typically been processed by a number of elements already [15], such as
image filters and buffering mechanisms. These elements may vary between the
training and operation environments leading to the trained function becoming
dependent on hidden features of the training environment not relevant in the
target system. In addition, typical reliability issues in the target hardware (e.g.,
random hardware failures) may not manifest themselves directly in obviously
erroneous outputs due to the data driven approach, where deviations of individual
parameters or calculations may have subtle but relevant effects on the overall
decision made by the neural network.
7.3.1.4 Sources of Evidence and Structuring the Assurance Case
The development of methods for demonstrating the performance of machine
learning functions to the level of integrity required by safety-critical systems is
currently an emerging field of research. It is expected that, analogous to traditional
algorithmic-based software approaches, a diverse set of complementary evidence
based on constructive measures, formal analysis, and test methods will be required
to make a robust assurance case. In this section, we discuss different categories of
potential evidence that can be used to support such an assurance case.
The choice of training data has a direct impact on accuracy of a machine
learning function. Criteria are therefore required in order to determine whether or
not the training data has the potential to lead to a sufficient level of performance,
including:
• Training data volume: A sufficient amount of training data is used to provide a
statistically relevant distribution of scenarios and to ensure a stabilization of a
strong coverage of weightings in the neural network.
• Coverage of known, critical scenarios: Domain experience based on wellunderstood physical properties of the system and environment as well as previous
validation exercises ensures the identification of classes of scenarios that should
exhibit similar behavior in the function.
• Minimization of unknown, critical scenarios: Some critical attributes of the input
space may not be known during system design [20]. A combination of systematic
identification of equivalence classes in the training data and statistical coverage
during training and validation will therefore be essential to minimize the residual
risk of insufficiencies due to inadequate training data.
Key components of demonstrating the correctness of traditional safety-critical
software are introspective techniques that include manual code review, static
analysis, code coverage and formal verification. These techniques allow for an
argument to be formulated on the detailed algorithmic design and implementation
but cannot be easily transferred to the machine learning paradigms. Other arguments
must therefore be found that make use of knowledge of the internal behavior of the
neural networks.
