380
W. Chang et al.
rarely or may be so dangerous that they are not well represented in the training
data. Consider, for example, the situation where a small child enters the road
ahead between two parked vehicles. This leads to the effect that critical situations
remain undertrained in the final function (scalable oversight). In addition, the
system should continue to perform accurately even if the operational environment
differs from the training environment (distributional shift) [14]. This effectively
can be formulated as the robustness of the system to react in a shift of distribution
between its training and operational environment. Distributional shift will be
inevitable in most open context systems, as the environment constantly changes
and can adapt to the behavior of actors within the system. For example, car
drivers will adjust their behavior within an environment in which autonomous
vehicles are present, vehicle and pedestrian appearances change over time, etc.
• Robustness of the trained function. Machine learning techniques are typically
chosen for their ability to approximate target functions based on a finite set of
training data. This has advantages over procedural techniques where the function
to be implemented may be too complex to specify or implement algorithmically
due to an open context environment or due to the unstructured nature of the
input data. In other words, when presented with new data, the function will
predict a correct answer based on already observed input/output pairs. An often
cited problem associated with neural networks, is the possibility of adversarial
perturbations [16, 17, 18]. An adversarial perturbation is an input sample that is
similar (at least to the human eye) to other samples but that leads to a completely
different categorization with a high confidence value. It has been shown that
such examples can be automatically generated and used to “trick” the network.
Although it is still unclear to what extent adversarial perturbations could occur
naturally or whether they would be exploited for malicious purposes, from a
safety validation perspective, they are useful for demonstrating that features
can be learnt by the network and assigned an incorrect relevance. Therefore,
methods are required to minimize the probability of such behavior especially
in critical driving situations. One of the factors that is often attributed to this
class of problems is that the set of possible functions is exponentially larger than
those that can be represented through machine learning techniques. Therefore,
the likelihood that a machine learning technique would select an appropriate
approximation appears at first glance very unlikely. The authors in [19] argue,
however, that deep learning is nevertheless effective because the function to be
approximated is rooted within the physical universe and physics favors certain
classes of exceptionally simple probability distributions that deep learning is
uniquely suited to model. The challenge, therefore, is how to ensure that the
machine learning algorithms focus on those physical properties of the inputs
relevant to the target function without becoming distracted by irrelevant features.
In other words, act within the same hierarchical dimensions as the target function
[19].
• Differences between the training and execution platforms. As discussed above,
machine learning functions can be sensitive to subtle changes in the input data.
When using machine learning to represent a function that is embedded as part of
W. Chang et al.
rarely or may be so dangerous that they are not well represented in the training
data. Consider, for example, the situation where a small child enters the road
ahead between two parked vehicles. This leads to the effect that critical situations
remain undertrained in the final function (scalable oversight). In addition, the
system should continue to perform accurately even if the operational environment
differs from the training environment (distributional shift) [14]. This effectively
can be formulated as the robustness of the system to react in a shift of distribution
between its training and operational environment. Distributional shift will be
inevitable in most open context systems, as the environment constantly changes
and can adapt to the behavior of actors within the system. For example, car
drivers will adjust their behavior within an environment in which autonomous
vehicles are present, vehicle and pedestrian appearances change over time, etc.
• Robustness of the trained function. Machine learning techniques are typically
chosen for their ability to approximate target functions based on a finite set of
training data. This has advantages over procedural techniques where the function
to be implemented may be too complex to specify or implement algorithmically
due to an open context environment or due to the unstructured nature of the
input data. In other words, when presented with new data, the function will
predict a correct answer based on already observed input/output pairs. An often
cited problem associated with neural networks, is the possibility of adversarial
perturbations [16, 17, 18]. An adversarial perturbation is an input sample that is
similar (at least to the human eye) to other samples but that leads to a completely
different categorization with a high confidence value. It has been shown that
such examples can be automatically generated and used to “trick” the network.
Although it is still unclear to what extent adversarial perturbations could occur
naturally or whether they would be exploited for malicious purposes, from a
safety validation perspective, they are useful for demonstrating that features
can be learnt by the network and assigned an incorrect relevance. Therefore,
methods are required to minimize the probability of such behavior especially
in critical driving situations. One of the factors that is often attributed to this
class of problems is that the set of possible functions is exponentially larger than
those that can be represented through machine learning techniques. Therefore,
the likelihood that a machine learning technique would select an appropriate
approximation appears at first glance very unlikely. The authors in [19] argue,
however, that deep learning is nevertheless effective because the function to be
approximated is rooted within the physical universe and physics favors certain
classes of exceptionally simple probability distributions that deep learning is
uniquely suited to model. The challenge, therefore, is how to ensure that the
machine learning algorithms focus on those physical properties of the inputs
relevant to the target function without becoming distracted by irrelevant features.
In other words, act within the same hierarchical dimensions as the target function
[19].
• Differences between the training and execution platforms. As discussed above,
machine learning functions can be sensitive to subtle changes in the input data.
When using machine learning to represent a function that is embedded as part of
