11 Applicability Domain: Towards a More Formal …
225
reliability of the prediction and given the use case, the user can set a desired threshold
below which the prediction is deemed unreliable. For instance, in the context of
virtual screening, a low reliability level might be acceptable, whereas in the case
of human safety assessment, the end user might accept only predictions with high
reliability. Reliability is a prediction attribute as opposed to applicability which is a
model property.
Reliability is not dependent on the result of the prediction. To better understand
this decoupling, let us compare the model to a group of people asked a question in
a specific domain and depositing their answer in a shared sealed envelope. If the
people are picked randomly in a public area, before we even look in the envelope, we
would consider the contents of the envelope to be less reliable than if we had chosen
a group of experts in the relevant domain. We can see the outcome of the algorithm
as a closed envelope, where the reliability captures how much we trust the contents
of the envelope prior to opening it.
The actual desired level of reliability is use-case-dependent and can be set by the
end user. The reliability metric must be calibrated and normalised. Methods based
on information density can be used for this purpose. If a prediction’s reliability falls
below the required threshold, the prediction is said to be “Outside the Reliability
Domain”. Whether the result in the envelope points strongly or weakly towards a
given class (classification model) or value (regression model) is not captured by the
reliability; it is expressed through the likelihood assigned to these outcomes.
11.2.3 Decidability
After checking that we can use the model to make a valid prediction and that this prediction is reliable enough for the intended use case, we can finally consider the actual
outcome of the model. Reusing the previous envelope analogy, we can now open this
envelope to see the actual answers and check the consistency across them. Agreement
amongst the supporting evidence is the main driver towards a clear conclusion of the
model. If the supporting data for a given prediction converge, the model can build a
more decisive prediction, whereas if the supporting information is self-contradicting,
the prediction will be equivocal and therefore less decisive (Fig. 11.9).
This result is usually expressed by the model in the form of a probability distribution across the different classes or values, or in the form of discrete likelihood levels
as in the case of expert systems. For instance:
• Naïve Bayes classifiers directly assign a posterior probability for each class.
• k-nearest neighbours (kNN) models can express a distribution of likelihood based
on the distribution of the labels of the k-nearest neighbours and their distance to
the query within the descriptor space.
• Similarly, a likelihood distribution can be derived from Random Forest predictions
based on the relative vote count (at the individual tree level) for each possible class
or value.
225
reliability of the prediction and given the use case, the user can set a desired threshold
below which the prediction is deemed unreliable. For instance, in the context of
virtual screening, a low reliability level might be acceptable, whereas in the case
of human safety assessment, the end user might accept only predictions with high
reliability. Reliability is a prediction attribute as opposed to applicability which is a
model property.
Reliability is not dependent on the result of the prediction. To better understand
this decoupling, let us compare the model to a group of people asked a question in
a specific domain and depositing their answer in a shared sealed envelope. If the
people are picked randomly in a public area, before we even look in the envelope, we
would consider the contents of the envelope to be less reliable than if we had chosen
a group of experts in the relevant domain. We can see the outcome of the algorithm
as a closed envelope, where the reliability captures how much we trust the contents
of the envelope prior to opening it.
The actual desired level of reliability is use-case-dependent and can be set by the
end user. The reliability metric must be calibrated and normalised. Methods based
on information density can be used for this purpose. If a prediction’s reliability falls
below the required threshold, the prediction is said to be “Outside the Reliability
Domain”. Whether the result in the envelope points strongly or weakly towards a
given class (classification model) or value (regression model) is not captured by the
reliability; it is expressed through the likelihood assigned to these outcomes.
11.2.3 Decidability
After checking that we can use the model to make a valid prediction and that this prediction is reliable enough for the intended use case, we can finally consider the actual
outcome of the model. Reusing the previous envelope analogy, we can now open this
envelope to see the actual answers and check the consistency across them. Agreement
amongst the supporting evidence is the main driver towards a clear conclusion of the
model. If the supporting data for a given prediction converge, the model can build a
more decisive prediction, whereas if the supporting information is self-contradicting,
the prediction will be equivocal and therefore less decisive (Fig. 11.9).
This result is usually expressed by the model in the form of a probability distribution across the different classes or values, or in the form of discrete likelihood levels
as in the case of expert systems. For instance:
• Naïve Bayes classifiers directly assign a posterior probability for each class.
• k-nearest neighbours (kNN) models can express a distribution of likelihood based
on the distribution of the labels of the k-nearest neighbours and their distance to
the query within the descriptor space.
• Similarly, a likelihood distribution can be derived from Random Forest predictions
based on the relative vote count (at the individual tree level) for each possible class
or value.
