11 Applicability Domain: Towards a More Formal …
223
the query instance is labelled as “Outside the Applicability Domain”. The AD of a
model is a necessary but not sufficient condition for being confident in a prediction;
it is also important to consider the reliability and the assertiveness of the prediction.
It is worth noting that at this stage we have not yet used the model to perform a
prediction.
11.2.2 Reliability
Once established that the model is applicable, it becomes legitimate and meaningful
to perform an actual prediction; we can call such a prediction a “valid” prediction.
Whether the prediction is reliable or not will depend on the quantity, quality and
relevance of the information available to the model. This supporting information can
be in the form of training examples for statistical models, or a knowledge base in the
case of expert systems. Typically, we would expect that a model based on training
data containing compounds similar to the query compound, will produce a more
reliable prediction than a model for which the query structure is an outlier. Indeed, it
is fair to assume that the quantity and quality of information available to the model
for a given query structure is provided by its nearest neighbours in the descriptor
space. The number of close neighbours and the average distance of these neighbours
capture the information density in the SAR region of the predicted structure. Since
some data points have a degree of redundancy, it is useful to also take into account
their dispersion; for instance, data points close to each other provide similar evidence
and their combination holds less information than two better dispersed data points
(more efficient domain coverage). Finally, the reliability of the data points themselves
may impact the quality of the information; typically, experimental data obtained
following good laboratory practice are likely to bear more information and less
noise. All these elements contribute to the level of evidence supporting a reliable
prediction (Fig. 11.7).
Different AD methods are based on this hypothesis including:
• Distance to model: These methods attempt to measure the distance in terms of
dissimilarity between the query compound and the training dataset. There are
many ways to define this distance; for instance, we could consider only the distance between the predicted structure and the closest structure in the training set
(Fig. 11.8a). Alternatively, it is possible to consider the distance between the query
structure and the centre of the training data using a virtual centroid point. Different
similarity metrics (Tanimoto, Euclidean, Manathan, etc.) can be used to compute
the actual value of the distance [25].
• Information density: The distance to the model is a coarse expression of the information available to the model, it does not consider internal data density variation
since the training data is treated as a whole. A more fine-grained approach consists
of measuring the average distance to the k closest compounds [26]. This approach
can be extended into a continuous estimation of the information density using a
Précédent

- 233/416

Suivant