40
R. Williams et al.
Fig. 3.1 Outline of the negative prediction methodology employed by an expert knowledge-based
system
that involved comparing the fragments present in a non-alerting chemical to a large
dataset of other chemicals with known activities [15].
The first method was based on the expectation that chemicals which have purposely been excluded from a structural alert based on a deactivating feature were
more likely to be inactive than those that simply contain no substructures that match
an alert, as more is known about the former than the latter. In practice, this method
was a poor indicator of the reliability of a negative prediction. The border between
active and inactive compounds was rarely unambiguous due to limitations in both
the quality and quantity of the available data. In contrast, the second method was
successful in highlighting cases where the negative prediction could be treated with
either a higher or lower degree of confidence. This was achieved by answering two
questions of a non-alerting chemical that an expert user might also ask (Fig. 3.1):
1. Is the chemical similar to known active chemicals that the expert system predicts
incorrectly (i.e. false negatives)?
2. Does the chemical contain any fragments that the expert system has not seen
before?
Both questions are answered by comparing the non-alerting chemical to a large,
curated dataset of chemicals with known activity data collected from the public
domain. When a chemical contains a fragment that is found exclusively in false
negative compounds in the dataset, this is flagged to the user as a misclassified
feature. Likewise, where a chemical contains a fragment that is not present at all in
the dataset, it is highlighted as an unclassified feature. The presence of either type of
feature is likely to reduce the confidence a user has in the negative in silico outcome.
Misclassified features are expected to decrease the accuracy of the prediction as the
non-alerting chemical is similar to other active compounds that the expert system
predicts poorly. Unclassified features are expected to increase the uncertainty around
the prediction as the non-alerting chemical resides in an unstudied area of chemical
space.
R. Williams et al.
Fig. 3.1 Outline of the negative prediction methodology employed by an expert knowledge-based
system
that involved comparing the fragments present in a non-alerting chemical to a large
dataset of other chemicals with known activities [15].
The first method was based on the expectation that chemicals which have purposely been excluded from a structural alert based on a deactivating feature were
more likely to be inactive than those that simply contain no substructures that match
an alert, as more is known about the former than the latter. In practice, this method
was a poor indicator of the reliability of a negative prediction. The border between
active and inactive compounds was rarely unambiguous due to limitations in both
the quality and quantity of the available data. In contrast, the second method was
successful in highlighting cases where the negative prediction could be treated with
either a higher or lower degree of confidence. This was achieved by answering two
questions of a non-alerting chemical that an expert user might also ask (Fig. 3.1):
1. Is the chemical similar to known active chemicals that the expert system predicts
incorrectly (i.e. false negatives)?
2. Does the chemical contain any fragments that the expert system has not seen
before?
Both questions are answered by comparing the non-alerting chemical to a large,
curated dataset of chemicals with known activity data collected from the public
domain. When a chemical contains a fragment that is found exclusively in false
negative compounds in the dataset, this is flagged to the user as a misclassified
feature. Likewise, where a chemical contains a fragment that is not present at all in
the dataset, it is highlighted as an unclassified feature. The presence of either type of
feature is likely to reduce the confidence a user has in the negative in silico outcome.
Misclassified features are expected to decrease the accuracy of the prediction as the
non-alerting chemical is similar to other active compounds that the expert system
predicts poorly. Unclassified features are expected to increase the uncertainty around
the prediction as the non-alerting chemical resides in an unstudied area of chemical
space.
