3 Modelling Simple Toxicity Endpoints …
41
These expectations were borne out by the data when the methodology (as implemented in Derek Nexus [16]) was tested against several proprietary datasets, for
both the mutagenicity and skin sensitisation endpoints [15, 17]. The general trend
observed was that misclassified and unclassified features occurred relatively infrequently. When misclassified features did occur, they tended to result in a drop in
negative predictivity compared to the entire dataset. This suggests that the identification of chemical fragments that were found elsewhere in known false negatives was
a useful similarity metric and worth highlighting. The presence of unclassified features, however, did not show the same trend in that the negative predictivity tended
to remain high. However, the observed negative predictivity values across several
proprietary mutagenicity datasets displayed a large interquartile range, indicating
the greater variability in the accuracy of predictions for chemicals containing novel
fragments.
Across the proprietary test sets the negative predictivity for chemicals containing either misclassified or unclassified features typically remained higher than the
prevalence of inactive chemicals in the datasets, indicating that their presence was
not a definitive flag for activity. Rather, they should be interpreted as weak arguments
against the negative prediction, indicating greater uncertainty in the predictions. It
is possible that additional expert review of the in silico negative predictions could
resolve some of this additional uncertainty. As such, the identification of misclassified and/or unclassified features serves a secondary purpose, by providing the user
with fragments within the chemical structure to focus on as a starting point for any
further assessment.
The endpoints that have been described so far are driven by a single reactivitybased MIE, reflecting the fact that to make negative predictions, expert systems have
leveraged the power of fragment-based approaches to model chemical reactivity. The
challenges which are expected when making negative predictions for more complex
endpoints (e.g. carcinogenicity or hepatotoxicity) are: the need to use more relevant
descriptors to model non-reactivity-based MIEs and creation of multiple models for
each individual MIE, to ensure that a chemical is not expected to initiate any of the
multiple pathways that could lead to the adverse outcome in question.
3.4 Moving to Quantitative Predictions and Weight
of Evidence Approaches
Another challenge that structural alerts do not address is the need to make quantitative
toxicity predictions. However, they can be used to group chemicals into categories
which react through the same toxicity mechanism, which can then provide a starting
point for making quantitative read across predictions within these categories. In
computational terminology these predictions can be described as k-nearest neighbour
(kNN) models, although in practice this is an example of how SARs and read across
can be used together to make interpretable quantitative toxicity predictions.
41
These expectations were borne out by the data when the methodology (as implemented in Derek Nexus [16]) was tested against several proprietary datasets, for
both the mutagenicity and skin sensitisation endpoints [15, 17]. The general trend
observed was that misclassified and unclassified features occurred relatively infrequently. When misclassified features did occur, they tended to result in a drop in
negative predictivity compared to the entire dataset. This suggests that the identification of chemical fragments that were found elsewhere in known false negatives was
a useful similarity metric and worth highlighting. The presence of unclassified features, however, did not show the same trend in that the negative predictivity tended
to remain high. However, the observed negative predictivity values across several
proprietary mutagenicity datasets displayed a large interquartile range, indicating
the greater variability in the accuracy of predictions for chemicals containing novel
fragments.
Across the proprietary test sets the negative predictivity for chemicals containing either misclassified or unclassified features typically remained higher than the
prevalence of inactive chemicals in the datasets, indicating that their presence was
not a definitive flag for activity. Rather, they should be interpreted as weak arguments
against the negative prediction, indicating greater uncertainty in the predictions. It
is possible that additional expert review of the in silico negative predictions could
resolve some of this additional uncertainty. As such, the identification of misclassified and/or unclassified features serves a secondary purpose, by providing the user
with fragments within the chemical structure to focus on as a starting point for any
further assessment.
The endpoints that have been described so far are driven by a single reactivitybased MIE, reflecting the fact that to make negative predictions, expert systems have
leveraged the power of fragment-based approaches to model chemical reactivity. The
challenges which are expected when making negative predictions for more complex
endpoints (e.g. carcinogenicity or hepatotoxicity) are: the need to use more relevant
descriptors to model non-reactivity-based MIEs and creation of multiple models for
each individual MIE, to ensure that a chemical is not expected to initiate any of the
multiple pathways that could lead to the adverse outcome in question.
3.4 Moving to Quantitative Predictions and Weight
of Evidence Approaches
Another challenge that structural alerts do not address is the need to make quantitative
toxicity predictions. However, they can be used to group chemicals into categories
which react through the same toxicity mechanism, which can then provide a starting
point for making quantitative read across predictions within these categories. In
computational terminology these predictions can be described as k-nearest neighbour
(kNN) models, although in practice this is an example of how SARs and read across
can be used together to make interpretable quantitative toxicity predictions.
