290
R. Huang
Fig. 14.2 Performance
distribution of human
adverse drug effect
prediction models built with
different datasets measured
by AUC-ROC
city endpoints performed moderately (average AUC-ROC = 0.56), similar to those
built with the in vitro assay data (average AUC-ROC = 0.55) for predicting ADEs in
human. This result again confirms that species differences, as well as data sparsity
and lack of consistency, limit the reliability of extrapolating animal in vivo toxicity
data to human in vivo effects.
Similar to the animal toxicity-based models, most models built with in vitro human
cell-based assay data did not show good predictive capacity of human ADEs either.
This low performance may be due to the limited biological space covered by the
current panel of Tox21 assays. Since many drugs in the 10K collection have target
and/or mechanism of action annotations available, we collected drug target annotations (DTAs) from the literature (2370 DTAs) and combined them with in vitro assay
data to build new models. These combined models showed remarkable improvements in predictive performance with average AUC-ROC for human ADE prediction increased from 0.55 to 0.67 (Fig. 14.2) [19]. In addition, we identified a small
subset of 58 DTAs that contributed the most to the prediction. Adding this set of 58
DTAs to in vitro assay data significantly improved the model performance, increasing the average AUC-ROC to 0.63 for human ADE prediction (Fig. 14.2) [19]. This
result shows that data on just a small set of additional DTAs (2% of the entire 2370
DTA set) can expand the biological space coverage sufficiently to produce predictive
models of human toxicity when combined with in vitro assay data. While the entire
DTA set improved the model performance by 22–28% on average, the selected set of
58 DTAs alone improved the model performance, on average, by 15–18%. In other
words, 2% of the DTA information could account for ~70% of the improvement
in the predictive capacity of the models. Most of the 58 selected targets/pathways
are GPCR targets, which are important drug targets non-specific activity on which
R. Huang
Fig. 14.2 Performance
distribution of human
adverse drug effect
prediction models built with
different datasets measured
by AUC-ROC
city endpoints performed moderately (average AUC-ROC = 0.56), similar to those
built with the in vitro assay data (average AUC-ROC = 0.55) for predicting ADEs in
human. This result again confirms that species differences, as well as data sparsity
and lack of consistency, limit the reliability of extrapolating animal in vivo toxicity
data to human in vivo effects.
Similar to the animal toxicity-based models, most models built with in vitro human
cell-based assay data did not show good predictive capacity of human ADEs either.
This low performance may be due to the limited biological space covered by the
current panel of Tox21 assays. Since many drugs in the 10K collection have target
and/or mechanism of action annotations available, we collected drug target annotations (DTAs) from the literature (2370 DTAs) and combined them with in vitro assay
data to build new models. These combined models showed remarkable improvements in predictive performance with average AUC-ROC for human ADE prediction increased from 0.55 to 0.67 (Fig. 14.2) [19]. In addition, we identified a small
subset of 58 DTAs that contributed the most to the prediction. Adding this set of 58
DTAs to in vitro assay data significantly improved the model performance, increasing the average AUC-ROC to 0.63 for human ADE prediction (Fig. 14.2) [19]. This
result shows that data on just a small set of additional DTAs (2% of the entire 2370
DTA set) can expand the biological space coverage sufficiently to produce predictive
models of human toxicity when combined with in vitro assay data. While the entire
DTA set improved the model performance by 22–28% on average, the selected set of
58 DTAs alone improved the model performance, on average, by 15–18%. In other
words, 2% of the DTA information could account for ~70% of the improvement
in the predictive capacity of the models. Most of the 58 selected targets/pathways
are GPCR targets, which are important drug targets non-specific activity on which
