268
M. Chen et al.
marone/amiodarone, which are defined by their molecular structure (tanimoto similarity > 0.5) and similar mode of action but discordant toxicity [65].
13.3.3 Conventional QSAR [13]
QSAR models have been extensively applied to predict drug-induced liver injury due
to their ability to produce rapid results without requiring physical drug substance [22,
24, 66, 67]. So far, most of the QSAR-DILI models’ report limited predictive performance, with accuracies of approximately 60% or less, especially when the models
are challenged by external validation sets. We implemented an improved strategy to
develop the QSAR model for predicting DILI in humans using a robust annotation of
DILI risk relying on FDA-approved drug labeling and applying an extensive modeling validation strategy to ensure the model performance was sustainable and better
than by chance.
Our conventional QSAR was developed by using a decision forest (DF) algorithm
to correlate the chemical structures with their DILI risk in humans based on a set of
drugs as the training set. The DF algorithm is a supervised machine learning technique utilizing a modified decision tree model by employing a consensus technique
to combine multiple heterogeneous decision trees to achieve a more accurate predictive model. The DF algorithm is developed by our laboratory, and the software
is publicly available @ https://www.fda.gov/ScienceResearch/BioinformaticsTools/
DecisionForest/default.htm. Meanwhile, the chemical structures of drugs were codified into a digital format (i.e., chemical descriptors) as the input for the machine learning algorithm DF. Here, we utilized the Mold2 molecular descriptors to transform
the 2-dimensional chemical structures into 777 chemical descriptors. Mold2 is also
developed by NCTR and freely available at https://www.fda.gov/ScienceResearch/
BioinformaticsTools/Mold2/default.htm.
The training set to develop the QSAR model included 197 drugs (NCTR training
set), which were annotated by FDA-approved drug labeling as discussed previously.
The drug label-based DILI annotation proved to be robust and consistent as compared
to other annotations [37], which is critical for the development of an improved QSAR
model. The developed models were evaluated by internal and external validations.
Internal validation employed a 2000 run of 10-fold cross-validation based on the
NCTR training set. External validation of the QSAR models was applied to 3 different
datasets with a total of 438 unique drugs: NCTR validation dataset with N = 190
drugs, Greene et al. dataset with N = 328 drugs, and Xu et al. dataset with N = 241
drugs. The validation results in Table 13.2 show that when using the NCTR annotated
training or validation set, the predictive performance of the QSAR model had an
accuracy of 69.7% for internal cross-validation and 68.9% for external validation.
Meanwhile, the external validation assessed by Greene and Xu et al. datasets was at
accuracies of 61.6 and 63.1%, respectively. The performances evaluated by different
datasets are largely consistent, the occasional variations might reflect the quality of
annotation, and the diverse drugs included in the datasets.
Précédent

- 277/416

Suivant