14 Predictive Modeling of Tox21 Data
287
using either the structure (structure-based models) or assay activity (activity-based
models) SOM clusters or both [17]. To build models using both the structure and
activity SOM clusters, each compound was reassigned to a “consensus cluster” such
that only compounds that belong to the same structure cluster and the same activity
cluster were assigned to the same “consensus cluster.” The consensus clusters were
used to build the structure–activity combined models. For each SOM cluster containing the training compounds, the enrichment of toxic compounds was determined
by a Fisher’s exact test. The −log 10 p-value from the Fisher’s exact test was used
as a measure of the toxic potential (toxicity score) of the compounds in this cluster,
and evaluated as a predictor of toxicity for test compounds that fall into the same
cluster. More significant p-values (larger −log p-values) indicate a larger probability
of toxicity. If a cluster was deficient of toxic compounds, i.e., the fraction of toxic
compounds in the cluster was smaller than the fraction of toxic compounds in the
whole library, the log 10 p-value was used instead. Here we denote the toxicity scores
obtained from the activity SOM as p-activity, those from the structure SOM as pstructure, and those using both the activity and structure SOMs as p-both. To test
model performance, the corresponding SOM cluster or consensus cluster was located
for each test set compound, and p-activity, p-structure, or p-both obtained from the
training set were retrieved. These statistics were compared with the true toxicity
outcome of the test compound to determine if the test compound should be counted
as a true positive (TP: toxic and score > cutoff), false positive (FP: non-toxic and
score > cutoff), true negative (TN: non-toxic and score ≤ cutoff), or false negative
(FN: toxic and score ≤ cutoff).
Models were built for the human adverse drug effects (ADEs) using assay activity
(activity-based models), compound structure (structure-based models), combinations
of structure and activity data with or without drug target annotations (DTAs), and
animal toxicity endpoints [19]. The Weighted Feature Significance (WFS) method
previously developed at NCATS [24] was applied to construct the models. Briefly,
WFS is a two-step scoring algorithm. In the first step, a Fisher’s exact test is used to
determine the significance of enrichment for each feature in the drugs with a certain
ADE compared to the ones without such ADE reported, and a p-value is calculated
for all the features present in the dataset. For assay activity data, each assay readout
was treated as a feature and the feature value was set to 1 for active compounds and
0 for inactive compounds. For animal in vivo toxicity data, each toxicity endpoint
was treated as a feature, and the feature value was set to 1 for toxic compounds and 0
for non-toxic compounds. For structure data, the feature value was set to 1 for drugs
containing that structural feature and 0 for drugs that do not have that feature. For
DTA data, each DTA was treated as a feature, and the feature value was set to 1 for
drugs that reported to have that DTA and 0 for drugs that not known to have the
DTA (see Table 14.2). If a feature is less frequent in the active compound set than
the non-active compound set, then its p-value is set to 1. These p-values form what
we call a “comprehensive” feature fingerprint, which is then used to score each drug
for its potential to cause a certain ADE according to Eq. (1), where p i is the p-value
for feature i; C is the set of all features present in a drug; M is the set of features
encoded in the “comprehensive” feature fingerprint (i.e., features present in at least
287
using either the structure (structure-based models) or assay activity (activity-based
models) SOM clusters or both [17]. To build models using both the structure and
activity SOM clusters, each compound was reassigned to a “consensus cluster” such
that only compounds that belong to the same structure cluster and the same activity
cluster were assigned to the same “consensus cluster.” The consensus clusters were
used to build the structure–activity combined models. For each SOM cluster containing the training compounds, the enrichment of toxic compounds was determined
by a Fisher’s exact test. The −log 10 p-value from the Fisher’s exact test was used
as a measure of the toxic potential (toxicity score) of the compounds in this cluster,
and evaluated as a predictor of toxicity for test compounds that fall into the same
cluster. More significant p-values (larger −log p-values) indicate a larger probability
of toxicity. If a cluster was deficient of toxic compounds, i.e., the fraction of toxic
compounds in the cluster was smaller than the fraction of toxic compounds in the
whole library, the log 10 p-value was used instead. Here we denote the toxicity scores
obtained from the activity SOM as p-activity, those from the structure SOM as pstructure, and those using both the activity and structure SOMs as p-both. To test
model performance, the corresponding SOM cluster or consensus cluster was located
for each test set compound, and p-activity, p-structure, or p-both obtained from the
training set were retrieved. These statistics were compared with the true toxicity
outcome of the test compound to determine if the test compound should be counted
as a true positive (TP: toxic and score > cutoff), false positive (FP: non-toxic and
score > cutoff), true negative (TN: non-toxic and score ≤ cutoff), or false negative
(FN: toxic and score ≤ cutoff).
Models were built for the human adverse drug effects (ADEs) using assay activity
(activity-based models), compound structure (structure-based models), combinations
of structure and activity data with or without drug target annotations (DTAs), and
animal toxicity endpoints [19]. The Weighted Feature Significance (WFS) method
previously developed at NCATS [24] was applied to construct the models. Briefly,
WFS is a two-step scoring algorithm. In the first step, a Fisher’s exact test is used to
determine the significance of enrichment for each feature in the drugs with a certain
ADE compared to the ones without such ADE reported, and a p-value is calculated
for all the features present in the dataset. For assay activity data, each assay readout
was treated as a feature and the feature value was set to 1 for active compounds and
0 for inactive compounds. For animal in vivo toxicity data, each toxicity endpoint
was treated as a feature, and the feature value was set to 1 for toxic compounds and 0
for non-toxic compounds. For structure data, the feature value was set to 1 for drugs
containing that structural feature and 0 for drugs that do not have that feature. For
DTA data, each DTA was treated as a feature, and the feature value was set to 1 for
drugs that reported to have that DTA and 0 for drugs that not known to have the
DTA (see Table 14.2). If a feature is less frequent in the active compound set than
the non-active compound set, then its p-value is set to 1. These p-values form what
we call a “comprehensive” feature fingerprint, which is then used to score each drug
for its potential to cause a certain ADE according to Eq. (1), where p i is the p-value
for feature i; C is the set of all features present in a drug; M is the set of features
encoded in the “comprehensive” feature fingerprint (i.e., features present in at least
