66
S. A. Khan et al.
of common compounds tested in CMap cell lines were also profiled by the NCI60
program. This presents the unique opportunity to study toxic effects by integrating
these two large-scale data sets (see Sect. 4.3.3). For each CMap drug-cell pair which
was also screened by NCI60, a dose-dependent toxicity score was computed such that
positive values indicated that the CMap instance is profiled at a drug concentration
higher than GI50, TGI, or LC50, and therefore suggest a dose-dependent cytotoxic
response.
4.3.2 Multi-view Toxicogenomic Using Group Factor
Analysis
To study the gene–toxicity relationships, we performed an integrated modeling of
the two data sets, CMap and NCI60. The CMap data comprised detailed gene level
differential expression profiles that represent the molecular response space across
11,350 genes, measured after 222 drug treatments across 3 cell lines. These data
were preprocessed as previously described [47]. To focus the analysis, the 1106
highest variance genes were selected to form an expression matrix consisting of 222
drug-cell samples ×1106 genes. The toxicity values described in Sect. 4.3.1 were
used to represent profiles of 222 drug-cell samples ×3 toxicity measures.
Group factor analysis (GFA) is designed to model the relationships between multiple data sets. Here, GFA was used to identify the toxicogenomic dependencies
between drug-induced gene expression changes and toxicity scores. These dependencies, once identified, can elicit insights into molecular mechanisms of toxicity.
GFA was run with large enough components as specified by Virtanen et al. [28],
identifying 8 shared components that capture cross-expression and toxicity relationships, as shown in Fig. 4.4, whereas a number of components found were specific to
one of the data sets only. The shared components model the dependencies between
the data sets while those specific to gene expression capture patterns that are not
correlated with toxicity and vice versa. The components 1 through 8 had varying
numbers of genes attached to them: 518, 748, 39, 90, 27, 45, 16, and 20. The first
two components included an excess of up-regulated genes (component 1:316) and
down-regulated (component 2:706) genes, respectively.
Functional analysis of the eight components was performed with Ingenuity Pathway Analysis (IPA) which indicated that the first two components captured the largest
number of biological mechanisms. The first component is highlighted here, as upregulated genes are most informative for biomarker analysis applications (Fig. 4.5).
Component 1 enriched for many organ toxicity-related gene lists, including hepatic
cholestasis and liver necrosis as well as functional pathways related to oxidative
stress, the p. 53 pathway activation and Nf-kappa B signaling and Toll-like receptor (TLR) activation. RELA, the NF-kappa B regulator, was predicted to be most
clearly effected (p < 10–16 and Z-score 3.5). Others included the TP53 (p < 10–10,
Z > 1), TLR-related ECSIT (p < 10–15, Z > 3.5), and NR3C1 (p < 10–14, Z < −05),
Précédent

- 80/416

Suivant