4 Matrix and Tensor Factorization Methods …
59
In vitro toxicological outcomes are often based on large-scale compound response
profiles, which summarize the responses in a particular cell context. For instance,
NCI-60 developmental therapeutics program uses several metrics to quantify doseresponses to a library of thousands of compounds across a panel of 59 human tumor
cell lines; such summary metrics include: GI50 (50% Growth Inhibition), TGI (Total
Growth Inhibition), and LC50 (50% Lethal Concentration) (https://dtp.cancer.gov/
discovery_development/nci-60/). In such a high-throughput setting, the computational task is to search for patterns of toxicity outcomes in correlation with genomic
and molecular profiles of the same panel of cell lines. However, cytotoxicity is not
a biologically uniform response. Cells use multiple mechanisms that depend on the
chemical or drug and the dose at which it is applied to respond to and counter the
effects of stressors. Transcriptomic profiling and subsequent analyses using component modeling approaches discussed herein can segment these responses into biologically intelligible and explainable sub-responses, while also providing predictive
models.
Therefore, advances in machine learning methodology allow study of toxicogenomic relationships in a more systematic fashion and reveal valuable drug-gene
associations. For instance, community efforts have shown great promise to improve
in silico predictions of drug sensitivity [9]. In another effort [10] carried out a personalized quantitative structure–activity relationship QSAR analysis by integrating gene
expression, drug structures, and drug response profiles using a non-linear machine
learning approach. Their study demonstrated the possibility to predict the drug sensitivity outcome for untested drugs even in new cell types. drug-pathway associations
can be identified using advanced machine learning methodologies that model the
complex molecular interactions [11]. Recently, integrative multitask sparse regression methods have been used to systematically identify biomarker combinations for
predicting drug outcomes [12]. Increasing evidence from recent studies thus poses
the hypothesis that common patterns in the activity profiles of genes and sensitivity/toxicity profiles of drugs can identify cellular response mechanisms and could be
used even in predicting the tissue type or cell context-specific toxicity outcome of
drug treatment.
This chapter is organized as follows: Sect. 4.2 introduces representative classes of
recent machine learning methods, with a specific emphasis on the matrix and tensor
factorization methods. Section 4.3 demonstrates the application of these methods to
identification of toxicogenomic relationships in example case studies, followed by
a discussion in Sect. 4.4. Section 4.5 concludes the chapter with current limitations
and future directions in these developments.
4.2 Machine Learning Methods
Machine learning algorithms search for patterns in data to extract useful information
[13, 14]. These algorithms learn a representation a.k.a. the model from existing data
samples and then utilize the model in different tasks. When applied to experimental
59
In vitro toxicological outcomes are often based on large-scale compound response
profiles, which summarize the responses in a particular cell context. For instance,
NCI-60 developmental therapeutics program uses several metrics to quantify doseresponses to a library of thousands of compounds across a panel of 59 human tumor
cell lines; such summary metrics include: GI50 (50% Growth Inhibition), TGI (Total
Growth Inhibition), and LC50 (50% Lethal Concentration) (https://dtp.cancer.gov/
discovery_development/nci-60/). In such a high-throughput setting, the computational task is to search for patterns of toxicity outcomes in correlation with genomic
and molecular profiles of the same panel of cell lines. However, cytotoxicity is not
a biologically uniform response. Cells use multiple mechanisms that depend on the
chemical or drug and the dose at which it is applied to respond to and counter the
effects of stressors. Transcriptomic profiling and subsequent analyses using component modeling approaches discussed herein can segment these responses into biologically intelligible and explainable sub-responses, while also providing predictive
models.
Therefore, advances in machine learning methodology allow study of toxicogenomic relationships in a more systematic fashion and reveal valuable drug-gene
associations. For instance, community efforts have shown great promise to improve
in silico predictions of drug sensitivity [9]. In another effort [10] carried out a personalized quantitative structure–activity relationship QSAR analysis by integrating gene
expression, drug structures, and drug response profiles using a non-linear machine
learning approach. Their study demonstrated the possibility to predict the drug sensitivity outcome for untested drugs even in new cell types. drug-pathway associations
can be identified using advanced machine learning methodologies that model the
complex molecular interactions [11]. Recently, integrative multitask sparse regression methods have been used to systematically identify biomarker combinations for
predicting drug outcomes [12]. Increasing evidence from recent studies thus poses
the hypothesis that common patterns in the activity profiles of genes and sensitivity/toxicity profiles of drugs can identify cellular response mechanisms and could be
used even in predicting the tissue type or cell context-specific toxicity outcome of
drug treatment.
This chapter is organized as follows: Sect. 4.2 introduces representative classes of
recent machine learning methods, with a specific emphasis on the matrix and tensor
factorization methods. Section 4.3 demonstrates the application of these methods to
identification of toxicogenomic relationships in example case studies, followed by
a discussion in Sect. 4.4. Section 4.5 concludes the chapter with current limitations
and future directions in these developments.
4.2 Machine Learning Methods
Machine learning algorithms search for patterns in data to extract useful information
[13, 14]. These algorithms learn a representation a.k.a. the model from existing data
samples and then utilize the model in different tasks. When applied to experimental
