70
S. A. Khan et al.
A limitation of current methods is the ability to handle missing values particularly
when considering overlap between different data sets. Compared to many other “Big
Data” study areas, biomedical data is less extensive and contains more missing values.
The LDA method used with the PTGS had the advantage that the entire CMap data
set could be used to derive the initial components, whereas the GFA method required
at least some overlap between all variables, reducing the amount of gene expression
data used. Tensor factorization methods are even less tolerant of missing values.
Therefore unique methodological considerations and trade-offs apply to each study.
A key outcome of this joint analysis is the ability to predict the toxicity outcomes
of compound treatment. The prediction of unexpected toxic effects is a challenging
and important goal in toxicology. The presented first steps in computational toxicogenomic open up a systematic way for genomics-driven prediction of toxic effects. In
addition, these provide novel mechanistic insights into the links between genomic
measurements of cells and toxicological profiles of drugs. Gene expression response
profiles of drugs present a popular systems-level view, while toxicity profiles summarize the drugs’ phenotypic behavior. Large repositories of gene expression and
drug sensitivity profiles such as those emerging from NCI60, CMap, CCLE, Sanger,
and LINCS profile cellular responses at several levels of detail in a cell contextspecific manner. With the emergence of heterogeneous and partially paired data sets,
joint factorizations are gaining popularity to identify novel dependency patterns, as
well as to design powerful predictive applications [48, 53]. These recent advances in
machine learning, and especially the methods described in Sect. 4.2, enable systematic analysis of such large data repositories to provide novel toxicogenomic insights
and predictions.
4.5 Conclusion and Future Directions
State-of-the-art machine learning methods have been presented here for modeling
various toxicogenomic relationships. These advanced computational methodologies
enable integration of disparate, high-dimensional data sources, including but not
limited to omics, drug screening, chemical structures, and drug-targets to achieve
novel toxicogenomic analysis in terms of:
(i) providing means for predicting personalized toxicity outcomes,
(ii) identifying toxic modes of action, and
(iii) enabling quantitative structure activity modeling.
The here presented works suggest novel directions for future analysis. From the
application perspective, matrix and tensor factorization methods can serve to stimulate integrative analysis of various toxicological and toxicogenomic data sets to
suggest novel hypotheses. For example, a joint analysis of omics, toxicity, and drugtarget data sets can help to identify disparate target-driven and toxic molecular mechanisms. Integrative analysis with drug-side effect repositories can help draw novel
interactions between disease, side effect, and toxicity mechanisms. From a holistic
Précédent

- 84/416

Suivant