60
S. A. Khan et al.
data, the model formulation and learning processes take into account various forms of
inherent noise and corruptions in the measurements to learn a cleaner representation
of the data. In toxicogenomics, similar to many other real-life applications, data is
expected to be noisy and high dimensional, contains missing values, and may also
include correlated variables, which all make direct analysis complicated [15]. In
such cases, the key feature of machine learning is to identify a low-dimensional,
hidden representation that captures and summarizes the relevant information for the
toxicogenomic modeling task. These summaries can then be used to understand the
compound’s MoA and/or predict the cellular outcomes of the drugs in different cell
contexts.
4.2.1 Matrix Factorization
In machine learning, matrix factorization (MF) is a well-established approach to summarize a data set through unobserved features that explain why some parts of the data
are similar. MF has applications in broad range of scientific domains, and it is widely
used in several applications, including prediction of missing values, dimensionality
reduction, as well as data visualization [16, 17]. This wide applicability comes from
the assumption that MF can be seen as a means of describing the underlying processes which generated the data. Specifically, MF assumes that measurements have
been produced by a combination of a number of latent processes and aims to identify
the factors (a.k.a. components) that describe these processes. Figure 4.1 shows a
visual illustration of matrix factorization, where a matrix X is factorized into distinct
low-dimensional components. This component decomposition is valuable for many
applications, as the different components can be related to separate mechanisms that
may have contributed to the data. Several matrix factorization methods have been
proposed for various applications, including factor analysis (FA), principal component analysis (PCA), and Latent Dirichlet Allocation (LDA, see Sect. 2.2) [18–20].
While FA and PCA are designed for continuous data sets, LDA is formulated for
discrete data sets.
Fig. 4.1 Visual
representation of matrix
factorization. The data
matrix X is factorized into
low-dimensional matrices Z
and W that capture the key
statistical patterns in the data
S. A. Khan et al.
data, the model formulation and learning processes take into account various forms of
inherent noise and corruptions in the measurements to learn a cleaner representation
of the data. In toxicogenomics, similar to many other real-life applications, data is
expected to be noisy and high dimensional, contains missing values, and may also
include correlated variables, which all make direct analysis complicated [15]. In
such cases, the key feature of machine learning is to identify a low-dimensional,
hidden representation that captures and summarizes the relevant information for the
toxicogenomic modeling task. These summaries can then be used to understand the
compound’s MoA and/or predict the cellular outcomes of the drugs in different cell
contexts.
4.2.1 Matrix Factorization
In machine learning, matrix factorization (MF) is a well-established approach to summarize a data set through unobserved features that explain why some parts of the data
are similar. MF has applications in broad range of scientific domains, and it is widely
used in several applications, including prediction of missing values, dimensionality
reduction, as well as data visualization [16, 17]. This wide applicability comes from
the assumption that MF can be seen as a means of describing the underlying processes which generated the data. Specifically, MF assumes that measurements have
been produced by a combination of a number of latent processes and aims to identify
the factors (a.k.a. components) that describe these processes. Figure 4.1 shows a
visual illustration of matrix factorization, where a matrix X is factorized into distinct
low-dimensional components. This component decomposition is valuable for many
applications, as the different components can be related to separate mechanisms that
may have contributed to the data. Several matrix factorization methods have been
proposed for various applications, including factor analysis (FA), principal component analysis (PCA), and Latent Dirichlet Allocation (LDA, see Sect. 2.2) [18–20].
While FA and PCA are designed for continuous data sets, LDA is formulated for
discrete data sets.
Fig. 4.1 Visual
representation of matrix
factorization. The data
matrix X is factorized into
low-dimensional matrices Z
and W that capture the key
statistical patterns in the data
