4 Matrix and Tensor Factorization Methods …
63
4.2.4 Tensor Factorization
A tensor is a multidimensional arrayX ∈ R
I 1 ×I 2 ×...×I j and a generalization of matrices and vectors to higher order spaces. Tensors are therefore useful for representing
data that has more than two dimensions. Such representation allows investigation of
relationships that span multidimensional constructs. Mathematically, a tensor is also
commonly defined as an element of space induced by the tensor product of vector
spaces.
In order to capture the highly structured patterns of a multidimensional data set,
tensor methods employ constrained formulations that help to avoid the overfitting
problem [32]. A key characteristic of these tensor formulations is that they have fewer
parameters than their matrix counterparts. Analogous to matrix factorizations presented in Sect. 4.2.2, there exist several tensor factorization methods that can be used
to discover underlying dependencies in the data [32]. CANDECOMP/PARAFAC
(CP; [33, 34]) and Tucker family [35] are the most widely used tensor decomposition methods. The interested reader is referred to [32] for a comprehensive review
of various tensor factorization methods.
Tensor factorizations have obtained significant success in a large number of
domains, including chemometrics, psychometrics, bioinformatics, and have shown
immense promise for advanced applications in toxicology and toxicogenomics. For
example, tensor factorization has been used to explore stimuli-variant gene expression patterns [36], as well as in integrating phenotypic responses from multiple
studies [37, 38], modeling dependencies between metabolic and gene expression
networks [39], as well as in joint QSAR and toxicogenomic analysis [40, 41].
CP factorization, also known as the canonical decomposition or parallel factor
analysis [33, 34], is the most widely used tensor factorization method. CP is a natural
extension of matrix factorization to arrays of order 3 or more as shown in Fig. 4.3.
The method can be seen as carrying out simultaneous factor analysis on multiple
slabs (matrices) of a tensor such that the factors of each slab differ just by a scale. CP
factorization is defined in a symmetric fashion over all the modes, such that a tensor is
decomposed into a sum of rank-one tensors, where each rank-one tensor is the outer
product of the latent vector in all modes. For a third order tensor X ∈ R
N ×D×L , a
rank-K CP is represented as:
Fig. 4.3 Visual
representation of CP
factorization of a third order
tensor. The data tensor X is
factorized into
low-dimensional matrices Z,
U, and W that capture the
key statistical patterns in the
data
63
4.2.4 Tensor Factorization
A tensor is a multidimensional arrayX ∈ R
I 1 ×I 2 ×...×I j and a generalization of matrices and vectors to higher order spaces. Tensors are therefore useful for representing
data that has more than two dimensions. Such representation allows investigation of
relationships that span multidimensional constructs. Mathematically, a tensor is also
commonly defined as an element of space induced by the tensor product of vector
spaces.
In order to capture the highly structured patterns of a multidimensional data set,
tensor methods employ constrained formulations that help to avoid the overfitting
problem [32]. A key characteristic of these tensor formulations is that they have fewer
parameters than their matrix counterparts. Analogous to matrix factorizations presented in Sect. 4.2.2, there exist several tensor factorization methods that can be used
to discover underlying dependencies in the data [32]. CANDECOMP/PARAFAC
(CP; [33, 34]) and Tucker family [35] are the most widely used tensor decomposition methods. The interested reader is referred to [32] for a comprehensive review
of various tensor factorization methods.
Tensor factorizations have obtained significant success in a large number of
domains, including chemometrics, psychometrics, bioinformatics, and have shown
immense promise for advanced applications in toxicology and toxicogenomics. For
example, tensor factorization has been used to explore stimuli-variant gene expression patterns [36], as well as in integrating phenotypic responses from multiple
studies [37, 38], modeling dependencies between metabolic and gene expression
networks [39], as well as in joint QSAR and toxicogenomic analysis [40, 41].
CP factorization, also known as the canonical decomposition or parallel factor
analysis [33, 34], is the most widely used tensor factorization method. CP is a natural
extension of matrix factorization to arrays of order 3 or more as shown in Fig. 4.3.
The method can be seen as carrying out simultaneous factor analysis on multiple
slabs (matrices) of a tensor such that the factors of each slab differ just by a scale. CP
factorization is defined in a symmetric fashion over all the modes, such that a tensor is
decomposed into a sum of rank-one tensors, where each rank-one tensor is the outer
product of the latent vector in all modes. For a third order tensor X ∈ R
N ×D×L , a
rank-K CP is represented as:
Fig. 4.3 Visual
representation of CP
factorization of a third order
tensor. The data tensor X is
factorized into
low-dimensional matrices Z,
U, and W that capture the
key statistical patterns in the
data
