78
Notice that this approach requires all variables to be informed at all locations, that is, the
dataset must be homotopic. Samples with missing variables are common in geological data
sets for many reasons. The missing data must be imputed (inferred) to permit the measured
data to be used to their full extent. Imputation methods for geological data should address
spatial structure and multivariate complexity. If some variables are missing, an imputation
process should be applied (Silva and Deutsch, 2016).
3.1 Log-ratio transform
Geological data are frequently reported in terms of the grades of different elements or the
mineralogical proportions present in the rock. These sets of variables form a closed array
or a composition, as their sum must add to the whole of the material. If all elements are
considered, they should sum to 100%. If mineralogical proportions are used, they should
add up to 1. This translates in a dependence between the variables, as there is always one less
degree of freedom in the system, than variables available. Correlations are also distorted by
this dependence, and this can lead to wrong inference and interpretations (Pawlowsky-Glahn
and Olea, 2004). This also occurs with sub-compositions, that is when only a subset of all the
variables that form the composition are used (Pawlowsky-Glahn and Egozcue, 2006).
Considering the data {
}
p p
X
n
α
( )
u α
= …
α
form a composition, compositional data analysis solves this closure problem by applying a log-ratio transformation of the data, so that
any further statistical manipulation respects the constraint that the sum adds to 100%
(Pawlowsky-Glahn and Egozcue, 2006).
The three most important transformations are:
1. Additive log-ratio (alr): it is the logarithm of the ratio between each component and one
of the variables, in our case, the filler variable, and was introduced by Aitchison (1982)
(see also Pawlowsky and Egozcue, 2006; Aitchison, 1986);
2. Centered log-ratio (clr): it is the logarithm of the ratio between each component and the
geometric mean of the parts, and was also introduced by Aitchison (1982); and
3. Isometric log-ratio (ilr): it is obtained by projecting the composition over an orthonormal
basis with (p – 1) dimensions. It was introduced by Egozcue et al. (2003).
In our methodology, we used the alr transform, therefore, the data are transformed to a
new vector variable, as follows:
Z
log
X
R
log
X
R
log
X p
X
( )
u α =
( )
u α
( )
u α
⎛
⎝
⎜
⎛ ⎛
⎝ ⎝
⎞
⎠
⎟
⎠ ⎠
( )
u α
( )
u α
( )
⎛
⎝
⎜
⎝ ⎝
⎞
⎠
⎟
⎞ ⎞
⎠ ⎠
…
−
2
X X
1
X X
l
( )
u α ⎞
⎟
⎞ ⎞
⎛
⎜
⎛ ⎛
1
,
R
, log
( )
⎝
⎜
⎝ ⎝
⎠
⎟
⎠ ⎠
,
α
( )
u α α
( )
α
⎛
⎝
⎜
⎛ ⎛
⎝ ⎝
⎞
⎠
⎟
⎞ ⎞
⎠ ⎠
⎛
⎝
⎜
⎛ ⎛
⎝ ⎝
⎞
⎠
⎟
⎞ ⎞
⎠ ⎠
∀ =
α
…
R(
n
1, , (4)
The p dimensional vector X(u α ), becomes a (p – 1) dimensional vector Z(u α ).
3.2 Principal component transform
The next step is to transform the log-ratios obtained in the previous step to linearly uncorrelated factors by using Principal Component Analysis. This principal component transformation finds a set of orthogonal linear axes that passes through the multivariate mean of
the log-ratio transformed variables, and is such that the variance of the projections of the
original log-ratios onto the first axis (called first principal component) is maximized. Axes
corresponding to subsequent principal components are determined orthogonal to the previous ones, and with maximum variance (Howarth, 2017).
Principal components are found after an eigen-decomposition of the covariance matrix of
the variable of interest (Wackernagel, 2003). The steps required are:
• Compute the mean vector of the variables in vector Z:
m (
)
m m
m
Z
Z p
Z
…
m m
Z m
Z
Z
Z
Z Z
m Z
(5)
Notice that this approach requires all variables to be informed at all locations, that is, the
dataset must be homotopic. Samples with missing variables are common in geological data
sets for many reasons. The missing data must be imputed (inferred) to permit the measured
data to be used to their full extent. Imputation methods for geological data should address
spatial structure and multivariate complexity. If some variables are missing, an imputation
process should be applied (Silva and Deutsch, 2016).
3.1 Log-ratio transform
Geological data are frequently reported in terms of the grades of different elements or the
mineralogical proportions present in the rock. These sets of variables form a closed array
or a composition, as their sum must add to the whole of the material. If all elements are
considered, they should sum to 100%. If mineralogical proportions are used, they should
add up to 1. This translates in a dependence between the variables, as there is always one less
degree of freedom in the system, than variables available. Correlations are also distorted by
this dependence, and this can lead to wrong inference and interpretations (Pawlowsky-Glahn
and Olea, 2004). This also occurs with sub-compositions, that is when only a subset of all the
variables that form the composition are used (Pawlowsky-Glahn and Egozcue, 2006).
Considering the data {
}
p p
X
n
α
( )
u α
= …
α
form a composition, compositional data analysis solves this closure problem by applying a log-ratio transformation of the data, so that
any further statistical manipulation respects the constraint that the sum adds to 100%
(Pawlowsky-Glahn and Egozcue, 2006).
The three most important transformations are:
1. Additive log-ratio (alr): it is the logarithm of the ratio between each component and one
of the variables, in our case, the filler variable, and was introduced by Aitchison (1982)
(see also Pawlowsky and Egozcue, 2006; Aitchison, 1986);
2. Centered log-ratio (clr): it is the logarithm of the ratio between each component and the
geometric mean of the parts, and was also introduced by Aitchison (1982); and
3. Isometric log-ratio (ilr): it is obtained by projecting the composition over an orthonormal
basis with (p – 1) dimensions. It was introduced by Egozcue et al. (2003).
In our methodology, we used the alr transform, therefore, the data are transformed to a
new vector variable, as follows:
Z
log
X
R
log
X
R
log
X p
X
( )
u α =
( )
u α
( )
u α
⎛
⎝
⎜
⎛ ⎛
⎝ ⎝
⎞
⎠
⎟
⎠ ⎠
( )
u α
( )
u α
( )
⎛
⎝
⎜
⎝ ⎝
⎞
⎠
⎟
⎞ ⎞
⎠ ⎠
…
−
2
X X
1
X X
l
( )
u α ⎞
⎟
⎞ ⎞
⎛
⎜
⎛ ⎛
1
,
R
, log
( )
⎝
⎜
⎝ ⎝
⎠
⎟
⎠ ⎠
,
α
( )
u α α
( )
α
⎛
⎝
⎜
⎛ ⎛
⎝ ⎝
⎞
⎠
⎟
⎞ ⎞
⎠ ⎠
⎛
⎝
⎜
⎛ ⎛
⎝ ⎝
⎞
⎠
⎟
⎞ ⎞
⎠ ⎠
∀ =
α
…
R(
n
1, , (4)
The p dimensional vector X(u α ), becomes a (p – 1) dimensional vector Z(u α ).
3.2 Principal component transform
The next step is to transform the log-ratios obtained in the previous step to linearly uncorrelated factors by using Principal Component Analysis. This principal component transformation finds a set of orthogonal linear axes that passes through the multivariate mean of
the log-ratio transformed variables, and is such that the variance of the projections of the
original log-ratios onto the first axis (called first principal component) is maximized. Axes
corresponding to subsequent principal components are determined orthogonal to the previous ones, and with maximum variance (Howarth, 2017).
Principal components are found after an eigen-decomposition of the covariance matrix of
the variable of interest (Wackernagel, 2003). The steps required are:
• Compute the mean vector of the variables in vector Z:
m (
)
m m
m
Z
Z p
Z
…
m m
Z m
Z
Z
Z
Z Z
m Z
(5)
