126
Some simplifications of the LMC were proposed in the literature. Almeida and Journel
(1994) suggested a Markov Model of coregionalization for modeling the cross-covariances. In
the Markov Model, the cross variogram model is either proportional to the variogram of the primary variable (MM1) or the secondary variable (MM2). These Markov Models are usually used
with collocated cokriging, which uses only the collocated secondary datum for estimation. The
problem with collocated cokriging is that the variance is inflated (Deutsch and Journel, 1998).
To avoid variance inflation, Babak and Deutsch (2009) proposed to use the secondary data at
the location being estimated and at the locations of the primary data. This approach is denoted
as intrinsic collocated cokriging (ICCK). The limitation of collocated and intrinsic collocated
cokriging is that the secondary data must be sampled at all the nodes of the grid. We are interested in situations where the secondary data is more sampled than the primary data, but it does
not cover the entire grid. One alternative is to perform the simulation hierarchically (Almeida
and Journel, 1994). In this case, the secondary is simulated first and used as secondary information with ICCK in the simulation of the primary variable. However, this approach demands
more computer memory and processing time than simulating directly the primary variable.
In situations where the secondary data are more sampled than the primary, multiple imputation (MI) techniques are available and consist of simulating first at the locations of the secondary and second at the nodes of the simulation grid (Barnett and Deutsch, 2015; Soares et al,
2017; Silva e Deutsch, 2018). Barnett and Deutsch (2015) proposed two multiple imputation
methodologies: parametric and non-parametric to impute missing geological data based on
Bayesian updating to generated isotopic datasets. The parametric method assumes that the
joint distribution is multivariate Gaussian while the non-parametric calculates the joint distribution from the scatter plots. Similarly, the methods used by Soares et al. (2017) and Silva
and Deutsch (2018) used the scatter plots to obtain the joint distribution. The problem is that
building a scatter plot requires a subset where the primary and secondary data are isotopic. As
a result, these methods are not applicable when the data set is completely heterotopic. In this
context, the parametric approach proposed by Barnett and Deutsch (2015) was used.
The parametric approach (Barnett and Deutsch, 2015) requires the coefficient of correlation between the primary and secondary variable. When these variables are completely heterotopic, the coefficient of correlation may be obtained by extrapolating the experimental cross
correlogram (Minnitt and Deutsch, 2014) to the zero distance. In this paper, we combine
the approach of Minnitt and Deutsch (2014) to obtain the coefficient of correlation with
the parametric multiple imputation method of Barnett and Deutsch (2015). The parametric
method presented by Barnett and Deutsch (2015) uses Bayesian Updating to account for the
secondary data. Bayesian Updating is a form of collocated cokriging (Doyen, 1996; Deutsch
and Zanon, 2004; Ren, 2007) and does not require the LMC, only the coefficient of correlation between the primary and secondary variable. We assume that obtaining the coefficient
of correlation is easier than modeling the LMC.
As data sources do not have the same quality and are totally heterotopic and spatially
correlated, it is difficult to establish the correlation between these data types. Minnitt and
Deutsch (2014) suggested to infer their statistical relationships using the experimental crosscorrelogram values extrapolated so the correlation coefficient at zero lag can be obtained.
This work proposes a framework to simulate hard and soft data considering all information regarding the relationship of the multiple variables without modelling the LMC (linear
model coregionalization). This framework combines two approaches: Bayesian Updating
and Sequential Gaussian Simulation. The simulated models were compared with the models
obtained by Sequential Gaussian Simulation (SGS) using only the original hard data. To
illustrate the methodology, we used a modified version of the Walker lake dataset (Isaaks
and Srivastava, 1989).
2 METHODOLOGY
This framework combines two approaches: Bayesian Updating (BU) and Sequential
Gaussian Simulation (SGS). Firstly, at each soft data location, the hard data values are
Some simplifications of the LMC were proposed in the literature. Almeida and Journel
(1994) suggested a Markov Model of coregionalization for modeling the cross-covariances. In
the Markov Model, the cross variogram model is either proportional to the variogram of the primary variable (MM1) or the secondary variable (MM2). These Markov Models are usually used
with collocated cokriging, which uses only the collocated secondary datum for estimation. The
problem with collocated cokriging is that the variance is inflated (Deutsch and Journel, 1998).
To avoid variance inflation, Babak and Deutsch (2009) proposed to use the secondary data at
the location being estimated and at the locations of the primary data. This approach is denoted
as intrinsic collocated cokriging (ICCK). The limitation of collocated and intrinsic collocated
cokriging is that the secondary data must be sampled at all the nodes of the grid. We are interested in situations where the secondary data is more sampled than the primary data, but it does
not cover the entire grid. One alternative is to perform the simulation hierarchically (Almeida
and Journel, 1994). In this case, the secondary is simulated first and used as secondary information with ICCK in the simulation of the primary variable. However, this approach demands
more computer memory and processing time than simulating directly the primary variable.
In situations where the secondary data are more sampled than the primary, multiple imputation (MI) techniques are available and consist of simulating first at the locations of the secondary and second at the nodes of the simulation grid (Barnett and Deutsch, 2015; Soares et al,
2017; Silva e Deutsch, 2018). Barnett and Deutsch (2015) proposed two multiple imputation
methodologies: parametric and non-parametric to impute missing geological data based on
Bayesian updating to generated isotopic datasets. The parametric method assumes that the
joint distribution is multivariate Gaussian while the non-parametric calculates the joint distribution from the scatter plots. Similarly, the methods used by Soares et al. (2017) and Silva
and Deutsch (2018) used the scatter plots to obtain the joint distribution. The problem is that
building a scatter plot requires a subset where the primary and secondary data are isotopic. As
a result, these methods are not applicable when the data set is completely heterotopic. In this
context, the parametric approach proposed by Barnett and Deutsch (2015) was used.
The parametric approach (Barnett and Deutsch, 2015) requires the coefficient of correlation between the primary and secondary variable. When these variables are completely heterotopic, the coefficient of correlation may be obtained by extrapolating the experimental cross
correlogram (Minnitt and Deutsch, 2014) to the zero distance. In this paper, we combine
the approach of Minnitt and Deutsch (2014) to obtain the coefficient of correlation with
the parametric multiple imputation method of Barnett and Deutsch (2015). The parametric
method presented by Barnett and Deutsch (2015) uses Bayesian Updating to account for the
secondary data. Bayesian Updating is a form of collocated cokriging (Doyen, 1996; Deutsch
and Zanon, 2004; Ren, 2007) and does not require the LMC, only the coefficient of correlation between the primary and secondary variable. We assume that obtaining the coefficient
of correlation is easier than modeling the LMC.
As data sources do not have the same quality and are totally heterotopic and spatially
correlated, it is difficult to establish the correlation between these data types. Minnitt and
Deutsch (2014) suggested to infer their statistical relationships using the experimental crosscorrelogram values extrapolated so the correlation coefficient at zero lag can be obtained.
This work proposes a framework to simulate hard and soft data considering all information regarding the relationship of the multiple variables without modelling the LMC (linear
model coregionalization). This framework combines two approaches: Bayesian Updating
and Sequential Gaussian Simulation. The simulated models were compared with the models
obtained by Sequential Gaussian Simulation (SGS) using only the original hard data. To
illustrate the methodology, we used a modified version of the Walker lake dataset (Isaaks
and Srivastava, 1989).
2 METHODOLOGY
This framework combines two approaches: Bayesian Updating (BU) and Sequential
Gaussian Simulation (SGS). Firstly, at each soft data location, the hard data values are
