Non-deterministic Calibration
183
In short, MCMC avoids computing the denominator of Eq. 12 and instead utilizes
the proportionality:
π(θ|y) ∝ π(y|θ)π(θ),
(14)
which can be computed at any point θ ∈ R p . An iterative sampling procedure is
implemented to form the Markov chain. Since realizations of the chain are samples
of the posterior by definition, a sample-based approximation of the posterior
density can be obtained. As with standard Monte Carlo sampling, this sample-based
estimate of the posterior density converges as the number of samples in the chain,
N → ∞. In practice, N << ∞, and thus MCMC yields an approximation of the
posterior density.
A multitude of algorithms exist for forming this Markov chain. In general, they
involve a proposal distribution, J (θ ∗ |θ k−1 ), that depends only on the previous
sample in the chain, θ k−1 . A common choice is J (θ ∗ |θ k−1 ) = N(θ k−1 , V ), a
normal distribution centered at the previous sample with some covariance, V . The
candidate sample, θ ∗ , is either accepted or rejected based on the value of the
acceptance ratio:
A(θ
∗ , θ
k−1 ) =
π(θ ∗ |y)
π(θ k−1 |y)
=
π(y|θ ∗ )π(θ ∗ )
π(y|θ k−1 )π(θ k−1 )
.
(15)
A new sample yielding A(θ ∗ , θ k−1 ) > 1 is always accepted into the chain as it has
a high posterior probability than the previous sample. Accepting a sample implies
θ k = θ ∗ . If not, the new sample is accepted with probability A(θ ∗ , θ k−1 ), meaning
that the sample is more likely to be accepted the closer π(θ ∗ |y) is to π(θ k−1 |y). If
rejected, θ k = θ k−1 . This process is iterated until chain convergence. 1
The dependence of J (·) on θ k−1 means MCMC algorithms require initialization.
In this work, a least squares optimization was conducted to deterministically fit the
model parameters to available data and generate an initial guess in a region of high
posterior probability. This method accelerates chain convergence by reducing the
time spent searching for this region of high probability by a random walk over the
parameter space. Adaptive tuning of the proposal covariance V is typically required
during the initial stage of chain development as well. The resulting, non-stationary
period of searching and tuning is referred to as the burn-in period. The end of the
burn-in is defined by the point at which the Markov chain reaches a stationary
condition. By definition, samples obtained from the burn-in period are not drawn
from the targeted posterior distribution. In practice, an initial percentage of the chain
is attributed to burn-in and discarded.
1 Diagnosing chain convergence can be challenging, and readers are referred to [6, 13] for more
information.
183
In short, MCMC avoids computing the denominator of Eq. 12 and instead utilizes
the proportionality:
π(θ|y) ∝ π(y|θ)π(θ),
(14)
which can be computed at any point θ ∈ R p . An iterative sampling procedure is
implemented to form the Markov chain. Since realizations of the chain are samples
of the posterior by definition, a sample-based approximation of the posterior
density can be obtained. As with standard Monte Carlo sampling, this sample-based
estimate of the posterior density converges as the number of samples in the chain,
N → ∞. In practice, N << ∞, and thus MCMC yields an approximation of the
posterior density.
A multitude of algorithms exist for forming this Markov chain. In general, they
involve a proposal distribution, J (θ ∗ |θ k−1 ), that depends only on the previous
sample in the chain, θ k−1 . A common choice is J (θ ∗ |θ k−1 ) = N(θ k−1 , V ), a
normal distribution centered at the previous sample with some covariance, V . The
candidate sample, θ ∗ , is either accepted or rejected based on the value of the
acceptance ratio:
A(θ
∗ , θ
k−1 ) =
π(θ ∗ |y)
π(θ k−1 |y)
=
π(y|θ ∗ )π(θ ∗ )
π(y|θ k−1 )π(θ k−1 )
.
(15)
A new sample yielding A(θ ∗ , θ k−1 ) > 1 is always accepted into the chain as it has
a high posterior probability than the previous sample. Accepting a sample implies
θ k = θ ∗ . If not, the new sample is accepted with probability A(θ ∗ , θ k−1 ), meaning
that the sample is more likely to be accepted the closer π(θ ∗ |y) is to π(θ k−1 |y). If
rejected, θ k = θ k−1 . This process is iterated until chain convergence. 1
The dependence of J (·) on θ k−1 means MCMC algorithms require initialization.
In this work, a least squares optimization was conducted to deterministically fit the
model parameters to available data and generate an initial guess in a region of high
posterior probability. This method accelerates chain convergence by reducing the
time spent searching for this region of high probability by a random walk over the
parameter space. Adaptive tuning of the proposal covariance V is typically required
during the initial stage of chain development as well. The resulting, non-stationary
period of searching and tuning is referred to as the burn-in period. The end of the
burn-in is defined by the point at which the Markov chain reaches a stationary
condition. By definition, samples obtained from the burn-in period are not drawn
from the targeted posterior distribution. In practice, an initial percentage of the chain
is attributed to burn-in and discarded.
1 Diagnosing chain convergence can be challenging, and readers are referred to [6, 13] for more
information.
