single position or orientation, but rather a whole set with different probabilities. The
assignments are initially fuzzy and become better defined during an iterative
learning process, most commonly using an optimization algorithm called expectation maximization. This algorithm starts from an initial model for the 3D density
underlying the tomograms. In the expectation step different rotations and translations of the model are sampled and compared with the subtomograms in order to
approximate their probabilities. Taking into consideration the probabilities of the
rotations and translations, subtomograms are averaged (maximization step) and the
resulting structure is used as a reference for the next round of alignment. In addition
to approximation of the hidden variables, the software RELION also uses an
empirical Bayesian approach to correct for the CTF [15]. This approach is very
successful in single particle analysis [16] and has recently been generalized to
subtomogram averaging.
The above ML-based subtomogram averaging approach is computationally
extremely demanding. To process large amounts of data in a reasonable timeframe a
simpler quasi-expectation maximization algorithm is typically used. Instead of
assigning a continuous probability density function to each hidden variable only a
single value is assigned—the one with the highest probability. In mathematical
terms the probability density function would be a delta function, i.e., non-zero only
for a single value. This binary assignment greatly simplifies the alignment process,
but the radius of convergence is typically decreased compared to ML-based
alignment; the iterative optimization process is more prone to getting trapped in
local minima. Nevertheless, this simpler algorithm is sufficient for accurate particle
alignment in particular for large complexes and will be described in more detail in
the following.
9.4.2 Sampling Strategies
Starting from an initial structural model, the iterative quasi-expectation maximization algorithm aims to approximate the values of experimentally undetermined
translations and rotations of the respective subtomogram. In each iteration, different
rotations and translations of subtomograms are sampled in order to estimate the
expected rotations and translations (quasi-expectation step). These values are
determined on the basis of a similarity score, commonly the cross-correlation
coefficient (CCC) of the subtomogram with the iteratively updated reference. The
computation of the CCC is typically constrained to commonly sampled regions in
Fourier space (‘constrained cross correlation’) in order to compensate for the
missing wedge effect [9]. Taking into consideration the rotations and translations
yielding the maximum CCCs, subtomograms are averaged and the resulting
structure is used as a reference for the next round of alignment (quasi-maximization
step). This procedure is repeated until rotations and translations of subtomograms
have converged.
9 Structural Biology in Situ Using Cryo-Electron …
243
Précédent

- 260/339

Suivant