3.3 Assembly Fitting
Even with the help of high-resolution candidate structures of the
individual components, predicting the overall architecture of a
complex remains highly challenging. The multi-component search
for the best-fitting assembly model is complicated by the existence
of many local maxima, corresponding to models that produce
densities similar to the experimental map, but with incorrectly
swapped components, or wrongly oriented ones. This is particularly salient at lower resolutions. Additional experimental sources
may help resolve this ambiguity, but may not fully lift it. If the initial
fitting of a candidate models in the map can be obtained, for
example by homology [29], the search space could be reduced
significantly, avoiding unnecessary calculations. Such an approach
was used more than a decade ago to generate models for the first
mammalian ribosome structure at 8.7 A ˚ resolution (PDB ID:
2ZKR, 2ZKQ) [30] and the 80S ribosome from T. lanuginosus at
8.9 A ˚ resolution (PDB ID: 3JYX, 3JYW, 3JYV) [31], where the fit
of each modeled protein was optimized locally, using Mod-EM
local exhaustive optimization. However, if there is no initial knowledge about the approximate position of a candidate individual or
multiple models in the map, efficient methods to conduct that
search, such as gmfit [32], MultiFit [33], SITUS [4], and
γ-TEMPy [34], are necessary.
3.3.1 Initial Placement
To obtain initial positions for the components in the map, an
analysis of the map is conducted, using vector quantization
(VQ) (Fig. 2) [35]. This method reduces the complex density
data to a limited set of points (also known as codebook vectors),
and iteratively updates the set of points to reduce the distortion, a
measure of the distance between each point and the part of the data
it represents, until convergence. This method is similar to the
k-means procedure. Other methods can produce the components’
placement, for example by using a gaussian mixture model to
obtain a reduced representation formed by a set of gaussian functions [32]. These methods provide powerful ways to obtain initial
placement for components in the map, but there is no guarantee
that the codebook vectors or gaussian centers will correspond to
the centroid of components. Therefore, the placement of components with respect to those vectors need to be refined, and the
initial centroid placements strongly affect the refinement [34].
3.3.2 IQP and wICP
After generating code vectors with VQ, it is possible to follow the
same procedure on individual map components and obtain a putative match with the IQP algorithm. This match can then be iteratively refined using the wICP procedure, to reassign the data points
to a center, recomputing the centers, and iterating until convergence. This is similar in spirit to the expectation-maximization
procedure. The combination of IQP and wICP showed results
qualitatively similar to gmfit [35].
196
Tristan Cragnolini et al.
Précédent

- 201/346

Suivant