9.5 Subtomogram Classification
The purpose of classification approaches is to disentangle compositional and conformational heterogeneity of the macromolecular complex of interest. However, the
incomplete spatial sampling originating from the missing wedge, the low SNR, the
unknown number of classes and the often-unbalanced class occupation (i.e., relative
abundances of conformers differ strongly) in tomographic data make a reliable
classification of subtomograms challenging. In this section, various classification
approaches are introduced that either rely on a prior, separate alignment of the
subtomograms or allow simultaneous alignment and classification based on multiple references.
9.5.1 Uncoupled Alignment and Classification
For classification by constrained principal-component analysis (CPCA), a matrix of
constrained correlation coefficients is computed for all pairs of pre-aligned subtomograms and subsequently analysed by principal-component analysis (PCA) and
k-means clustering [20, 21]. A major advantage of this approach is that it is conceptually easy and its independence from the alignment process allows straightforward focusing on specific areas of interest. Due to the pairwise cross-correlation,
this approach is computationally very expensive and only applicable to mediumsized datasets (<10,000 subtomograms). However, in many instances the analysis
can be performed in smaller subsets of the data and the subtomograms can often be
downsampled (‘binned’) to distinguish their main distinctive features. In order to
recover also small populations of structurally or conformationally distinct macromolecular complexes, the number of output classes typically strongly oversamples
the number of expected distinct classes in the data. Redundant classes can then be
merged to reduce the number of classes and increase the SNRs of the individual
structures.
9.5.2 Simultaneous Alignment and Classification
Other commonly used methods base on simultaneous subtomogram alignment to different reference structures (‘multi-reference procedures’). Classification is achieved by
assigning the subtomogram to the reference class with the highest score [13, 14, 17].
Since the computational effort of simultaneous alignment to multiple references is much
larger than for a single reference, acceleration of the alignment by FRM is strongly
beneficial [22]. A further problem of multi-reference alignment is that the noise outside
the area of structural variation can strongly decrease the classification accuracy. To
overcome this limitation, multi-reference alignment has recently been extended to
246
S. Pfeffer and F. Förster
The purpose of classification approaches is to disentangle compositional and conformational heterogeneity of the macromolecular complex of interest. However, the
incomplete spatial sampling originating from the missing wedge, the low SNR, the
unknown number of classes and the often-unbalanced class occupation (i.e., relative
abundances of conformers differ strongly) in tomographic data make a reliable
classification of subtomograms challenging. In this section, various classification
approaches are introduced that either rely on a prior, separate alignment of the
subtomograms or allow simultaneous alignment and classification based on multiple references.
9.5.1 Uncoupled Alignment and Classification
For classification by constrained principal-component analysis (CPCA), a matrix of
constrained correlation coefficients is computed for all pairs of pre-aligned subtomograms and subsequently analysed by principal-component analysis (PCA) and
k-means clustering [20, 21]. A major advantage of this approach is that it is conceptually easy and its independence from the alignment process allows straightforward focusing on specific areas of interest. Due to the pairwise cross-correlation,
this approach is computationally very expensive and only applicable to mediumsized datasets (<10,000 subtomograms). However, in many instances the analysis
can be performed in smaller subsets of the data and the subtomograms can often be
downsampled (‘binned’) to distinguish their main distinctive features. In order to
recover also small populations of structurally or conformationally distinct macromolecular complexes, the number of output classes typically strongly oversamples
the number of expected distinct classes in the data. Redundant classes can then be
merged to reduce the number of classes and increase the SNRs of the individual
structures.
9.5.2 Simultaneous Alignment and Classification
Other commonly used methods base on simultaneous subtomogram alignment to different reference structures (‘multi-reference procedures’). Classification is achieved by
assigning the subtomogram to the reference class with the highest score [13, 14, 17].
Since the computational effort of simultaneous alignment to multiple references is much
larger than for a single reference, acceleration of the alignment by FRM is strongly
beneficial [22]. A further problem of multi-reference alignment is that the noise outside
the area of structural variation can strongly decrease the classification accuracy. To
overcome this limitation, multi-reference alignment has recently been extended to
246
S. Pfeffer and F. Förster
