5 Methods
5.1 Candidate Model
Generation
Having a model close to the final solution increases the speed and
efficacy of automated fitting. For Identification of experimental
candidate models, the Protein Data Bank (PDB) is a useful resource
that contains atomic models derived from experimental data
(mainly X-ray crystallography, NMR spectroscopy, and cryo-EM).
It is possible to search for structures by sequence (see Note 4),
keyword, experimental technique, and author information among
other features. A search by sequence can yield experimental models
suitable for use as initial models.
Once an initial model has been identified it is important to
consider how complete the model is, as it is common for experimental models to lack flexible regions such as loops. In such cases, it
is possible to use further experimental structures to model these
regions. To identify further structures from the PDB, a sequence or
keyword search can sometimes yield meaningful results. Alternatively, the PDBeFold server (see Note 5), which compares the 3D
fold of input structures with all models in the PDB, can be used to
identify structural/sequence homologs. This server takes a PDB
coordinate file as input. It is also possible to use a PDB identifier
(e.g., 5ogc) to search.
If no suitable experimental structure is available, comparative/
ab initio models can be used. Many methods have been proposed to
generate models from the protein sequence, and blind competitions have shown that modern methods can achieve very high
accuracy in some cases [51]. MODELLER [2, 78] and SWISSMODEL [21] software suites are among the most popular for
comparative modeling. Rosetta provides both ab initio and comparative modeling tools. Other popular tools are I-Tasser [79] and
RaptorX [80]. State of the art ab initio (as showcased in recent
CASP competitions) often use co-evolution to more accurately
predict contacts and help in modeling [81].
5.2 Rigid-Body
Fitting
There are many programs that rigidly fit a candidate model into a
map by optimizing the CCC. Possibly one of the most commonly
used is the fit-in-map function implemented in UCSF-Chimera.
The method requires an initial manual alignment (see Note 6) and
the cryo-EM map. User input should be defined as use map
simulated from atoms: Average map resolution. Optimize: correlation. Allow: rotation, shift, and move whole molecules. Generally,
this method works well when the initial approximate placement of
the candidate model is obvious (see Note 6).
The Mod-EM method is implemented in the latest version of
MODELLER along with the python scripts for executing
Mod-EM (see Note 7). It provides options for both local and global
search for a single candidate model. Mod-EM requires the
212
Tristan Cragnolini et al.
5.1 Candidate Model
Generation
Having a model close to the final solution increases the speed and
efficacy of automated fitting. For Identification of experimental
candidate models, the Protein Data Bank (PDB) is a useful resource
that contains atomic models derived from experimental data
(mainly X-ray crystallography, NMR spectroscopy, and cryo-EM).
It is possible to search for structures by sequence (see Note 4),
keyword, experimental technique, and author information among
other features. A search by sequence can yield experimental models
suitable for use as initial models.
Once an initial model has been identified it is important to
consider how complete the model is, as it is common for experimental models to lack flexible regions such as loops. In such cases, it
is possible to use further experimental structures to model these
regions. To identify further structures from the PDB, a sequence or
keyword search can sometimes yield meaningful results. Alternatively, the PDBeFold server (see Note 5), which compares the 3D
fold of input structures with all models in the PDB, can be used to
identify structural/sequence homologs. This server takes a PDB
coordinate file as input. It is also possible to use a PDB identifier
(e.g., 5ogc) to search.
If no suitable experimental structure is available, comparative/
ab initio models can be used. Many methods have been proposed to
generate models from the protein sequence, and blind competitions have shown that modern methods can achieve very high
accuracy in some cases [51]. MODELLER [2, 78] and SWISSMODEL [21] software suites are among the most popular for
comparative modeling. Rosetta provides both ab initio and comparative modeling tools. Other popular tools are I-Tasser [79] and
RaptorX [80]. State of the art ab initio (as showcased in recent
CASP competitions) often use co-evolution to more accurately
predict contacts and help in modeling [81].
5.2 Rigid-Body
Fitting
There are many programs that rigidly fit a candidate model into a
map by optimizing the CCC. Possibly one of the most commonly
used is the fit-in-map function implemented in UCSF-Chimera.
The method requires an initial manual alignment (see Note 6) and
the cryo-EM map. User input should be defined as use map
simulated from atoms: Average map resolution. Optimize: correlation. Allow: rotation, shift, and move whole molecules. Generally,
this method works well when the initial approximate placement of
the candidate model is obvious (see Note 6).
The Mod-EM method is implemented in the latest version of
MODELLER along with the python scripts for executing
Mod-EM (see Note 7). It provides options for both local and global
search for a single candidate model. Mod-EM requires the
212
Tristan Cragnolini et al.
