76
conserved N- and C-terminal regions and insertions that are unique
to individual protein sequences, since PAML (the program used for
reconstruction of ancestral protein sequences) will attempt to reconstruct one ancestral residue for each column in the alignment.
1. Align the protein sequences from Subheading 3.1 using the
MUSCLE server [21] (see Note 6).
2. Open the alignment in SeaView [22].
3. Inspect the alignment carefully. If necessary, make manual
adjustments to improve the alignment, particularly in regions
with many gaps.
4. Trim the alignment, removing insertions that would not be
present in the ancestral proteins of interest (see Note 7).
Remove poorly conserved N- and C- terminal extensions,
including the N-terminal signal peptides in the case of SBPs.
5. Save the alignment in PHYLIP interleaved format.
Phylogenetic analysis and reconstruction of ancestral protein
sequences using the maximum-likelihood method requires a probabilistic model of protein evolution, which is needed to evaluate
the likelihood of different evolutionary scenarios. The main component of this evolutionary model is an empirical substitution
matrix that specifies the frequency of each amino acid and the substitution rate between each pair of amino acids. Some modifications are also available to make the basic evolutionary model more
realistic; for example, different sites in the protein are often allowed
to evolve at different rates by modeling rate heterogeneity using
the discrete-gamma model (denoted +G) [23].
The first step in maximum-likelihood phylogenetic analysis is
to choose the evolutionary model most appropriate for the given
sequence dataset, which can be achieved using programs such as
ProtTest [24] (see Note 8). ProtTest evaluates different models
using the Akaike information criterion (AIC), a metric that
accounts for the likelihood of the tree inferred using the model
(the goodness of fit) and the number of model parameters (to
penalize overfitting).
1. Open the ProtTest graphical interface and load the alignment
file.
2. Select “Compute likelihood scores” under the Analysis menu.
3. Select the substitution matrices to evaluate. At a minimum,
select the general-purpose JTT, WAG, and LG matrices; most
of the remaining matrices are based on specific types of proteins (for example, MtArt is based on sequence data from
invertebrate mitochondrial proteins).
4. Ensure that the +I, +G, +I+G, and empirical amino acid frequency options are selected to evaluate these modifications.
3.3 Model Selection
Ben E. Clifton et al.
Précédent

- 81/332

Suivant