The proportion of sites under selection in a sequence alignment
is a sign of how the evolutionary forces are acting on the sequence;
however, for a given gene only a small fraction of codon positions
will be subject to strong positive selection pressure. This may
happen, for instance, in pathogen epitopes where certain positively
selected codons may lead to escape variants able to evade the
immune system [30, 92, 93]. The most widely used methods for
the identification of sites under positive selection estimate the rates
of fixation of nonsynonymous and synonymous mutations at each
codon site based on maximum-likelihood substitution models.
Site-to-site variation in selection pressure across coding regions
can be analyzed with the popular software HyPhy [66] and
PAML [94]. The Codeml program of PAML and REL methods
in HyPhy constitute very similar approaches for site-to-site selection analysis [95]. These methods take phylogenies into account;
therefore, in order to conduct the analysis one needs a multiple
sequence alignment and a corresponding phylogenetic tree. Accurate phylogenies can be constructed using the DNAPARS program
(parsimony algorithm) in the PHYLIP package or the PhyML
program (maximum likelihood) [69]. However several other algorithms can be employed to do the same. Prior to the selection
analysis, recombination signals between sequences in the alignment
of orthologous sequences should be detected, because recombination events have a profound effect on evolutionary inferences and
on the identification of codon sites under positive selection
[96]. The GARD algorithm implemented in the HyPhy package
uses a genetic algorithm to screen multiple sequence alignments for
recombination and identify putative recombination breakpoints
[97]. Once recombination has been identified, it is necessary to
split the alignment into nonrecombinant sequence fragments and
the selection analyses can be run separately for each fragment
[95]. Figure 2 provides as an example a sequence conservation
analysis for two candidate proteins predicted as pathogen-specific
antigens in B. pseudomallei [16], pathogenic gram-negative bacteria responsible for melioidosis. This analysis showed that the protein BPSL1626 is highly conserved in this species [16], but the
protein BPSS1727 suffers extensive variability because it is subject
to recombination and selection pressure. The figure also shows a
pipeline proposal for genome-wide identification of positive selection and recombination in bacterial pangenomes.
3.7 Scripting
and Automatization
Most of the programs used in the antigen-identification protocol
presented here have a command line interface or are accessible as
web services and can be run in a noninteractive mode, facilitating
their inclusion in a pipeline for the automation of the entire process. It is thus possible to script the whole sequence of tasks so that
after the selection of species/strains no additional user intervention
is required. The output of such a pipeline will be a final list of
Bacterial Pan-Proteome-Based Antigen Discovery
55
is a sign of how the evolutionary forces are acting on the sequence;
however, for a given gene only a small fraction of codon positions
will be subject to strong positive selection pressure. This may
happen, for instance, in pathogen epitopes where certain positively
selected codons may lead to escape variants able to evade the
immune system [30, 92, 93]. The most widely used methods for
the identification of sites under positive selection estimate the rates
of fixation of nonsynonymous and synonymous mutations at each
codon site based on maximum-likelihood substitution models.
Site-to-site variation in selection pressure across coding regions
can be analyzed with the popular software HyPhy [66] and
PAML [94]. The Codeml program of PAML and REL methods
in HyPhy constitute very similar approaches for site-to-site selection analysis [95]. These methods take phylogenies into account;
therefore, in order to conduct the analysis one needs a multiple
sequence alignment and a corresponding phylogenetic tree. Accurate phylogenies can be constructed using the DNAPARS program
(parsimony algorithm) in the PHYLIP package or the PhyML
program (maximum likelihood) [69]. However several other algorithms can be employed to do the same. Prior to the selection
analysis, recombination signals between sequences in the alignment
of orthologous sequences should be detected, because recombination events have a profound effect on evolutionary inferences and
on the identification of codon sites under positive selection
[96]. The GARD algorithm implemented in the HyPhy package
uses a genetic algorithm to screen multiple sequence alignments for
recombination and identify putative recombination breakpoints
[97]. Once recombination has been identified, it is necessary to
split the alignment into nonrecombinant sequence fragments and
the selection analyses can be run separately for each fragment
[95]. Figure 2 provides as an example a sequence conservation
analysis for two candidate proteins predicted as pathogen-specific
antigens in B. pseudomallei [16], pathogenic gram-negative bacteria responsible for melioidosis. This analysis showed that the protein BPSL1626 is highly conserved in this species [16], but the
protein BPSS1727 suffers extensive variability because it is subject
to recombination and selection pressure. The figure also shows a
pipeline proposal for genome-wide identification of positive selection and recombination in bacterial pangenomes.
3.7 Scripting
and Automatization
Most of the programs used in the antigen-identification protocol
presented here have a command line interface or are accessible as
web services and can be run in a noninteractive mode, facilitating
their inclusion in a pipeline for the automation of the entire process. It is thus possible to script the whole sequence of tasks so that
after the selection of species/strains no additional user intervention
is required. The output of such a pipeline will be a final list of
Bacterial Pan-Proteome-Based Antigen Discovery
55
