3.6 Antigen
Variability and Epitope
Conservation Among
Different Genomes
As already mentioned, an ideal vaccine would be based on conserved antigens eliciting a protective memory immune response. In
this regard, antigen variability resulting from bacterial genetic
diversity is a major obstacle to vaccine broad-spectrum efficacy
[85]. The availability of multiple genomic sequences from different
strains of the same pathogenic bacterial species greatly facilitates the
analysis of sequence variability in vaccine candidate encoding genes.
Sequence variability in multiple alignments of ortholog protein
sequences may be calculated as a simple percentage of the number
of variable positions in the alignment or, better, as the average
amino acid identity (overall uncorrected p-distance). However,
when variation among sequences is large and the alignment algorithm inserts gaps, these simple methods will inaccurately estimate
antigen diversity in a bacterial population. To identify conserved
epitopes, the gaps are often removed and the ungapped multiple
alignments are used for further analysis. Distance methods for the
analysis of multiple sequence alignments can be found in many
open source software and freely available packages (see Table 1).
The analysis of amino acid diversity can be also done site-by-site by
estimating the absolute site variability within a multiple sequence
alignment. PVS (Protein Variability Server) is a web-based tool that
provides absolute sequence variability estimates per site in multiple
sequence alignments in terms of Shannon entropy, Simpson diversity index and Wu-Kabat variability coefficient [63]. Shannon
entropy has been used to assess the sequence variation of viral
proteomes in reverse vaccinology approaches [86].
The study of antigenic variability can be also approached from
phylogenetics, since the large diversity found in bacterial pathogens
has partly emerged as an evolutionary strategy to evade host immunity. The ability of bacteria to constantly and rapidly evolve by
natural selection via recombination and mutation has shaped the
high diversity found in protein antigenic regions [87]. Genomewide evidence for positive selection and recombination in bacterial
genomes has demonstrated that genes coding for putative antigens
and virulence factors are prone to natural selective pressure [88–
91]. Nonsynonymous mutations are translated into differences at
protein level and can directly affect protein function and their
recognition by the immune system receptors. Therefore, the proportion of synonymous and nonsynonymous differences can be
used to test the action of natural selection on protein-coding
genes and it has been widely used as an indicator of positive
(Darwinian) selection. DnaSP (DNA Sequence Polymorphism)
and MEGA (Molecular Evolutionary Genetics Analysis) are freely
available software packages for the analysis of sequence polymorphisms using multiple sequence alignments [62, 64]. They can estimate sequence diversity and several measures of DNA sequence
variation within populations in synonymous or nonsynonymous
sites.
54
Daniel Yero et al.
Variability and Epitope
Conservation Among
Different Genomes
As already mentioned, an ideal vaccine would be based on conserved antigens eliciting a protective memory immune response. In
this regard, antigen variability resulting from bacterial genetic
diversity is a major obstacle to vaccine broad-spectrum efficacy
[85]. The availability of multiple genomic sequences from different
strains of the same pathogenic bacterial species greatly facilitates the
analysis of sequence variability in vaccine candidate encoding genes.
Sequence variability in multiple alignments of ortholog protein
sequences may be calculated as a simple percentage of the number
of variable positions in the alignment or, better, as the average
amino acid identity (overall uncorrected p-distance). However,
when variation among sequences is large and the alignment algorithm inserts gaps, these simple methods will inaccurately estimate
antigen diversity in a bacterial population. To identify conserved
epitopes, the gaps are often removed and the ungapped multiple
alignments are used for further analysis. Distance methods for the
analysis of multiple sequence alignments can be found in many
open source software and freely available packages (see Table 1).
The analysis of amino acid diversity can be also done site-by-site by
estimating the absolute site variability within a multiple sequence
alignment. PVS (Protein Variability Server) is a web-based tool that
provides absolute sequence variability estimates per site in multiple
sequence alignments in terms of Shannon entropy, Simpson diversity index and Wu-Kabat variability coefficient [63]. Shannon
entropy has been used to assess the sequence variation of viral
proteomes in reverse vaccinology approaches [86].
The study of antigenic variability can be also approached from
phylogenetics, since the large diversity found in bacterial pathogens
has partly emerged as an evolutionary strategy to evade host immunity. The ability of bacteria to constantly and rapidly evolve by
natural selection via recombination and mutation has shaped the
high diversity found in protein antigenic regions [87]. Genomewide evidence for positive selection and recombination in bacterial
genomes has demonstrated that genes coding for putative antigens
and virulence factors are prone to natural selective pressure [88–
91]. Nonsynonymous mutations are translated into differences at
protein level and can directly affect protein function and their
recognition by the immune system receptors. Therefore, the proportion of synonymous and nonsynonymous differences can be
used to test the action of natural selection on protein-coding
genes and it has been widely used as an indicator of positive
(Darwinian) selection. DnaSP (DNA Sequence Polymorphism)
and MEGA (Molecular Evolutionary Genetics Analysis) are freely
available software packages for the analysis of sequence polymorphisms using multiple sequence alignments [62, 64]. They can estimate sequence diversity and several measures of DNA sequence
variation within populations in synonymous or nonsynonymous
sites.
54
Daniel Yero et al.
