3.4.2 HLA Selection
and Population Analysis
The repertoire of MHC molecules and their frequencies varies for
different human population groups; therefore, an adequate selection of the relevant HLA alleles should be made if a specific population is targeted. The Allele Frequency Net Database [76] can be
used to select the appropriate alleles in such cases. Otherwise,
IEDB provides two lists for MHC class I and II allele frequencies
and reference sets with maximal population coverage [77–79].
3.4.3 Structure-Based
Identification of Antibody
Binding Regions
B-cell epitopes are not necessarily lineal and continuous but may
include amino acids far from each other in the protein sequence.
There are no good data-driven predictors of B-cell epitopes, and
one generally needs to resort to structure-based methods predicting protein regions susceptible of binding to other proteins, for
example an antibody. Here, we propose two structure-based methods BEPPE [59] and EDP [60] that have been successfully used to
predict antibody epitopes in a series of studies on B. pseudomallei
antigens [80–83]. These methods use relatively inexpensive
approximations (compared to, for example, molecular dynamics
simulations [84]) to evaluate the internal energy distribution and
topology of the protein or its surface desolvation free energy and
predict from these quantities the regions more likely to bind an
antibody (or other protein). Since they require the structure of the
protein and a nonnegligible computation time, they cannot be used
in a high-throughput manner. However, they have proven very
helpful to validate potential candidates and eventually assist antigen
redesign.
3.5 Similarity
with Human Proteins
and Human
Microbiome Proteins
It has been suggested that linear B-cell epitopes found in pathogen
proteins do not share in general any sequence identity with human
proteins, and when they do it is always with proteins known to be
autoantigens [11]. In any case, clearly, one should avoid the selection of antigens including epitopes with any significant similarity to
a human sequence. Every potential antigen candidate should be
therefore screened against the whole human proteome and any
matching region annotated for later evaluation.
To avoid the selection of candidates that may lead to an
immune cross-reaction against bacteria of the human microbiome,
one may include in the initial subtractive analysis (see Subheading
3.2.1) at least one genome of those species (or serotypes within a
species) known to be phylogenetically close to the target species
(or serotype) and part of the human microbiome. For example,
when targeting Enteropathogenic E. coli one should include in the
analysis a number of genomes representative of commensal E. coli
strains and preselect proteins not common to the two groups.
Bacterial Pan-Proteome-Based Antigen Discovery
53
and Population Analysis
The repertoire of MHC molecules and their frequencies varies for
different human population groups; therefore, an adequate selection of the relevant HLA alleles should be made if a specific population is targeted. The Allele Frequency Net Database [76] can be
used to select the appropriate alleles in such cases. Otherwise,
IEDB provides two lists for MHC class I and II allele frequencies
and reference sets with maximal population coverage [77–79].
3.4.3 Structure-Based
Identification of Antibody
Binding Regions
B-cell epitopes are not necessarily lineal and continuous but may
include amino acids far from each other in the protein sequence.
There are no good data-driven predictors of B-cell epitopes, and
one generally needs to resort to structure-based methods predicting protein regions susceptible of binding to other proteins, for
example an antibody. Here, we propose two structure-based methods BEPPE [59] and EDP [60] that have been successfully used to
predict antibody epitopes in a series of studies on B. pseudomallei
antigens [80–83]. These methods use relatively inexpensive
approximations (compared to, for example, molecular dynamics
simulations [84]) to evaluate the internal energy distribution and
topology of the protein or its surface desolvation free energy and
predict from these quantities the regions more likely to bind an
antibody (or other protein). Since they require the structure of the
protein and a nonnegligible computation time, they cannot be used
in a high-throughput manner. However, they have proven very
helpful to validate potential candidates and eventually assist antigen
redesign.
3.5 Similarity
with Human Proteins
and Human
Microbiome Proteins
It has been suggested that linear B-cell epitopes found in pathogen
proteins do not share in general any sequence identity with human
proteins, and when they do it is always with proteins known to be
autoantigens [11]. In any case, clearly, one should avoid the selection of antigens including epitopes with any significant similarity to
a human sequence. Every potential antigen candidate should be
therefore screened against the whole human proteome and any
matching region annotated for later evaluation.
To avoid the selection of candidates that may lead to an
immune cross-reaction against bacteria of the human microbiome,
one may include in the initial subtractive analysis (see Subheading
3.2.1) at least one genome of those species (or serotypes within a
species) known to be phylogenetically close to the target species
(or serotype) and part of the human microbiome. For example,
when targeting Enteropathogenic E. coli one should include in the
analysis a number of genomes representative of commensal E. coli
strains and preselect proteins not common to the two groups.
Bacterial Pan-Proteome-Based Antigen Discovery
53
