predictions of antigen processing may be performed. In addition,
depending on the specific pathogen, its immunology and the target
community, one needs to consider the set of relevant HLA alleles
and the accuracy of corresponding predictors, which go from reasonably good to very poor depending on the amount of experimental data available for their training. The accuracy of the predictor for
every particular allele should be documented in the program.
When dealing with multiple strains in a pangenome approach,
the selection of genomes for subsequent analysis should pay careful
attention to data quality and the representativeness of the sample.
Strain selection should represent as much as possible the global
pathogen population. If strain selection is biased towards a particular lineage or genotype, this will lead to an overestimation of the
size of the core genome and an underestimation of antigen variability. The quality of the genome sequences is also critical. Low
quality genome assembly and annotation can lead to data loss
during comparative genomics analysis and a decrease in the proportion of core genes. To avoid the loss of an orthologous group due
to incomplete genome sequences in the sample, one option is to
make the core genome search less restrictive. For example, by
applying a cutoff value for gene prevalence in the sample close to
100% or simply defining core genes as those genes that are present
in all strains except one or two (soft core proteome). Recently, a
strategy based on the prevalence of essential genes to remove
genome sequences with poor or low coverage has been reported
[8]. The presence of paralogous gene families in the sample also
complicates the computational determination of the core genome
and hinders an accurate estimation of sequence variability and the
proportion of sites under selection. In addition, paralog genes
encoding proteins appear to play a role in DNA recombination
and antigenic variation [98–100]. All things considered, it is reasonable to filter out any cluster containing paralogous genes from
the list of candidate antigens.
Most antigen discovery strategies are aimed at the selection of
conserved epitopes in pathogen proteins. Theoretically, an ideal
vaccine should provide a broad coverage. From an evolutionary
perspective, conserved antigens are often assumed to be less immunogenic than highly variable ones [101]. In other words, immunodominant antigens are frequently also the most variable. Studies of
vaccine antigen diversity in natural populations have revealed
regions of proteins under selective pressure as immune-system
targets [14, 102, 103]. Conversely, other studies have identified
and validated highly conserved and protective bacterial vaccine
antigens [8, 104]. In light of this, sequence variability (as many
other properties) cannot be straightforwardly used as a criterion for
antigen prioritization but must be evaluated as per case. Nevertheless, extensive variability in a protein alignment should be a sufficient sign to filter out a protein during the selection process.
Bacterial Pan-Proteome-Based Antigen Discovery
57
Précédent

- 72/595

Suivant