total number of organisms used (core-proteome) to one (genomespecific protein). This can be done with several programs, of which
we propose three in Table 1 that are representative for different
methods used in available software [44–46]. We recommend a
complete-graph-based reciprocal-best-hit method (e.g., [47])
since, in addition to determining the orthologous groups, it will
facilitate the removal of paralogs (the presence of paralogs is hinted
by the generation of noncomplete graphs) and the identification of
proteins truly specific of a subgroup of genomes (see below).
3.2.1 Subtractive
Pangenome Analysis
When the objective is to select orthologous groups made of proteins present only in a defined subgroup of genomes (e.g., pathogens), only complete graph–based methods are recommended.
Due to the way each method determines the orthologous groups,
neither arbitrary similarity thresholds nor Markov clustering
(MCL) nor similarity-based clique methods, where similar proteins
may end up in different clusters, will guarantee the “non-presence”
of similar proteins in other organisms. When using any of the latter
three methods, one may have the temptation to calculate the core
proteome for the whole panel of strains, then calculate the core
proteome for a selected subgroup and finally subtract the first result
from the second. Note that when doing this, all proteins in the
resulting list will be in all members of the selected subgroup and
will not be in all members of the full panel (as intended) but may
still be present in some members of the full panel that do not
belong to the selected subgroup.
3.3 Annotations
All available data for each protein has to be taken into account. This
data will be crucial when the final selection is made, not only to
decide if the protein is a good vaccine candidate but also to determine if, with ease, it can be produced in vitro and purified.
3.3.1 Protein Function
It may be extracted from different sources, but since UniProt [48]
contains almost all available data for each protein, it is our primary
source for annotations. From here, Gene Ontology (GO) codes
[71], Pfam domains [72], and Interpro data [73] can be assigned to
each orthologous group. Lately, many proteins have been removed
from UniProt due to redundancy issues, but it is highly probable
that at least one protein per orthologous group remains in the
database, and its annotated data may then be made extensive to
all other members of the orthologous group. If this were not the
case, some functional data can be extracted from the initially downloaded gene bank file.
3.3.2 Subcellular
Localization
The analysis of subcellular localization has been central to reverse
vaccinology since its conception [6]. It is important to select proteins that are accessible to antibodies on the pathogen’s surface if
one is looking for humoral immunity. Although modern vaccine
Bacterial Pan-Proteome-Based Antigen Discovery
51
Précédent

- 66/595

Suivant