Table 1
Bioinformatics tools, web servers, and databases relevant to the antigen discovery pipeline described
herein
Tool [reference]
URL (http or ftp)
Description
CLI
run
a
3.1. Selection of species and strains and acquisition of genomic sequences
NCBI genomes
[41]
https://www.ncbi.nlm.nih.
gov/genome
NCBI complete genomes database
Yes
ENA assembly
search portal
[42]
https://www.ebi.ac.uk/ena/
data/warehouse/search?
portal¼assembly
ENA complete genomes database accession
portal
Yes
PATRIC [43]
https://www.patricbrc.org
Pathosystems Resource Integration Center,
provides integrated data and analysis tools
to support biomedical research on
bacterial infectious diseases
Yes
3.2. Comparative analysis
CD-HIT [44]
http://weizhongli-lab.org/
cd-hit/
Sequence clustering by similarity, based on
all vs. all sequence comparison and an
arbitrary cutoff
Yes
OrthoMCL [45]
https://orthomcl.org/
orthomcl/
MCL clustering of blast searches
Yes
OMA [46]
http://omabrowser.org/
standalone
Score-based clique selection from graphs
constructed by reciprocal best hit of
all vs. all genome/proteome comparisons
Yes
Complete
Reciprocal Best
Hit [47]
Selection of complete graph constructed by
reciprocal best hit of all vs. all genome/
proteome comparisons
Yes
3.3.1. Annotations: Protein function
UniProt [48]
https://www.uniprot.org
Main site for protein sequence and functional
information
Yes
NCBI gb file
[41, 49]
https://www.ncbi.nlm.nih.
gov/genome
Complete genome annotated file used as
source of all required data
Yes
3.3.2. Annotations: Subcellular localization
PSORTb [50]
https://www.psort.org/
psortb/
Bacterial cellular localization prediction tool Yes
3.3.3. Annotations: Protein solubility related data
TMHMM [51]
http://www.cbs.dtu.dk/
services/TMHMM/
Prediction of transmembrane helices
Yes
FoldIndex [52]
https://fold.weizmann.ac.il/
fldbin/findex
Prediction of unfolded/unstructured
regions
Yes
Aggrescan [53]
http://bioinf.uab.es/
aggrescan/
Prediction of aggregation hot spots
Yes
(continued)
48
Daniel Yero et al.
Bioinformatics tools, web servers, and databases relevant to the antigen discovery pipeline described
herein
Tool [reference]
URL (http or ftp)
Description
CLI
run
a
3.1. Selection of species and strains and acquisition of genomic sequences
NCBI genomes
[41]
https://www.ncbi.nlm.nih.
gov/genome
NCBI complete genomes database
Yes
ENA assembly
search portal
[42]
https://www.ebi.ac.uk/ena/
data/warehouse/search?
portal¼assembly
ENA complete genomes database accession
portal
Yes
PATRIC [43]
https://www.patricbrc.org
Pathosystems Resource Integration Center,
provides integrated data and analysis tools
to support biomedical research on
bacterial infectious diseases
Yes
3.2. Comparative analysis
CD-HIT [44]
http://weizhongli-lab.org/
cd-hit/
Sequence clustering by similarity, based on
all vs. all sequence comparison and an
arbitrary cutoff
Yes
OrthoMCL [45]
https://orthomcl.org/
orthomcl/
MCL clustering of blast searches
Yes
OMA [46]
http://omabrowser.org/
standalone
Score-based clique selection from graphs
constructed by reciprocal best hit of
all vs. all genome/proteome comparisons
Yes
Complete
Reciprocal Best
Hit [47]
Selection of complete graph constructed by
reciprocal best hit of all vs. all genome/
proteome comparisons
Yes
3.3.1. Annotations: Protein function
UniProt [48]
https://www.uniprot.org
Main site for protein sequence and functional
information
Yes
NCBI gb file
[41, 49]
https://www.ncbi.nlm.nih.
gov/genome
Complete genome annotated file used as
source of all required data
Yes
3.3.2. Annotations: Subcellular localization
PSORTb [50]
https://www.psort.org/
psortb/
Bacterial cellular localization prediction tool Yes
3.3.3. Annotations: Protein solubility related data
TMHMM [51]
http://www.cbs.dtu.dk/
services/TMHMM/
Prediction of transmembrane helices
Yes
FoldIndex [52]
https://fold.weizmann.ac.il/
fldbin/findex
Prediction of unfolded/unstructured
regions
Yes
Aggrescan [53]
http://bioinf.uab.es/
aggrescan/
Prediction of aggregation hot spots
Yes
(continued)
48
Daniel Yero et al.
