are no unplaced scaffolds); Chromosome (a chromosome is present
but can have gaps, unplaced scaffolds or contigs); Scaffold (some
contigs are joined); and Contigs (assembled sequences are unconnected and unlocalized). Thus, the number of genes and coding
sequences (CDSs) of different genomes of the same species can
show artifactual differences, making the selection of proteins present in all strains of the species (core proteome) or specific for a
selected group practically impossible. To circumvent this problem,
the strains of a species with CDS number and genome size deviating largely from those of strains with “Complete” status need to be
removed from the initial panel of strains. At the same time, a
maximum allowed number of contigs and a minimum N50 should
be used for strains with “Contigs” or “Scaffold” status, to ensure a
minimum quality of the genomes used.
How this step is completed will highly depend on the source
(s) for the genome sequences. Although many databases exist, in
Table 1 we list the three main ones. We recommend NCBI genomes [41] since it usually contains the largest number of species/
strains and one can retrieve data in an unattended manner (with
scripts) following the same steps that one would perform interactively through a web browser, thus allowing the combination of
both retrieval options.
3.2 Comparative
Analysis
In the previous step all selected genomes were downloaded, preferably in a full gene bank format. Corresponding files will contain the
full genome sequence with all annotated genes and their associated
features, in a way that is easy to manually visualize and automatically
extract for later processing (i.e., gene and protein sequences and
annotated data such as protein function or their database links).
Comparative analysis is then done at the CDS level, determining the protein orthologous groups with sizes ranging from the
Table 1
(continued)
Tool [reference]
URL (http or ftp)
Description
CLI
run
a
Datamonkey 2.0
[67]
https://www.datamonkey.
org/
A web interface to run HyPhy standard
analyses
No
PHYLIP package
[68]
http://evolution.genetics.
washington.edu/phylip.
html
Phylogenies (evolutionary trees) by the
parsimony algorithm
Yes
PhyML [69]
http://www.atgcmontpellier.fr/phyml/
Phylogenies (evolutionary trees) based on
the maximum-likelihood principle
Yes
a
The software can be run through a Command Line Interface (CLI) and can be therefore integrated in a computational
pipeline for the automation of the entire process
50
Daniel Yero et al.
Précédent

- 65/595

Suivant