species. Therefore, in this latter case, DNA/DNA
hybridization between the genomes of two strains is required
(Fig. 6.6). The lack of equivalence agreed between the two
techniques lies in the fact that 16S rDNA corresponds to one
locus among the thousands of loci present in a genome.
Therefore, it is recommended to use a multi-locus approach
such as AFLP (amplified fragment length polymorphism)
and correlate the results with data from DNA/DNA
hybridization (Stackebrandt et al. 2002) (Fig. 6.7).
Phylogenetic analysis of molecular markers* (e.g., 16S
rRNA) is a multistep process (Fig. 6.8), the first of which is to
compare the sequence of interest to investigate whether
homologous sequences are present in sequence public
databases (e.g., general databases such as the nr database at
the NCBI or more specialized databases such as the RDP II or
silva in the case of 16S rRNA). This initial research is based
on sequence similarity comparisons among the sequences of
the database and the studied sequence. Indeed, the closer
related the two strains are, the closer the sequences of their
genes (including 16S rRNA) will be. One of the most frequently used software to perform these searches is the BLAST
(Basic Local Alignment Search Tool). It helps to identify
similar regions between two protein or nucleic acid sequences
(Fig. 6.8a). The identification of most similar sequences
provides only a crude taxonomic indication (genus level or
higher). To get a finer taxonomic affiliation, it is thus important to perform a phylogenetic analysis. For the phylogenetic
analysis, all homologous sequences from organisms close to
the one analyzed are extracted from the database and aligned
using multiple alignment software such as Clustal omega or
Muscle (Fig. 6.8b, c). The resulting alignments are trimmed in
order to remove regions where homology between sites is
ambiguous (i.e., poorly conserved regions containing
insertions/deletions that cannot be unambiguously positioned). The remaining sites are used to infer the phylogenetic
trees (Fig. 6.8d). One of the tree reconstruction methods
commonly used in taxonomy is the neighbor-joining (Saitou
and Nei 1987) because of its simplicity and speed. The
neighbor-joining method belongs to the family of distance
methods because the first step of neighbor-joining requires the
quantification of evolutionary distances among all pairs of
sequences studied. This estimate of the evolutionary distances
is based upon the use of an evolutionary model. The calculation of evolutionary distances between each pair of sequences
allows the establishment of a pairwise distance matrix
(Fig. 6.8e). These distances are then represented as phylogenetic trees where branch lengths are proportional to evolutionary distances among sequences (Fig. 6.8f). One of the
advantages of the neighbor-joining is the ability to process a
large number of sequences in a very short time (few seconds
CE2105 (Rps)
IO2102 (Rdv/Rba)
IO2103 (Rdv/Rba)
CE2104 (Rdv/Rba)
EP2104 (Rbi/Rps)
EP2101 (Rdv/Rba)
EP2105 (Rbi/Rps)
CE2102 (Rbi/Rps)
CE2401 (Prs)
CE2208 (Chr)
CE2206 (Chr)
CE2203 (Tcs)
CE2209 (Tca)
EP2208 (Tca)
EP2204 (Tca)
EP2202 (Tca)
EP2205 (Chr)
EP2201 (Chr)
IO2201 (Chr)
IO2205 (Chr)
EP2209 (Chr)
IO2204 (Chr)
IO2203 (Chr)
CE2201 (Chr)
CE2205 (Chr)
c
Fig. 6.5 (continued)
6 Taxonomy and Phylogeny of Prokaryotes
157
hybridization between the genomes of two strains is required
(Fig. 6.6). The lack of equivalence agreed between the two
techniques lies in the fact that 16S rDNA corresponds to one
locus among the thousands of loci present in a genome.
Therefore, it is recommended to use a multi-locus approach
such as AFLP (amplified fragment length polymorphism)
and correlate the results with data from DNA/DNA
hybridization (Stackebrandt et al. 2002) (Fig. 6.7).
Phylogenetic analysis of molecular markers* (e.g., 16S
rRNA) is a multistep process (Fig. 6.8), the first of which is to
compare the sequence of interest to investigate whether
homologous sequences are present in sequence public
databases (e.g., general databases such as the nr database at
the NCBI or more specialized databases such as the RDP II or
silva in the case of 16S rRNA). This initial research is based
on sequence similarity comparisons among the sequences of
the database and the studied sequence. Indeed, the closer
related the two strains are, the closer the sequences of their
genes (including 16S rRNA) will be. One of the most frequently used software to perform these searches is the BLAST
(Basic Local Alignment Search Tool). It helps to identify
similar regions between two protein or nucleic acid sequences
(Fig. 6.8a). The identification of most similar sequences
provides only a crude taxonomic indication (genus level or
higher). To get a finer taxonomic affiliation, it is thus important to perform a phylogenetic analysis. For the phylogenetic
analysis, all homologous sequences from organisms close to
the one analyzed are extracted from the database and aligned
using multiple alignment software such as Clustal omega or
Muscle (Fig. 6.8b, c). The resulting alignments are trimmed in
order to remove regions where homology between sites is
ambiguous (i.e., poorly conserved regions containing
insertions/deletions that cannot be unambiguously positioned). The remaining sites are used to infer the phylogenetic
trees (Fig. 6.8d). One of the tree reconstruction methods
commonly used in taxonomy is the neighbor-joining (Saitou
and Nei 1987) because of its simplicity and speed. The
neighbor-joining method belongs to the family of distance
methods because the first step of neighbor-joining requires the
quantification of evolutionary distances among all pairs of
sequences studied. This estimate of the evolutionary distances
is based upon the use of an evolutionary model. The calculation of evolutionary distances between each pair of sequences
allows the establishment of a pairwise distance matrix
(Fig. 6.8e). These distances are then represented as phylogenetic trees where branch lengths are proportional to evolutionary distances among sequences (Fig. 6.8f). One of the
advantages of the neighbor-joining is the ability to process a
large number of sequences in a very short time (few seconds
CE2105 (Rps)
IO2102 (Rdv/Rba)
IO2103 (Rdv/Rba)
CE2104 (Rdv/Rba)
EP2104 (Rbi/Rps)
EP2101 (Rdv/Rba)
EP2105 (Rbi/Rps)
CE2102 (Rbi/Rps)
CE2401 (Prs)
CE2208 (Chr)
CE2206 (Chr)
CE2203 (Tcs)
CE2209 (Tca)
EP2208 (Tca)
EP2204 (Tca)
EP2202 (Tca)
EP2205 (Chr)
EP2201 (Chr)
IO2201 (Chr)
IO2205 (Chr)
EP2209 (Chr)
IO2204 (Chr)
IO2203 (Chr)
CE2201 (Chr)
CE2205 (Chr)
c
Fig. 6.5 (continued)
6 Taxonomy and Phylogeny of Prokaryotes
157
