cost-effective and straightforward approach. The
chloroplast genome of Potentilla micrantha was
the first one to be sequenced by using PacBio
long reads. The qualified reads after
error-correction represent 320-fold chloroplast
genome depth with a mean length of
1902 bp. A single contig was generated from
long-read assembly covering the entire genome
without any ambiguities. The chloroplast genome
assembled from PacBio long reads were consistent with that of Illumina short reads, whereas the
PacBio assembly was more continuous and
resolved 187 ambiguities existed in Illumina
assembly (Ferrarini et al. 2013). There was a
stronger positive correlation between the coverage and GC content in NGS, whereas Pacbio
long reads did not show any obvious GC-bias.
Still, TGS has its own limitation given its high
inherent error rate that used to require NGS reads
to correct (Ferrarini et al. 2013). With the falling
of sequencing cost and the random errors in
TGS, the sequences could be revised by a consensus deduced from deep sequencing coverage.
The accuracy could be achieved to 99.9% without any impact on the assembly. It is worth to
mention that Pacbio greatly facilitated the integration of inverted repeats (IRs). The unique
sequence from the junctions of IRs and the small
variation in IRs permitted them to be assigned
unambiguously, which were extremely challenging for NGS. The two IRs with the individual length of 25,530 bp in P. micrantha were
almost identical but with three nucleotide difference that was enough to correctly locate the IRs.
However, it was a non-trivial task for untangling
IRs from short-read assembly that required
tedious steps of manual operations (Ku et al.
2013; Ferrarini et al. 2013).
Another two chloroplast genomes from the
species Picea glauca (a gymnosperm) and Sinningia speciosa (a eudicot angiosperm) were also
assembled by taking advantage of long reads.
The pipeline of Organelle_PBA was specifically
designed to assemble chloroplast and mitochondrial genome by computationally selecting the
organelle long reads from total DNA sequencing
(Soorni et al. 2017). The simple-to-use program
and the long reads promoted the organelle
genome assembly performance and resolve the
inverted repeats. Furthermore, the application of
PacBio long reads would enhance our understanding of the complex structure and function of
chloroplast genome. It is believed that more plant
chloroplast genomes, including duckweeds, will
be sequenced by the long-read sequencing technology in near future.
10.4 Chloroplast Genome
Applications
10.4.1 Chloroplast Genome
Sequences for Plant
Barcode
Chloroplast DNA sequence data are a powerful
tool for plant identification and deciphering
genetic relationships among plant species. DNA
barcode is an efficient molecular identification
system to tell species apart by a universal marker.
The chloroplast-derived markers are the most
popular for identifying the genetic distances in
plants (Fig. 10.2). However, no single locus
could distinguish between all plant species. It is
even more challenging in duckweeds to perform
species identification and taxonomic studies due
to the highly reduced morphology and small
plant size (Wang et al. 2010).
The atpF-atpH chloroplast marker was proposed to be the best one with its high success of
PCR amplification and high power to discriminate duckweeds (Wang et al. 2010). To further
barcode sibling species that was lack of enough
sequence variation and polymorphism, the combination of two plastid sequences rpl16 and
rps16 recognized the species in Wolffia genus
(Landolt 1994; Bog et al. 2013). However, it
failed to delineate all 11 species of because of
highly closely inter- and intra-specific genetic
distances. The study of using two barcodes
(atpF-atpH and psbK-psbI spacer regions) could
distinguish 30 of the 37 duckweed species. The
increase of resolution from the whole chloroplast
genome as a single-locus DNA barcode becomes
the most feasible way to discriminate plant
species.
10 Duckweed Chloroplast Genome Sequencing and Annotation
111
chloroplast genome of Potentilla micrantha was
the first one to be sequenced by using PacBio
long reads. The qualified reads after
error-correction represent 320-fold chloroplast
genome depth with a mean length of
1902 bp. A single contig was generated from
long-read assembly covering the entire genome
without any ambiguities. The chloroplast genome
assembled from PacBio long reads were consistent with that of Illumina short reads, whereas the
PacBio assembly was more continuous and
resolved 187 ambiguities existed in Illumina
assembly (Ferrarini et al. 2013). There was a
stronger positive correlation between the coverage and GC content in NGS, whereas Pacbio
long reads did not show any obvious GC-bias.
Still, TGS has its own limitation given its high
inherent error rate that used to require NGS reads
to correct (Ferrarini et al. 2013). With the falling
of sequencing cost and the random errors in
TGS, the sequences could be revised by a consensus deduced from deep sequencing coverage.
The accuracy could be achieved to 99.9% without any impact on the assembly. It is worth to
mention that Pacbio greatly facilitated the integration of inverted repeats (IRs). The unique
sequence from the junctions of IRs and the small
variation in IRs permitted them to be assigned
unambiguously, which were extremely challenging for NGS. The two IRs with the individual length of 25,530 bp in P. micrantha were
almost identical but with three nucleotide difference that was enough to correctly locate the IRs.
However, it was a non-trivial task for untangling
IRs from short-read assembly that required
tedious steps of manual operations (Ku et al.
2013; Ferrarini et al. 2013).
Another two chloroplast genomes from the
species Picea glauca (a gymnosperm) and Sinningia speciosa (a eudicot angiosperm) were also
assembled by taking advantage of long reads.
The pipeline of Organelle_PBA was specifically
designed to assemble chloroplast and mitochondrial genome by computationally selecting the
organelle long reads from total DNA sequencing
(Soorni et al. 2017). The simple-to-use program
and the long reads promoted the organelle
genome assembly performance and resolve the
inverted repeats. Furthermore, the application of
PacBio long reads would enhance our understanding of the complex structure and function of
chloroplast genome. It is believed that more plant
chloroplast genomes, including duckweeds, will
be sequenced by the long-read sequencing technology in near future.
10.4 Chloroplast Genome
Applications
10.4.1 Chloroplast Genome
Sequences for Plant
Barcode
Chloroplast DNA sequence data are a powerful
tool for plant identification and deciphering
genetic relationships among plant species. DNA
barcode is an efficient molecular identification
system to tell species apart by a universal marker.
The chloroplast-derived markers are the most
popular for identifying the genetic distances in
plants (Fig. 10.2). However, no single locus
could distinguish between all plant species. It is
even more challenging in duckweeds to perform
species identification and taxonomic studies due
to the highly reduced morphology and small
plant size (Wang et al. 2010).
The atpF-atpH chloroplast marker was proposed to be the best one with its high success of
PCR amplification and high power to discriminate duckweeds (Wang et al. 2010). To further
barcode sibling species that was lack of enough
sequence variation and polymorphism, the combination of two plastid sequences rpl16 and
rps16 recognized the species in Wolffia genus
(Landolt 1994; Bog et al. 2013). However, it
failed to delineate all 11 species of because of
highly closely inter- and intra-specific genetic
distances. The study of using two barcodes
(atpF-atpH and psbK-psbI spacer regions) could
distinguish 30 of the 37 duckweed species. The
increase of resolution from the whole chloroplast
genome as a single-locus DNA barcode becomes
the most feasible way to discriminate plant
species.
10 Duckweed Chloroplast Genome Sequencing and Annotation
111
