efficient sequencing approach yields an average
sequence depth of 55Â to 186Â and produced
high-quality draft genomes with an estimation of
88–94% completion (Cronn et al. 2008). Three
Brassica rapa accessions sequenced by Illumina
short reads also generated complete chloroplast
genomes from total DNA without any experimental purification (Wu et al. 2012).
Three more duckweed chloroplast genomes (S.
polyrhiza, W. lingulata, and W. australiana) were
sequenced by using SOLiD platform with a read
length of 50 bp. The short reads derived from
cpDNA were filtered electronically using the
reference of L. minor. The genomes were de novo
assembled using the SOLiD System de novo
Accessory Tools 2.0 (http://solidsoftwaretools.
com/gf/project/denovo/) in conjunction with the
velvet assembly engine (Zerbino and Birney
2008). The remaining gaps were closed with
Sanger sequencing of 13–29 additional PCR
products. The chloroplast genome sizes of duckweeds had a range of 165,955–169,353 bp. They
contained two copies of *31-Kb of inverted
repeats, separated by a *90-Kb LSC and a
*10-Kb SSC (Table 10.2). The borders between
IR and the single-copy region were variable in
duckweeds. The accessibility of duckweed
chloroplast genomes and the sequence divergence
indicated their taxonomic and phylogenetic relationships, which would provide the references for
further species identification and chloroplast
genetic engineering.
The overall structure of the duckweed
chloroplast genomes was conserved. They had
similar gene copy number and gene order. Still,
there were numbers of rearrangements and
sequences polymorphisms, like INDELs
(insertion/deletion) and single-nucleotide variations. The sequence alignment and comparison
revealed multiple variation hotspots, like the
10-Kb regions between rpoB and psbD in the
genomes. The rich sequence variation occurred
in the noncoding intergenic regions, while IR
regions showed lower sequence divergence than
the single-copy regions. There was a 505-bp
deletion in S. polyrhiza compared to the 100-bp
deletion in W. lingulata. It was also found a
353-bp insertion occurred at 31 Kb of the intergenic petN-psbM region of W. Australiana
(Wang and Messing 2011).
The genomes were annotated by DOGMA
(Wyman et al. 2004). There were 83–85
protein-coding genes in three duckweeds, 37 for
rRNA genes and 8 for tRNA (Table 10.2).
Generally, the chloroplast genome was conserved in gene number and organization with
respect to the reference genome of L. minor.
10.3.3 Other Chloroplast Genomes
Sequenced by TGS
Technology
Given the high-throughput and the increase of
computational capacity for NGS, the number of
released chloroplast genome will double in a very
short period and accelerate the development of the
chloroplast genomes. However, the reads produced by NGS are relatively short (*150 bp),
leading to the fragmented assembly and requiring
the inevitable steps of gap filling with PCR
amplification and Sanger sequencing (Dohm et al.
2008). The technology of third-generation
sequencing (TGS), also called single-molecule
real-time sequencing (SMRT) was invented by
Pacific Biosciences in 2009 (Eid et al. 2009; Koren
et al. 2013). The way of sequencing-by-synthesis
could produce long reads and cover the full template without amplification bias. The current
average read length generated by PacBio Sequel is
exceptionally long (>20 Kb), spanning over the
repeat regions and increasing genome contiguity
(Ardui et al. 2018). It has been widely used in plant
sciences, even for large and complex nuclear
genomes (Li et al. 2017). Several genomes in terms
of continuity and accuracy showed improved
quality, including U. gibba (82 Mb) (Lan et al.
2017), O. thomaeum (245 Mb) (VanBuren et al.
2015), C. quinoa (1500 Mb) (Jarvis et al. 2017),
Zea mays (2300 Mb) (Jiao et al. 2017), and H.
annuus (3000 Mb) (Badouin et al. 2017; Li et al.
2017).
The revolutionary PacBio long reads also
provide the chloroplast genome sequencing a
110
Y. Zhang and W. Wang
sequence depth of 55Â to 186Â and produced
high-quality draft genomes with an estimation of
88–94% completion (Cronn et al. 2008). Three
Brassica rapa accessions sequenced by Illumina
short reads also generated complete chloroplast
genomes from total DNA without any experimental purification (Wu et al. 2012).
Three more duckweed chloroplast genomes (S.
polyrhiza, W. lingulata, and W. australiana) were
sequenced by using SOLiD platform with a read
length of 50 bp. The short reads derived from
cpDNA were filtered electronically using the
reference of L. minor. The genomes were de novo
assembled using the SOLiD System de novo
Accessory Tools 2.0 (http://solidsoftwaretools.
com/gf/project/denovo/) in conjunction with the
velvet assembly engine (Zerbino and Birney
2008). The remaining gaps were closed with
Sanger sequencing of 13–29 additional PCR
products. The chloroplast genome sizes of duckweeds had a range of 165,955–169,353 bp. They
contained two copies of *31-Kb of inverted
repeats, separated by a *90-Kb LSC and a
*10-Kb SSC (Table 10.2). The borders between
IR and the single-copy region were variable in
duckweeds. The accessibility of duckweed
chloroplast genomes and the sequence divergence
indicated their taxonomic and phylogenetic relationships, which would provide the references for
further species identification and chloroplast
genetic engineering.
The overall structure of the duckweed
chloroplast genomes was conserved. They had
similar gene copy number and gene order. Still,
there were numbers of rearrangements and
sequences polymorphisms, like INDELs
(insertion/deletion) and single-nucleotide variations. The sequence alignment and comparison
revealed multiple variation hotspots, like the
10-Kb regions between rpoB and psbD in the
genomes. The rich sequence variation occurred
in the noncoding intergenic regions, while IR
regions showed lower sequence divergence than
the single-copy regions. There was a 505-bp
deletion in S. polyrhiza compared to the 100-bp
deletion in W. lingulata. It was also found a
353-bp insertion occurred at 31 Kb of the intergenic petN-psbM region of W. Australiana
(Wang and Messing 2011).
The genomes were annotated by DOGMA
(Wyman et al. 2004). There were 83–85
protein-coding genes in three duckweeds, 37 for
rRNA genes and 8 for tRNA (Table 10.2).
Generally, the chloroplast genome was conserved in gene number and organization with
respect to the reference genome of L. minor.
10.3.3 Other Chloroplast Genomes
Sequenced by TGS
Technology
Given the high-throughput and the increase of
computational capacity for NGS, the number of
released chloroplast genome will double in a very
short period and accelerate the development of the
chloroplast genomes. However, the reads produced by NGS are relatively short (*150 bp),
leading to the fragmented assembly and requiring
the inevitable steps of gap filling with PCR
amplification and Sanger sequencing (Dohm et al.
2008). The technology of third-generation
sequencing (TGS), also called single-molecule
real-time sequencing (SMRT) was invented by
Pacific Biosciences in 2009 (Eid et al. 2009; Koren
et al. 2013). The way of sequencing-by-synthesis
could produce long reads and cover the full template without amplification bias. The current
average read length generated by PacBio Sequel is
exceptionally long (>20 Kb), spanning over the
repeat regions and increasing genome contiguity
(Ardui et al. 2018). It has been widely used in plant
sciences, even for large and complex nuclear
genomes (Li et al. 2017). Several genomes in terms
of continuity and accuracy showed improved
quality, including U. gibba (82 Mb) (Lan et al.
2017), O. thomaeum (245 Mb) (VanBuren et al.
2015), C. quinoa (1500 Mb) (Jarvis et al. 2017),
Zea mays (2300 Mb) (Jiao et al. 2017), and H.
annuus (3000 Mb) (Badouin et al. 2017; Li et al.
2017).
The revolutionary PacBio long reads also
provide the chloroplast genome sequencing a
110
Y. Zhang and W. Wang
