6.1.2 PacBio and Oxford Nanopore
Sequencing Providing
an Opportunity to Crack
the Complex Duckweed
Genomes
PacBio single-molecule real-time (SMRT)
sequencing is the most popular TGS technology
in the market. Different with NGS, PacBio
sequencing simulates the natural processes of
DNA replication, enables real-time sequencing of
DNA molecules through zero-mode waveguide
pores (ZMWs) and phosphorylated nucleotides,
and does not require a pause between read steps,
and each step of template amplification will
generate a light pulse that can be identified as a
different labeled nucleotide (Schadt et al. 2010).
The sequencing technology can generate extremely long reads with an average read length of
more than 15 Kb. Some reads can reach up to
100 Kb that is comparable to a BAC clone
(Pacific 2018a). However, the main concern is
the high random sequencing errors in the long
noisy reads. The latest sequencing platform for
PacBio is Sequel System, which is featured with
one million ZMWs per SMRT cell instead of
150,000 in RS II, and therefore, produces seven
times more reads per SMRT cell than RS II
(Pacific 2018b). The increased throughput and
the decreased cost make PacBio SMRT
sequencing technology available to any individual laboratory. Given long-read lengths and
GC-free preference, PacBio reads allow assemblers to span repeat regions. It has been used in
de novo assembling multiple plant genomes and
dramatically improved the genome reconstruction (VanBuren et al. 2015; Jiao et al. 2017b; Lan
et al. 2017).
The recently published review has summarized the broad applications of PacBio SMRT
sequencing technology in genome sequencing, as
well as the comparisons with the next-generation
sequencing (Li et al. 2018). It was found that
PacBio long reads can assemble unprecedented
contiguity genomes and more complete high
repetitive regions, such as LTR retrotransposon,
centromeres, and telomeric repeats (Li et al.
2018). The latest study reported that using
single-molecule real-time sequencing and a
meta-assembly approach obtained one of the
most comprehensive plant genome of Rosa chinensis that had a contig N50 of 24 Mb. Therefore, PacBio SMRT sequencing has shown a
great promise to solve the complex duckweed
genomes. But, there are still no released duckweed genomes sequenced by PacBio reads.
However, there was a pioneer work done for
Lemna minor 8627. The additional input of
long-read sequencing significantly increased the
contiguity of genome assembly that the contig
N50 was extended from 65 to 222 Kb (Ernst
2016).
Oxford Nanopore is another long-read
sequencing platform and has similar characteristics as PacBio long reads, which can produce
reads up to hundreds of kilobases but with a high
error rate. Nanopore sequencing determines
DNA sequences by the ionic current changes
when DNA strands pass through the tiny Nanopores in the flow cell (Li et al. 2018). Thus, it can
produce ultra-long-read lengths the same as the
DNA molecule lengths. The recent study showed
that Nanopore sequencing can produce sequence
reads up to one Mb (Willing et al. 2015), which
will effectively solve the duckweeds genomes
with the highly repetitive region, like Wolffiella
and Wolffia. Compared to PacBio SMRT, Oxford
Nanopore sequencing technology can produce
ultra-long reads which are more productive to
assemble high-continuity genomes, even though
it shows overall lower data quality (Weirather
et al. 2017; Jain et al. 2018). However, upon now
there are few reports about Nanopore genome
sequencing plant genome.
6.1.3 Bionano, 10X Genomics,
and Hi-C for Scaffolding
Duckweed Genomes
Due to the redundant sequence and complexity of
the plant genome, it is almost impossible to
assemble the genome only by sequencing reads.
After obtaining contigs from sequence assembly,
they are often ordered and oriented into scaffolds
by using large fragment libraries like BAC, MP,
6 Strategies and Tools for Sequencing Duckweeds
69
Précédent

- 83/191

Suivant