ALLPATHS-LG and SSPACE (Boetzer et al.
2011).
6.2.2 De Novo Assembling Tools
for PacBio or Nanopore
Long Reads
There are still no released duckweed genomes
sequenced by the third-generation sequencing of
PacBio or Nanopore platforms that generate long
reads. Here, we recommend four popular de novo
assembling tools for long reads, and expect they
can benefit the assembly of duckweed genomes
in the near future.
Certain alignment algorithms have been
developed to effectively discover overlaps among
noisy long reads, such as DALIGNER (Myers
2014), BLASR (Chaisson and Tesler 2012),
MHAP (Berlin et al. 2015), GraphMap (Sovic
et al. 2016), and Minimap (Li 2016).
DALIGNER was the first tool designed specifically for finding the overlaps between noisy long
reads with the filtering or removals of k-mers
redundancy. The program could increase the
computation speed, decrease memory usage, and
mitigate the effect of repetitive sequences. Given
the risk of filtering important k-mers, the
parameters (-k, -w, -h, -t) need to be customized.
Another method to improve the alignment efficiency is to split the large data set into small
blocks based on the total number of base pairs
and read lengths by using the DBsplit utility
(Myers 2014).
FALCON is an overlap–layout–consensus
(OLC) genome assembler based on DALIGNER,
which only supports PacBio data. The pipeline
includes six steps to construct contigs, while the
correction and polish of raw reads are the steps that
consume the most computational resources. The
script “fc_run.py” with a configuration file can
complete the whole assembly process. The configuration file contains input files, optimal
parameters, and computation resources. Several
parameters need to be well considered due to their
greater impacts on genome assembly, for instance,
length_cutoff which controls the threshold during
the error correction, and length_cutoff_pr which
sets the cutoff value used for the later assembly
overlapping step. Other parameters (-k, -w, -h, -t)
are used to optimize k-mers. The parameters are
determined by the sequencing depth and the
characteristics of the sequenced species. The recommendations can be found in this Web site
(https://pb-falcon.readthedocs.io/en/latest/parameters.html#parameters). It is suggested to choose
a smaller length_cutoff in the initial computation
run, then adjust length_cutoff_pr for a better
assembly. If the coverage of the corrected
high-quality reads longer than the cutoff length is
more than 20x, we are recommended to set the
min_cov to 5, max_cov to three times of coverage
and the max_diff to twice of coverage (Chin 2016).
Several studies have shown that FALCON has
advantages in assembling highly complex plant
genomes, such as maize and opium (Jiao et al.
2017a, b; Guo et al. 2018). It needs to be considered that FALCON requires a high computational
cost to complete the long-read correction and the
overlapping detection due to its alignment
algorithm.
Hierarchical genome assembly process
(HGAP) is developed from FALCON with the
integration of the polish step by using arrow or
quiver. A small genome with hundreds of Mb
could be assembled via a web-based graphical
user interface of HGAP, while a large genome
needs to be run through the environment of the
UNIX command line. HGAP has been widely
used in the assemblies of multiple plant genomes,
as well as of small genomes (VanBuren et al.
2015; Jiao et al. 2017b; Lan et al. 2017).
Canu, a successor of the Celera Assembler
(Denisov et al. 2008), can assemble both PacBio
and Nanopore sequencing reads. By the fact of
the optimized algorithms in the initial overlapping and correction process, Canu is often able to
generate a complete plant assembly less time
than FALCON (Koren et al. 2017). The genome
size is very critical parameter in Canu, which
decides how sensitive the mhap overlapper
should be. The rawErrorRate and correctedErrorRate are another two main parameters
which are involved in overlap detection. A more
accurate assembly will be achieved with a
preferably smaller parameter than the default
72
X. Xiang and C. Li
2011).
6.2.2 De Novo Assembling Tools
for PacBio or Nanopore
Long Reads
There are still no released duckweed genomes
sequenced by the third-generation sequencing of
PacBio or Nanopore platforms that generate long
reads. Here, we recommend four popular de novo
assembling tools for long reads, and expect they
can benefit the assembly of duckweed genomes
in the near future.
Certain alignment algorithms have been
developed to effectively discover overlaps among
noisy long reads, such as DALIGNER (Myers
2014), BLASR (Chaisson and Tesler 2012),
MHAP (Berlin et al. 2015), GraphMap (Sovic
et al. 2016), and Minimap (Li 2016).
DALIGNER was the first tool designed specifically for finding the overlaps between noisy long
reads with the filtering or removals of k-mers
redundancy. The program could increase the
computation speed, decrease memory usage, and
mitigate the effect of repetitive sequences. Given
the risk of filtering important k-mers, the
parameters (-k, -w, -h, -t) need to be customized.
Another method to improve the alignment efficiency is to split the large data set into small
blocks based on the total number of base pairs
and read lengths by using the DBsplit utility
(Myers 2014).
FALCON is an overlap–layout–consensus
(OLC) genome assembler based on DALIGNER,
which only supports PacBio data. The pipeline
includes six steps to construct contigs, while the
correction and polish of raw reads are the steps that
consume the most computational resources. The
script “fc_run.py” with a configuration file can
complete the whole assembly process. The configuration file contains input files, optimal
parameters, and computation resources. Several
parameters need to be well considered due to their
greater impacts on genome assembly, for instance,
length_cutoff which controls the threshold during
the error correction, and length_cutoff_pr which
sets the cutoff value used for the later assembly
overlapping step. Other parameters (-k, -w, -h, -t)
are used to optimize k-mers. The parameters are
determined by the sequencing depth and the
characteristics of the sequenced species. The recommendations can be found in this Web site
(https://pb-falcon.readthedocs.io/en/latest/parameters.html#parameters). It is suggested to choose
a smaller length_cutoff in the initial computation
run, then adjust length_cutoff_pr for a better
assembly. If the coverage of the corrected
high-quality reads longer than the cutoff length is
more than 20x, we are recommended to set the
min_cov to 5, max_cov to three times of coverage
and the max_diff to twice of coverage (Chin 2016).
Several studies have shown that FALCON has
advantages in assembling highly complex plant
genomes, such as maize and opium (Jiao et al.
2017a, b; Guo et al. 2018). It needs to be considered that FALCON requires a high computational
cost to complete the long-read correction and the
overlapping detection due to its alignment
algorithm.
Hierarchical genome assembly process
(HGAP) is developed from FALCON with the
integration of the polish step by using arrow or
quiver. A small genome with hundreds of Mb
could be assembled via a web-based graphical
user interface of HGAP, while a large genome
needs to be run through the environment of the
UNIX command line. HGAP has been widely
used in the assemblies of multiple plant genomes,
as well as of small genomes (VanBuren et al.
2015; Jiao et al. 2017b; Lan et al. 2017).
Canu, a successor of the Celera Assembler
(Denisov et al. 2008), can assemble both PacBio
and Nanopore sequencing reads. By the fact of
the optimized algorithms in the initial overlapping and correction process, Canu is often able to
generate a complete plant assembly less time
than FALCON (Koren et al. 2017). The genome
size is very critical parameter in Canu, which
decides how sensitive the mhap overlapper
should be. The rawErrorRate and correctedErrorRate are another two main parameters
which are involved in overlap detection. A more
accurate assembly will be achieved with a
preferably smaller parameter than the default
72
X. Xiang and C. Li
