mcFISH with 96 BACs probes from 20 chromosome pairs (Cao et al. 2016). Hi-C technique
could be an alternative way to reconstruct the
chromosomes of duckweed genomes.
Still, the technologies of Bionano, Hi-C, and
10X Genomics are lack of the fine resolution to
improve contig length compared with PacBio or
Nanopore sequencing. Currently, PacBio or
Nanopore sequencing combined with Bionano
optical map or 10X Genomics and Hi-C technology will be the best choice to optimize the
assembly of complex duckweed genomes.
6.2 Tools for Assembling
Duckweeds
6.2.1 De Novo Assembling Tools
for Illumina Paired-End
Reads
The selection of sequencing platforms for duckweeds is determined by their characteristics of
the genomes. In addition, the bioinformatics
programs for assembling genomes are also needed to customize. Here, we will introduce the
most common programs in terms of duckweed
genome assembly by using next-generation
sequencing.
SOAPdenovo (http://soap.genomics.org.cn/
soapdenovo.html) (Luo et al. 2012) is a de
novo assembly software developed by BGI,
based on a de Bruijn graph-based algorithm. It is
featured in the fast speed of genome assembly
and the long scaffold N50 value, but the error
rate is higher than other programs. Therefore,
SOAPdenovo is widely used for assembly of
large genomes, such as barley (Mascher et al.
2017), maize (Hirsch et al. 2016), and quinoa
(Jarvis et al. 2017).
ALLPATHS-LG
[(http://software.
broadinstitute.org/allpaths-lg/blog/)
(Gnerre
et al. 2011)], the short-read genome assembler
from the Computational Research and Development group at the Broad Institute, also based on a
de Bruijn graph-based algorithm, has a high
assembly accuracy, but consumes a significant
memory and CPU resource (Henson et al. 2012).
The following issues need to be considered.
ALLPATHS-LG needs sequence data from
multiple libraries with various insert sizes,
including at least one paired reads from an
“overlapping” fragment library. For example, the
paired reads are overlapped with a read length of
*100 bp from inserted fragments of
*180 bp. Genome assembly requires the
sequence coverage of 100X given the raw reads
data (before error correction and filtering).
ALLPATHS-LG does not support distributed
computing using MPI, but makes use of shared
memory parallelization. The usage of peak
memory is roughly 1.7 bytes per read base,
resulting in 17 G of memory for handling 10 Gb
of input data. Different from other de novo
assembly software, ALLPATHS-LG could
determine the optimal K value after a series of
self-trainings during the run (http://software.
broadinstitute.org/allpaths-lg/blog/?page_id=336
). ALLPATHS-LG supports the hybrid assembly
with PacBio long reads, but it is still limited to
assemble the small microbial genomes (Koren
et al. 2012; Shibata et al. 2013; Koren and
Phillippy 2015), which is not applied to any
animal and plant genomes yet.
MaSuRCA (Zimin et al. 2013) was derived
from the Celera Assembler (Myers et al. 2000),
combining the algorithm of the de Bruijn graph
and overlap–layout–consensus (OLC) approaches. MaSuRCA supports not only Illumina
short reads, but also a hybrid of short and long
reads. The MaSuRCA genome assembler has
been widely used in the field of large animal and
plant genomes (Chibucos et al. 2013) (Zimin
et al. 2014; Zimin et al. 2017).
To achieve the most continuous assembly of
Lemna minor 5500 genome, three programs were
evaluated including SOAPdenovo2, CLC bio,
and MaSuRCA. The draft genome generated by
MaSuRCA is the best compared to that of
SOAPdenovo2 and CLC bio (Van Hoeck et al.
2015). A high-quality draft genome of Spirodela
polyrhiza 9509 containing 774 scaffolds with an
N50 length of 4.3 Mb and a contig N50 length of
19 kb was reached using a combination of
6 Strategies and Tools for Sequencing Duckweeds
71
could be an alternative way to reconstruct the
chromosomes of duckweed genomes.
Still, the technologies of Bionano, Hi-C, and
10X Genomics are lack of the fine resolution to
improve contig length compared with PacBio or
Nanopore sequencing. Currently, PacBio or
Nanopore sequencing combined with Bionano
optical map or 10X Genomics and Hi-C technology will be the best choice to optimize the
assembly of complex duckweed genomes.
6.2 Tools for Assembling
Duckweeds
6.2.1 De Novo Assembling Tools
for Illumina Paired-End
Reads
The selection of sequencing platforms for duckweeds is determined by their characteristics of
the genomes. In addition, the bioinformatics
programs for assembling genomes are also needed to customize. Here, we will introduce the
most common programs in terms of duckweed
genome assembly by using next-generation
sequencing.
SOAPdenovo (http://soap.genomics.org.cn/
soapdenovo.html) (Luo et al. 2012) is a de
novo assembly software developed by BGI,
based on a de Bruijn graph-based algorithm. It is
featured in the fast speed of genome assembly
and the long scaffold N50 value, but the error
rate is higher than other programs. Therefore,
SOAPdenovo is widely used for assembly of
large genomes, such as barley (Mascher et al.
2017), maize (Hirsch et al. 2016), and quinoa
(Jarvis et al. 2017).
ALLPATHS-LG
[(http://software.
broadinstitute.org/allpaths-lg/blog/)
(Gnerre
et al. 2011)], the short-read genome assembler
from the Computational Research and Development group at the Broad Institute, also based on a
de Bruijn graph-based algorithm, has a high
assembly accuracy, but consumes a significant
memory and CPU resource (Henson et al. 2012).
The following issues need to be considered.
ALLPATHS-LG needs sequence data from
multiple libraries with various insert sizes,
including at least one paired reads from an
“overlapping” fragment library. For example, the
paired reads are overlapped with a read length of
*100 bp from inserted fragments of
*180 bp. Genome assembly requires the
sequence coverage of 100X given the raw reads
data (before error correction and filtering).
ALLPATHS-LG does not support distributed
computing using MPI, but makes use of shared
memory parallelization. The usage of peak
memory is roughly 1.7 bytes per read base,
resulting in 17 G of memory for handling 10 Gb
of input data. Different from other de novo
assembly software, ALLPATHS-LG could
determine the optimal K value after a series of
self-trainings during the run (http://software.
broadinstitute.org/allpaths-lg/blog/?page_id=336
). ALLPATHS-LG supports the hybrid assembly
with PacBio long reads, but it is still limited to
assemble the small microbial genomes (Koren
et al. 2012; Shibata et al. 2013; Koren and
Phillippy 2015), which is not applied to any
animal and plant genomes yet.
MaSuRCA (Zimin et al. 2013) was derived
from the Celera Assembler (Myers et al. 2000),
combining the algorithm of the de Bruijn graph
and overlap–layout–consensus (OLC) approaches. MaSuRCA supports not only Illumina
short reads, but also a hybrid of short and long
reads. The MaSuRCA genome assembler has
been widely used in the field of large animal and
plant genomes (Chibucos et al. 2013) (Zimin
et al. 2014; Zimin et al. 2017).
To achieve the most continuous assembly of
Lemna minor 5500 genome, three programs were
evaluated including SOAPdenovo2, CLC bio,
and MaSuRCA. The draft genome generated by
MaSuRCA is the best compared to that of
SOAPdenovo2 and CLC bio (Van Hoeck et al.
2015). A high-quality draft genome of Spirodela
polyrhiza 9509 containing 774 scaffolds with an
N50 length of 4.3 Mb and a contig N50 length of
19 kb was reached using a combination of
6 Strategies and Tools for Sequencing Duckweeds
71
