Therefore, a sequencing approach with
39Â coverage of the W. australiana strain
DWC304 using the Illumina next-generation
sequencing platform (HiSeq PE150) appears
sub-standardly. However, our focus was not on
clarifying the genome sequence as precisely as
possible, but our focus was on finding suitable
target sequences for the genome editing experiments described below.
It is not surprising that the degree of
heterozygosity is very low, which reflects the
preferential vegetative propagation of all duckweeds. As is true for the sequenced genomes of
the other genera (Van Hoeck et al. 2015; Wang
et al. 2014; Cao et al. 2016; Michael et al. 2017),
the percentage of repetitive sequences is quite
high. This repetition complicates the assembly
and annotation of the genome data. In contrast to
the genome projects mentioned above, there was
hardly any transcriptome data, besides an
EST-library of about 1,988 sequences
(SAMN00222771, ID: 222771) and the complete
sequence of the plastome (Wang and Messing
2011). The assembly was performed using the
well-established protocol described by Li et al.
(2010), and k-mer 17 analysis revealed a genome
size of 385 Mbp for W. australiana (Marçais and
Kingsford 2011) and as described here, http://
koke.asrc.kanazawa-u.ac.jp.
Due to the absence of the necessary computing capacity, a recently established, web-based
pipeliner, the Genome Sequence Annotation
Server, was chosen for the annotation (Humann
et al. 2017). The GenSAS pipeliner (www.
gensas.org) combines many tools and is precisely configurable. Because one may upload
DNA evidence data, we added gene sets from
Liliopsidae, ESTs from Lemnaceae, and 1,790
ESTs from a W. australiana EST project
(JZ896467.1) in addition to the predefined plant
reference genome dataset (Humann et al. 2017).
After eliminating repetitive sequences (RepeatMasker and RepeatModeler), the alignment can
be done using BLAT, nucleotide BLAST, and
PASA. The data obtained was fine-tuned using
gene modelers like Augustus and SNAP. A set of
protein sequence-based annotation tools like
BLASTp, InterProScan, Pfam, SignalP, and
TargetP was applied, followed by the creation of
the official gene set. This service allows even
small workgroups without access to mainframes
the annotation of genome data with a
user-friendly interface.
The gene set obtained from GenSAS was used
for further functional genome analysis on the
protein level using the BLAST2GO software
package (www.blast2go.com), which resulted in
the identification of 18,617 protein sequences
from W. australiana.
However, our main objective was to identify
the genome context of the adh1 gene (alcohol
dehydrogenase 1), whose gene product can be
used for selection using prop-2-en-1-ol, also
known as allyl alcohol (Widholm and Kishinami
1988). Inactivation of ADH1 enzyme enables the
plant to grow on allyl alcohol, which is otherwise
toxic to the plant. It should be emphasized that
there are typically several isoenzymes in a plant
genome and this protein class shows only relatively weak homologies. Therefore, a PCR-based
amplification of the gene from the genomic DNA
was not possible; this is why genome sequencing
was required. After identification of the adh1
locus in W. australiana, we were able to use this
locus as a selection system for targeted genome
editing events without the need for other selection markers.
17.4 Genome Editing
of W. australiana
Genome editing has become a major force in
modern biotechnology. Its popularity is due to its
underlying technologies which allow for the easy
study of genes and their functions through
knock-out/knock-ins or through the regulation of
gene expression. Furthermore, genome editing
allows for specific insertion of the gene of
interest into a predefined site of the genome
(knock-in). Because the knock-in scenario is
somewhat tricky, we focused on a proof of
principle approach using a knock-out strategy.
We combined our knock-out strategy with the
option to establish an in planta selection method,
which would enable us to omit selection marker
170
T. Reinard et al.
39Â coverage of the W. australiana strain
DWC304 using the Illumina next-generation
sequencing platform (HiSeq PE150) appears
sub-standardly. However, our focus was not on
clarifying the genome sequence as precisely as
possible, but our focus was on finding suitable
target sequences for the genome editing experiments described below.
It is not surprising that the degree of
heterozygosity is very low, which reflects the
preferential vegetative propagation of all duckweeds. As is true for the sequenced genomes of
the other genera (Van Hoeck et al. 2015; Wang
et al. 2014; Cao et al. 2016; Michael et al. 2017),
the percentage of repetitive sequences is quite
high. This repetition complicates the assembly
and annotation of the genome data. In contrast to
the genome projects mentioned above, there was
hardly any transcriptome data, besides an
EST-library of about 1,988 sequences
(SAMN00222771, ID: 222771) and the complete
sequence of the plastome (Wang and Messing
2011). The assembly was performed using the
well-established protocol described by Li et al.
(2010), and k-mer 17 analysis revealed a genome
size of 385 Mbp for W. australiana (Marçais and
Kingsford 2011) and as described here, http://
koke.asrc.kanazawa-u.ac.jp.
Due to the absence of the necessary computing capacity, a recently established, web-based
pipeliner, the Genome Sequence Annotation
Server, was chosen for the annotation (Humann
et al. 2017). The GenSAS pipeliner (www.
gensas.org) combines many tools and is precisely configurable. Because one may upload
DNA evidence data, we added gene sets from
Liliopsidae, ESTs from Lemnaceae, and 1,790
ESTs from a W. australiana EST project
(JZ896467.1) in addition to the predefined plant
reference genome dataset (Humann et al. 2017).
After eliminating repetitive sequences (RepeatMasker and RepeatModeler), the alignment can
be done using BLAT, nucleotide BLAST, and
PASA. The data obtained was fine-tuned using
gene modelers like Augustus and SNAP. A set of
protein sequence-based annotation tools like
BLASTp, InterProScan, Pfam, SignalP, and
TargetP was applied, followed by the creation of
the official gene set. This service allows even
small workgroups without access to mainframes
the annotation of genome data with a
user-friendly interface.
The gene set obtained from GenSAS was used
for further functional genome analysis on the
protein level using the BLAST2GO software
package (www.blast2go.com), which resulted in
the identification of 18,617 protein sequences
from W. australiana.
However, our main objective was to identify
the genome context of the adh1 gene (alcohol
dehydrogenase 1), whose gene product can be
used for selection using prop-2-en-1-ol, also
known as allyl alcohol (Widholm and Kishinami
1988). Inactivation of ADH1 enzyme enables the
plant to grow on allyl alcohol, which is otherwise
toxic to the plant. It should be emphasized that
there are typically several isoenzymes in a plant
genome and this protein class shows only relatively weak homologies. Therefore, a PCR-based
amplification of the gene from the genomic DNA
was not possible; this is why genome sequencing
was required. After identification of the adh1
locus in W. australiana, we were able to use this
locus as a selection system for targeted genome
editing events without the need for other selection markers.
17.4 Genome Editing
of W. australiana
Genome editing has become a major force in
modern biotechnology. Its popularity is due to its
underlying technologies which allow for the easy
study of genes and their functions through
knock-out/knock-ins or through the regulation of
gene expression. Furthermore, genome editing
allows for specific insertion of the gene of
interest into a predefined site of the genome
(knock-in). Because the knock-in scenario is
somewhat tricky, we focused on a proof of
principle approach using a knock-out strategy.
We combined our knock-out strategy with the
option to establish an in planta selection method,
which would enable us to omit selection marker
170
T. Reinard et al.
