PacBio and Oxford Nanopore reads have a considerably higher error rate (up to 15% errors)
than Illumina technology (Mardis 2017).
Although this error rate will likely improve as
new protocols become available, it is problematic for accurate genome sequencing and
assembly. Approaches for assembly using
these long reads are either high coverage
sequencing (Chin et al. 2013) or a hybrid
approach that uses Illumina reads to correct
sequencing and assembly errors (Walker et al.
2014). As these long-read NGS technologies
mature further, it is likely that obtaining
genome assemblies with telomere-to-telomere
chromosomes will become trivial and affordable within a few years.
III. Genome Annotation
Sequencing a genome is only the first step, and
the even more important next step is to annotate the genome. This process generally
includes the identification of regions of repetitive DNA, the prediction of genes, and a function prediction for these genes and domains.
These individual steps can be strung together
into a pipeline. Several pipelines exist for
eukaryotic genome annotation, and two examples of frequently used pipelines for fungal
genome annotation are MAKER (Cantarel
et al. 2008) and the pipeline used by the US
DOE Joint Genome Institute (Haridas et al.
2018). This section describes the steps of
genome annotation in more detail.
A. Repeats
The term “repeat” may refer to various types of
sequences: “low-complexity regions” (sometimes called “simple repeats”) such as a homopolymeric run of nucleotides, as well as
transposable (mobile) elements (transposons)
(Kapitonov and Jurka 2008). These transposable elements can essentially copy themselves
and thus spread throughout the genome. They
can be subdivided into two classes, depending
on their mode of proliferation (Wicker et al.
2007). Class I elements use an RNAintermediate (reminiscent of a retrovirus) and
move via a “copy-paste” mechanism. They
include long interspersed nuclear elements
(LINEs), short interspersed nuclear elements
(SINEs), and long terminal repeats (LTRs).
Class II elements move via a DNA intermediate
and include helitrons and terminal inverted
repeats (Kapitonov and Jurka 2001, 2008).
Since repetitive regions are markedly different from gene-coding regions, it is common
practice to “mask” the repetitive regions before
commencing gene prediction. Masking ensures
that any spurious open reading frames that may
be present in the repeats will not confound (the
training of) the gene predictor. Repeats can be
identified in a newly sequenced genome using
either homology-based or de novo tools.
Homology-based tools rely on a database of
known repetitive elements such as Repbase
(Jurka et al. 2005) and a search algorithm such
as RepeatMasker (Smit et al. 2015). Novel or
genome-specific repeats can be identified using
de novo tools such as Repeatscout (Price et al.
2005), which looks for sequences that occur
repeatedly throughout the genome. Since transposable elements tend to be relatively AT-rich,
their proliferation can result in large AT-rich
regions. Those regions can be distinguished
from gene-coding GC-rich regions by tools
such as OcculterCut (Testa et al. 2016).
The repetitive content of the genome varies
widely between fungi. For example, the very
compact 13.6 Mbp genome of the fern pathogen
Mixia osmundae has a repetitive content of
<1% (Toome et al. 2014), whereas the
177.6 Mbp genome of the mycorrhizal ascomycete Cenococcum geophilum consist for 81% of
repetitive sequences (Peter et al. 2016). Repetitive sequences are usually predominantly found
in centromeric and sub-telomeric regions but
may be spread throughout the assembly. Generally, self-replicating repeats are considered
deleterious since their spread may interrupt
genes. Fungi have evolved a defense mechanism
that recognizes repeats and inactivates these by
causing point mutations (repeat-induced point
mutations, or RIP) (Clutterbuck 2011; Castanera et al. 2016). Intriguingly, genome sequencing of several plant pathogens has revealed that
pathogenesis-related genes frequently co9 Fungal Genomics
209
than Illumina technology (Mardis 2017).
Although this error rate will likely improve as
new protocols become available, it is problematic for accurate genome sequencing and
assembly. Approaches for assembly using
these long reads are either high coverage
sequencing (Chin et al. 2013) or a hybrid
approach that uses Illumina reads to correct
sequencing and assembly errors (Walker et al.
2014). As these long-read NGS technologies
mature further, it is likely that obtaining
genome assemblies with telomere-to-telomere
chromosomes will become trivial and affordable within a few years.
III. Genome Annotation
Sequencing a genome is only the first step, and
the even more important next step is to annotate the genome. This process generally
includes the identification of regions of repetitive DNA, the prediction of genes, and a function prediction for these genes and domains.
These individual steps can be strung together
into a pipeline. Several pipelines exist for
eukaryotic genome annotation, and two examples of frequently used pipelines for fungal
genome annotation are MAKER (Cantarel
et al. 2008) and the pipeline used by the US
DOE Joint Genome Institute (Haridas et al.
2018). This section describes the steps of
genome annotation in more detail.
A. Repeats
The term “repeat” may refer to various types of
sequences: “low-complexity regions” (sometimes called “simple repeats”) such as a homopolymeric run of nucleotides, as well as
transposable (mobile) elements (transposons)
(Kapitonov and Jurka 2008). These transposable elements can essentially copy themselves
and thus spread throughout the genome. They
can be subdivided into two classes, depending
on their mode of proliferation (Wicker et al.
2007). Class I elements use an RNAintermediate (reminiscent of a retrovirus) and
move via a “copy-paste” mechanism. They
include long interspersed nuclear elements
(LINEs), short interspersed nuclear elements
(SINEs), and long terminal repeats (LTRs).
Class II elements move via a DNA intermediate
and include helitrons and terminal inverted
repeats (Kapitonov and Jurka 2001, 2008).
Since repetitive regions are markedly different from gene-coding regions, it is common
practice to “mask” the repetitive regions before
commencing gene prediction. Masking ensures
that any spurious open reading frames that may
be present in the repeats will not confound (the
training of) the gene predictor. Repeats can be
identified in a newly sequenced genome using
either homology-based or de novo tools.
Homology-based tools rely on a database of
known repetitive elements such as Repbase
(Jurka et al. 2005) and a search algorithm such
as RepeatMasker (Smit et al. 2015). Novel or
genome-specific repeats can be identified using
de novo tools such as Repeatscout (Price et al.
2005), which looks for sequences that occur
repeatedly throughout the genome. Since transposable elements tend to be relatively AT-rich,
their proliferation can result in large AT-rich
regions. Those regions can be distinguished
from gene-coding GC-rich regions by tools
such as OcculterCut (Testa et al. 2016).
The repetitive content of the genome varies
widely between fungi. For example, the very
compact 13.6 Mbp genome of the fern pathogen
Mixia osmundae has a repetitive content of
<1% (Toome et al. 2014), whereas the
177.6 Mbp genome of the mycorrhizal ascomycete Cenococcum geophilum consist for 81% of
repetitive sequences (Peter et al. 2016). Repetitive sequences are usually predominantly found
in centromeric and sub-telomeric regions but
may be spread throughout the assembly. Generally, self-replicating repeats are considered
deleterious since their spread may interrupt
genes. Fungi have evolved a defense mechanism
that recognizes repeats and inactivates these by
causing point mutations (repeat-induced point
mutations, or RIP) (Clutterbuck 2011; Castanera et al. 2016). Intriguingly, genome sequencing of several plant pathogens has revealed that
pathogenesis-related genes frequently co9 Fungal Genomics
209
