localize with repetitive sequences in these species. A potential evolutionary benefit of this colocalization is a higher rate of mutation due to
RIP, which in turn may lead to a higher rate of
evolution. This may allow these pathogens to
adapt more quickly to the host plant’s defenses
(Rouxel et al. 2011; Ohm et al. 2012).
B. Gene Prediction
Genes are (arguably) the most important functional elements in a fungal genome. However,
their accurate identification is non-trivial. The
presence of introns in fungal genes precludes
simple scanning for open reading frames
(ORFs), which is a common initial approach
in gene prediction in prokaryotes. The structure of protein-coding genes varies widely
between eukaryotes (Yandell and Ence 2012)
and even between fungi. Differences include
GC content of the coding regions, splicing
acceptor and donor sites, intron length, number of introns per gene, gene length, etc. For
example, the ascomycete yeast S. cerevisiae has
6576 predicted genes with a median gene length
of 1071 bp, of which 4.2% contain an intron
(Goffeau et al. 1996). In contrast, the basidiomycete mushroom-forming fungus Schizophyllum commune has 16204 predicted genes with a
median gene length of 1517 bp, of which 86.3%
contain an intron (Ohm et al. 2010). Therefore,
the gene-finding approach needs to be tailored
to each organism individually.
Gene prediction algorithms can be divided
into two categories: evidence-driven and ab
initio approaches. Evidence-driven predictors
take external evidence to identify the locations
of protein-coding genes. This evidence usually
takes the form of sequenced cDNA (Haas et al.
2003) or homology with known proteins of
related species (Birney et al. 2004). Sequenced
cDNA (in this context usually referred to as
Expressed Sequence Tags or ESTs) are aligned
to the assembly, and exons and intron splice
sites are inferred. This approach has the advantage that it uses evidence specific to the organism but has the disadvantage that unexpressed
genes are less likely to be identified correctly.
Homology-based gene predictors rely on the
alignment of known proteins from related
organisms to identify exons. Advantages of
this approach are that it is cheap (since no
cDNA sequencing is required) but has the disadvantage that organism-specific genes are less
likely to be identified correctly. An ab initio
approach uses a mathematical model of the
gene structure to predict genes. These algorithms require training, which means that they
need to learn what a gene looks like (e.g., typical gene length, intron length, GC content of
coding regions, etc.) from a subset of known
genes. This poses a problem, since for most
fungal genomes there is no prior knowledge
available. Modern approaches use a hybrid
strategy in which RNA-Seq data is used as evidence to train an ab initio gene predictor. The
algorithm BRAKER, for example, only requires
aligned RNA-Seq reads and a genome assembly
and no other prior knowledge (Hoff et al. 2016).
It uses these data to train the ab initio predictors GeneMark (Lomsadze et al. 2014) and
Augustus (Stanke and Waack 2003) and subsequently generates a high-quality gene prediction.
Various gene prediction algorithms may
predict different genes at the same locus.
Although these sometimes represent alternative
splicing variants (especially when the gene predictor uses expression data as evidence), it is
more likely that only one variant is correct.
Various methods have been published that
aim to select the correct gene prediction at
each locus; examples include MAKER (Cantarel
et al. 2008), the US DOE Joint Genome Institute
pipeline (Haridas et al. 2018), and FunGAP
(Min et al. 2017).
The quality and completeness of the set of
predicted genes can be assessed by determining
the percentage of highly conserved eukaryotic
genes that are found in the predicted gene set.
Since these highly conserved genes (histones,
DNA polymerase, etc.) are expected to be present among the genes of the newly sequenced
fungus, their absence can be indicative of an
incompleteness of the genome assembly or the
gene prediction. CEGMA (Core Eukaryotic
Genes Mapping Approach) was initially a popular tool to determine completeness (Parra
et al. 2007). However, a key issue with CEGMA
210
R. A. Ohm
Précédent

- 226/461

Suivant