was that the conserved genes were identified
from only six eukaryotic species. BUSCO takes
a clade-specific approach that is based on more
eukaryotic genomes, and fungi-specific conserved gene sets are available (Sima ˜o et al.
2015). More recently, FGMP (Fungal Genome
Mapping Project) was developed that provides
a computational framework and sequence
resource specifically designed to assess the
completeness of fungal genomes (Cisse ´ and
Stajich 2019). It is based on 246 fungal genomes
and can be used to assess assembly and annotation completeness as well as suggest assembly
improvements.
C. Functional Annotation of the Predicted
Genes
Once a reliable set of genes has been predicted,
the next step is to determine the putative role of
the encoded proteins. This is referred to as
functional annotation of the predicted proteins.
It is important to note, however, that automated function predictions should be interpreted with care. Lab experiments may be
required to definitively confirm the function
of individual genes (e.g., an enzyme activity
assay to confirm the predicted activity of a
putative enzyme).
Functional annotation usually starts with
homology searches in a database of known
proteins, for example, using Blast (Altschul
et al. 1990) to search for homologs in GenBank
(Clark et al. 2016) or UniProt/Swiss-Prot (Bateman et al. 2017). Moreover, conserved protein
domains can be identified using InterPro
(Hunter et al. 2009), which comprises a collection of domain databases that includes PFAM
(Finn et al. 2016). Cellular localization of the
proteins can be predicted using SignalP (Petersen et al. 2011), TMHMM (Krogh et al. 2001),
and WoLF PSORT (Horton et al. 2007). Proteases/peptidases can be identified by homology to known enzymes in the MEROPS
database (Rawlings et al. 2014). More generally,
Gene Ontology (GO) aims to provide a hierarchical functional annotation of the predicted
proteins, based on their molecular function,
cellular localization, and the biological process
they are involved in Ashburner et al. (2000).
Similarly, KEGG (Kyoto Encyclopedia of
Genes and Genomes) provides a classification
system into metabolic pathways, including predicted enzyme activities based on the Enzyme
Commission (EC) system (Kanehisa and Goto
2000).
Several functional annotation approaches
have been developed that aim to identify
genes that are involved in the lifestyle of fungi.
The CAZy (carbohydrate-active enzymes) database focuses on enzymes that assemble, modify,
or break down polysaccharides (Lombard et al.
2014). CAZymes are especially important in the
context of plant biomass breakdown, for example, in lignocellulose degradation and plant disease (further discussed below). Fungi are
known to produce a wide variety of secondary
metabolites and other natural products (further
discussed below). The genes involved in this
process are frequently clustered in the genome,
and these biosynthetic gene clusters can be
identified by tools like AntiSMASH (Blin et al.
2017) or SMURF (Khaldi et al. 2010).
D. Data Visualization, Analysis, and Manual
Curation
Large amounts of data are generated by genome
sequencing and annotation. These can be challenging to interpret unless they are visualized.
Genome sequencing consortia and/or institutes
generally make the data accessible to the public
by means of a centrally hosted web database,
which allows users to analyze the genome
sequence, gene predictions, and functional
annotations. Examples include the genusspecific websites Saccharomyces Genome Database (SGD) and the Aspergillus Genome Database (AspGD) (Cherry et al. 2012; Cerqueira
et al. 2014). MycoCosm hosts all fungal genome
portals of the US DOE Joint Genome Institute
(Grigoriev et al. 2014). FungiDB hosts numerous published fungi (Basenko et al. 2018). Upon
publication of a genome, the data is generally
submitted to NCBI GenBank, which has therefore amassed a large collection of fungal
genome data (Clark et al. 2016).
9 Fungal Genomics
211
from only six eukaryotic species. BUSCO takes
a clade-specific approach that is based on more
eukaryotic genomes, and fungi-specific conserved gene sets are available (Sima ˜o et al.
2015). More recently, FGMP (Fungal Genome
Mapping Project) was developed that provides
a computational framework and sequence
resource specifically designed to assess the
completeness of fungal genomes (Cisse ´ and
Stajich 2019). It is based on 246 fungal genomes
and can be used to assess assembly and annotation completeness as well as suggest assembly
improvements.
C. Functional Annotation of the Predicted
Genes
Once a reliable set of genes has been predicted,
the next step is to determine the putative role of
the encoded proteins. This is referred to as
functional annotation of the predicted proteins.
It is important to note, however, that automated function predictions should be interpreted with care. Lab experiments may be
required to definitively confirm the function
of individual genes (e.g., an enzyme activity
assay to confirm the predicted activity of a
putative enzyme).
Functional annotation usually starts with
homology searches in a database of known
proteins, for example, using Blast (Altschul
et al. 1990) to search for homologs in GenBank
(Clark et al. 2016) or UniProt/Swiss-Prot (Bateman et al. 2017). Moreover, conserved protein
domains can be identified using InterPro
(Hunter et al. 2009), which comprises a collection of domain databases that includes PFAM
(Finn et al. 2016). Cellular localization of the
proteins can be predicted using SignalP (Petersen et al. 2011), TMHMM (Krogh et al. 2001),
and WoLF PSORT (Horton et al. 2007). Proteases/peptidases can be identified by homology to known enzymes in the MEROPS
database (Rawlings et al. 2014). More generally,
Gene Ontology (GO) aims to provide a hierarchical functional annotation of the predicted
proteins, based on their molecular function,
cellular localization, and the biological process
they are involved in Ashburner et al. (2000).
Similarly, KEGG (Kyoto Encyclopedia of
Genes and Genomes) provides a classification
system into metabolic pathways, including predicted enzyme activities based on the Enzyme
Commission (EC) system (Kanehisa and Goto
2000).
Several functional annotation approaches
have been developed that aim to identify
genes that are involved in the lifestyle of fungi.
The CAZy (carbohydrate-active enzymes) database focuses on enzymes that assemble, modify,
or break down polysaccharides (Lombard et al.
2014). CAZymes are especially important in the
context of plant biomass breakdown, for example, in lignocellulose degradation and plant disease (further discussed below). Fungi are
known to produce a wide variety of secondary
metabolites and other natural products (further
discussed below). The genes involved in this
process are frequently clustered in the genome,
and these biosynthetic gene clusters can be
identified by tools like AntiSMASH (Blin et al.
2017) or SMURF (Khaldi et al. 2010).
D. Data Visualization, Analysis, and Manual
Curation
Large amounts of data are generated by genome
sequencing and annotation. These can be challenging to interpret unless they are visualized.
Genome sequencing consortia and/or institutes
generally make the data accessible to the public
by means of a centrally hosted web database,
which allows users to analyze the genome
sequence, gene predictions, and functional
annotations. Examples include the genusspecific websites Saccharomyces Genome Database (SGD) and the Aspergillus Genome Database (AspGD) (Cherry et al. 2012; Cerqueira
et al. 2014). MycoCosm hosts all fungal genome
portals of the US DOE Joint Genome Institute
(Grigoriev et al. 2014). FungiDB hosts numerous published fungi (Basenko et al. 2018). Upon
publication of a genome, the data is generally
submitted to NCBI GenBank, which has therefore amassed a large collection of fungal
genome data (Clark et al. 2016).
9 Fungal Genomics
211
