11.7 Functional Annotation of de novo Transcripts
Functional annotation of the de novo transcripts involves identifying biological information, such as metabolic activity, cellular and physiological functions of predicted genes, or
gene products/proteins [59]. In general, functional annotation can be either performed
using conventional homology search or using a gene ontology (GO-term) based mapping
[60, 61]. In the homology search, closely related protein sequences are initially identified
by using a BLASTp UniProtKB database search [62] and based on protein domains using
the Pfam database [63]. After integrating BLASTp and Pfam outputs, remaining functional
annotation is done using BLASTx and HMMER (http://hmmer.org/). Additionally,
Rnammer [64] and SignalP [65] are used for predicting ribosomal RNA and signal peptide
sequences, respectively.
In the gene ontology-based annotation, GO-terms associated with hits obtained from
Blast results are retrieved and catalogued into biological process ontologies, molecular
function ontologies, or cellular component ontologies. The ontology data provides information about the functions and physiological activities of identified gene products. Another
widely used tool, Blast2GO [66] utilizes statistics of GO-term frequencies for analyzing the
enrichment of GO annotations. Additionally, the Kyoto Encyclopedia of Genes and
Genomes (KEGG) [67] pathways are used for predicting interactions between gene
products and related metabolic activities [32].
11.8 Post-alignment/assembly Assessment and Statistics
After alignment and read mapping, an assessment is done to analyze the quality of the
alignment based on information, such as total number of processed reads, % of mapped
reads, or SJs identified with uniquely mapped reads. Before any downstream analysis can
be done, the output files require some post-processing, including file format conversions
(SAM/BAM), sorting, indexing, and merging. There are a variety of tools available for
post-processing, that is SAMtools [68], BAMtools [69], Sambamba [70], and Biobambam
[71]. SAMtools mainly incorporates methods for file conversions from SAM- into BAMformat or vice versa. This is important because the BAM file format is one of the main file
formats for several downstream tools. Besides, SAMtools is frequently used for sorting and
listing alignments in BAM files, e.g. based on mapping quality and statistics. Many of these
tasks are summarized in our RNA-Seq workflow available on GitHub: https://github.com/
grimmlab/BookChapter-RNA-Seq-Analyses.
160
R. Bharti and D. G. Grimm
Functional annotation of the de novo transcripts involves identifying biological information, such as metabolic activity, cellular and physiological functions of predicted genes, or
gene products/proteins [59]. In general, functional annotation can be either performed
using conventional homology search or using a gene ontology (GO-term) based mapping
[60, 61]. In the homology search, closely related protein sequences are initially identified
by using a BLASTp UniProtKB database search [62] and based on protein domains using
the Pfam database [63]. After integrating BLASTp and Pfam outputs, remaining functional
annotation is done using BLASTx and HMMER (http://hmmer.org/). Additionally,
Rnammer [64] and SignalP [65] are used for predicting ribosomal RNA and signal peptide
sequences, respectively.
In the gene ontology-based annotation, GO-terms associated with hits obtained from
Blast results are retrieved and catalogued into biological process ontologies, molecular
function ontologies, or cellular component ontologies. The ontology data provides information about the functions and physiological activities of identified gene products. Another
widely used tool, Blast2GO [66] utilizes statistics of GO-term frequencies for analyzing the
enrichment of GO annotations. Additionally, the Kyoto Encyclopedia of Genes and
Genomes (KEGG) [67] pathways are used for predicting interactions between gene
products and related metabolic activities [32].
11.8 Post-alignment/assembly Assessment and Statistics
After alignment and read mapping, an assessment is done to analyze the quality of the
alignment based on information, such as total number of processed reads, % of mapped
reads, or SJs identified with uniquely mapped reads. Before any downstream analysis can
be done, the output files require some post-processing, including file format conversions
(SAM/BAM), sorting, indexing, and merging. There are a variety of tools available for
post-processing, that is SAMtools [68], BAMtools [69], Sambamba [70], and Biobambam
[71]. SAMtools mainly incorporates methods for file conversions from SAM- into BAMformat or vice versa. This is important because the BAM file format is one of the main file
formats for several downstream tools. Besides, SAMtools is frequently used for sorting and
listing alignments in BAM files, e.g. based on mapping quality and statistics. Many of these
tasks are summarized in our RNA-Seq workflow available on GitHub: https://github.com/
grimmlab/BookChapter-RNA-Seq-Analyses.
160
R. Bharti and D. G. Grimm
