ungapped aligners are used, since the alignment process does not essentially incorporate a
range of gaps. However, as the alignment process utilizes a relatively reduced reference
size in the form of selected transcriptomic sequences, it is not particularly useful for
identifying novel exons, their expression patterns or splicing events [43, 44]. This is simply
because all the available isoforms are aligned to the same exon of a gene multiple times,
due to pre-existing splicing information. Bowtie2 [45] and MAQ [46] are the two most
widely used aligners for transcriptome-based alignment/mapping. Bowtie2 utilizes a
Burrows–Wheeler Transform (BWT) based FM index (Full-text index in Minute space)
to perform fast and memory efficient alignments. On the other hand, MAQ identifies
ungapped matches with low mismatch scores and assign them to a threshold-based
mapping algorithm. This generates phred-scaled quality scores for each read alignment
and a final alignment is produced by collating the best scoring matches for every read.
Reference genomes, e.g. for mouse, together with annotations can be downloaded from
GENCODE [47] using the following command:
11.6.1.1 Choice of Reference-Based Alignment Program
The choice of a reference-based aligner for RNA-Seq data depends mainly on the splicing
information or simply the type of organism [48]. For RNA-Seq reads originating from
organisms without introns, that is prokaryotes/archaea, ungapped aligners or aligners that
are not splice-aware are utilized. However, if the reads are aligned to intronic genomes
(eukaryotic), splice-aware aligners like TopHat2 [41], STAR [39], or BBMap [38] are used
that incorporate a three-step alignment procedure. In the first step, sequencing reads are
aligned to a reference genome or transcriptome sequence. Next, overlapping reads present
at each locus are collated into a graph representing all isoforms. In the final step, the graph
is resolved, and all isoforms associated with individual genes are identified [35]. In our
example (see GitHub and Fig. 11.2), we use the commonly used tools TopHat2 and STAR.
TopHat2: As mentioned in Chap. 9, Sect. 9.2.3.4 TopHat2 is a fast and memory
efficient tool that utilizes Bowtie2 for alignment. Genomic annotations are used together
with the available transcriptome information to produce precise spliced alignments. In
addition, TopHat2 allows excluding pseudogene sequences and shows minimum tolerance
to mismatches and non-alignment of low-quality bases. TopHat2 incorporates a multistep
alignment algorithm that consists of an optional transcriptional alignment followed by a
152
R. Bharti and D. G. Grimm
range of gaps. However, as the alignment process utilizes a relatively reduced reference
size in the form of selected transcriptomic sequences, it is not particularly useful for
identifying novel exons, their expression patterns or splicing events [43, 44]. This is simply
because all the available isoforms are aligned to the same exon of a gene multiple times,
due to pre-existing splicing information. Bowtie2 [45] and MAQ [46] are the two most
widely used aligners for transcriptome-based alignment/mapping. Bowtie2 utilizes a
Burrows–Wheeler Transform (BWT) based FM index (Full-text index in Minute space)
to perform fast and memory efficient alignments. On the other hand, MAQ identifies
ungapped matches with low mismatch scores and assign them to a threshold-based
mapping algorithm. This generates phred-scaled quality scores for each read alignment
and a final alignment is produced by collating the best scoring matches for every read.
Reference genomes, e.g. for mouse, together with annotations can be downloaded from
GENCODE [47] using the following command:
11.6.1.1 Choice of Reference-Based Alignment Program
The choice of a reference-based aligner for RNA-Seq data depends mainly on the splicing
information or simply the type of organism [48]. For RNA-Seq reads originating from
organisms without introns, that is prokaryotes/archaea, ungapped aligners or aligners that
are not splice-aware are utilized. However, if the reads are aligned to intronic genomes
(eukaryotic), splice-aware aligners like TopHat2 [41], STAR [39], or BBMap [38] are used
that incorporate a three-step alignment procedure. In the first step, sequencing reads are
aligned to a reference genome or transcriptome sequence. Next, overlapping reads present
at each locus are collated into a graph representing all isoforms. In the final step, the graph
is resolved, and all isoforms associated with individual genes are identified [35]. In our
example (see GitHub and Fig. 11.2), we use the commonly used tools TopHat2 and STAR.
TopHat2: As mentioned in Chap. 9, Sect. 9.2.3.4 TopHat2 is a fast and memory
efficient tool that utilizes Bowtie2 for alignment. Genomic annotations are used together
with the available transcriptome information to produce precise spliced alignments. In
addition, TopHat2 allows excluding pseudogene sequences and shows minimum tolerance
to mismatches and non-alignment of low-quality bases. TopHat2 incorporates a multistep
alignment algorithm that consists of an optional transcriptional alignment followed by a
152
R. Bharti and D. G. Grimm
