11.6 RNA-Seq Analysis
There are two main strategies for RNA-Seq analysis, one that is based on reference-based
alignment and the second that is based on a de novo assembly. In addition, there are hybrid
approaches that combine both strategies [4]. In the following, we will describe details for
both, the reference-based alignment and the de novo assembly.
11.6.1 Reference-Based Alignment
A reference-based alignment is a strategy to align (map) individual reads (based on
sequence similarities) against a chosen reference genome or a transcriptome sequence
[31], as illustrated in Fig. 11.3. One aim of this approach is to quantitate transcript
abundance at a genomic locus [32]. Here, each identified mapping location must indicate
the origin of transcript as well as the total number of reads corresponding to the same
transcript present in the dataset.
Reference-Genome Based Alignment A reference genome sequence is utilized for
aligning the reads. For an accurate mapping it is important to consider that reads might
originate from either cDNA corresponding to spliced transcripts or from non-spliced
transcripts [33, 34]. In the case of spliced transcripts, contiguous read sequences are
separated by intervening splice junction (SJ) boundaries and are split into two fragments
and assigned separately. Thus, there are two main alignment approaches for aligning reads
against a reference genome, that is spliced alignment and unspliced alignment [35, 36]. In
addition, the alignment of reads against a reference genome leads to unaligned regions and
gaps. Apparently, splice-aware aligners are generally suitable for most reference-genome
based alignments. As the information on SJs in the reference genome is crucial, most
aligners also refer to known annotated SJ databases to confirm the presence or absence of
splice sites [37]. In addition, most aligners also perform sequence matching of terminal
nucleotides in aligned reads with donor–acceptor sites of known splicing sites. Currently a
number of splice-aware alignment tools are available, such as BBMap [38], STAR [39],
GMAP [40], and TopHat2 [41]. STAR utilizes a sequential maximum mappable seed search
algorithm that consolidates into seed clustering and stitching steps to generate alignments.
In contrast, TopHat2 incorporates an exhaustive initial step where alignments are scanned
for possible exon–exon junctions, which are eventually utilized in the final step for
generating the final alignment. Importantly, parameters of alignment tools have to be
carefully chosen to gain optimal alignment results.
Reference-Transcriptome Based Alignment This alignment method is mainly utilized
when a well-annotated transcriptome or a set of known transcripts are available. Here,
sequencing reads can be directly aligned in a continuous, error-free manner without
involving computationally intensive steps, due to the availability of splicing information
for reference transcripts [42]. For this purpose, aligners that are not splice-aware or
150
R. Bharti and D. G. Grimm
There are two main strategies for RNA-Seq analysis, one that is based on reference-based
alignment and the second that is based on a de novo assembly. In addition, there are hybrid
approaches that combine both strategies [4]. In the following, we will describe details for
both, the reference-based alignment and the de novo assembly.
11.6.1 Reference-Based Alignment
A reference-based alignment is a strategy to align (map) individual reads (based on
sequence similarities) against a chosen reference genome or a transcriptome sequence
[31], as illustrated in Fig. 11.3. One aim of this approach is to quantitate transcript
abundance at a genomic locus [32]. Here, each identified mapping location must indicate
the origin of transcript as well as the total number of reads corresponding to the same
transcript present in the dataset.
Reference-Genome Based Alignment A reference genome sequence is utilized for
aligning the reads. For an accurate mapping it is important to consider that reads might
originate from either cDNA corresponding to spliced transcripts or from non-spliced
transcripts [33, 34]. In the case of spliced transcripts, contiguous read sequences are
separated by intervening splice junction (SJ) boundaries and are split into two fragments
and assigned separately. Thus, there are two main alignment approaches for aligning reads
against a reference genome, that is spliced alignment and unspliced alignment [35, 36]. In
addition, the alignment of reads against a reference genome leads to unaligned regions and
gaps. Apparently, splice-aware aligners are generally suitable for most reference-genome
based alignments. As the information on SJs in the reference genome is crucial, most
aligners also refer to known annotated SJ databases to confirm the presence or absence of
splice sites [37]. In addition, most aligners also perform sequence matching of terminal
nucleotides in aligned reads with donor–acceptor sites of known splicing sites. Currently a
number of splice-aware alignment tools are available, such as BBMap [38], STAR [39],
GMAP [40], and TopHat2 [41]. STAR utilizes a sequential maximum mappable seed search
algorithm that consolidates into seed clustering and stitching steps to generate alignments.
In contrast, TopHat2 incorporates an exhaustive initial step where alignments are scanned
for possible exon–exon junctions, which are eventually utilized in the final step for
generating the final alignment. Importantly, parameters of alignment tools have to be
carefully chosen to gain optimal alignment results.
Reference-Transcriptome Based Alignment This alignment method is mainly utilized
when a well-annotated transcriptome or a set of known transcripts are available. Here,
sequencing reads can be directly aligned in a continuous, error-free manner without
involving computationally intensive steps, due to the availability of splicing information
for reference transcripts [42]. For this purpose, aligners that are not splice-aware or
150
R. Bharti and D. G. Grimm
