11.3 RNA-Seq Library Preparation
The basic steps in RNA-Seq library preparation include efficient ribosomal RNA (rRNA)
removal from samples followed by cDNA synthesis for producing directional RNA-Seq
libraries [18, 19]. Few of the widely used RNA-Seq library preparation kits include Illumina
TruSeq (https://www.illumina.com/products/by-type/sequencing-kits/library-prep-kits/
truseq-rna-v2.html), Bioo Scientific NEXTFlex (http://shop.biooscientific.com/nextflexsmall-rna-seq-kit-v3/), and New England Biolabs NEB Next Ultra (https://www.neb.com/
products/e7370-nebnext-ultra-dna-library-prep-kit-for-illumina#Product%20Information).
In general, after removal of rRNA fractions, the remaining RNA is fragmented and reverse
transcribed using random primers with 5’-tagging sequences. Next, the 5’-tagged cDNA are
re-tagged at their 3’ ends by a terminal-tagging process to produce double-tagged, singlestranded cDNA. In addition, Illumina adaptor sequences are added using limited-cycle PCR
that ultimately results in a directional, amplified library. Finally, the amplified RNA-Seq
library is purified and is further utilized for cluster generation and sequencing.
11.4 Choice of Sequencing Platform
Several NGS platforms (see Chap. 4) have been successfully implemented for RNA-Seq
analysis in the past few years. Currently the three most widely used NGS platforms for
RNA-Seq are the Illumina HiSeq, Ion Torrent, and SOLiD systems [20]. Although the
nucleotide detection methodology varies for each platform, they follow similar library
preparation steps. Either sequencing platform generates between 10 and 100 million reads,
with typical read lengths of 300–500 bp. However, more recent sequencing technologies,
such as Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (MinION) can
sequence full-length transcripts on a transcriptome-wide scale and can produce long reads
ranging between 700 and 2000 bp [21]. Moreover, single-cell RNA-Seq (scRNA-Seq) has
emerged recently as a new tool for precisely performing transcriptomic analysis at the level
of individual cells [22]. Different RNA-Seq experiments require different read lengths and
sequencing depths. Importantly, data generated by different RNA-Seq platforms vary and
might affect the results and interpretations. Thus, the choice of the sequencing platform is
crucial and highly depends on the study design and the objectives. An overview of platform
specific differences is summarized in Table 11.1 and Chap. 4.
11.5 Quality Check (QC) and Sequence Pre-processing
After the sequencing experiment, a quality check (QC) of raw reads is required to filter out
poor-quality reads resulting from errors in library preparation, sequencing errors, PCR artifacts,
untrimmed adapter sequences, and presence of contaminating sequences [24]. Presence of lowquality reads often affect the downstream processing and interpretation of obtained results.
Several tools, including FastQC [25], htSeqTools [26], and SAMStat [27], to assess the quality
11 Design and Analysis of RNA Sequencing Data
147
The basic steps in RNA-Seq library preparation include efficient ribosomal RNA (rRNA)
removal from samples followed by cDNA synthesis for producing directional RNA-Seq
libraries [18, 19]. Few of the widely used RNA-Seq library preparation kits include Illumina
TruSeq (https://www.illumina.com/products/by-type/sequencing-kits/library-prep-kits/
truseq-rna-v2.html), Bioo Scientific NEXTFlex (http://shop.biooscientific.com/nextflexsmall-rna-seq-kit-v3/), and New England Biolabs NEB Next Ultra (https://www.neb.com/
products/e7370-nebnext-ultra-dna-library-prep-kit-for-illumina#Product%20Information).
In general, after removal of rRNA fractions, the remaining RNA is fragmented and reverse
transcribed using random primers with 5’-tagging sequences. Next, the 5’-tagged cDNA are
re-tagged at their 3’ ends by a terminal-tagging process to produce double-tagged, singlestranded cDNA. In addition, Illumina adaptor sequences are added using limited-cycle PCR
that ultimately results in a directional, amplified library. Finally, the amplified RNA-Seq
library is purified and is further utilized for cluster generation and sequencing.
11.4 Choice of Sequencing Platform
Several NGS platforms (see Chap. 4) have been successfully implemented for RNA-Seq
analysis in the past few years. Currently the three most widely used NGS platforms for
RNA-Seq are the Illumina HiSeq, Ion Torrent, and SOLiD systems [20]. Although the
nucleotide detection methodology varies for each platform, they follow similar library
preparation steps. Either sequencing platform generates between 10 and 100 million reads,
with typical read lengths of 300–500 bp. However, more recent sequencing technologies,
such as Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (MinION) can
sequence full-length transcripts on a transcriptome-wide scale and can produce long reads
ranging between 700 and 2000 bp [21]. Moreover, single-cell RNA-Seq (scRNA-Seq) has
emerged recently as a new tool for precisely performing transcriptomic analysis at the level
of individual cells [22]. Different RNA-Seq experiments require different read lengths and
sequencing depths. Importantly, data generated by different RNA-Seq platforms vary and
might affect the results and interpretations. Thus, the choice of the sequencing platform is
crucial and highly depends on the study design and the objectives. An overview of platform
specific differences is summarized in Table 11.1 and Chap. 4.
11.5 Quality Check (QC) and Sequence Pre-processing
After the sequencing experiment, a quality check (QC) of raw reads is required to filter out
poor-quality reads resulting from errors in library preparation, sequencing errors, PCR artifacts,
untrimmed adapter sequences, and presence of contaminating sequences [24]. Presence of lowquality reads often affect the downstream processing and interpretation of obtained results.
Several tools, including FastQC [25], htSeqTools [26], and SAMStat [27], to assess the quality
11 Design and Analysis of RNA Sequencing Data
147
