What You Will Learn in This Chapter
In this chapter, we introduce the concept of RNA-Seq analyses. First, we start to
provide an overview of a typical RNA-Seq experiment that includes extraction of
sample RNA, enrichment, and cDNA library preparation. Next, we review tools for
quality control and data pre-processing followed by a standard workflow to perform
RNA-Seq analyses. For this purpose, we discuss two common RNA-Seq strategies,
that is a reference-based alignment and a de novo assembly approach. We learn how
to do basic downstream analyses of RNA-Seq data, including quantification of
expressed genes, differential gene expression (DE) between different groups as
well as functional gene analysis. Eventually, we provide a best-practice example
for a reference-based RNA-Seq analysis from beginning to end, including all necessary tools and steps on GitHub: https://github.com/grimmlab/BookChapter-RNASeq-Analyses.
11.1 Introduction
The central dogma of molecular biology integrates the flow of information encoded in
DNA via transcription into RNA molecules that eventually translates into proteins inside a
cell (Fig. 11.1, also see Chap. 1, Section 0). Any alteration, e.g. due to genetic, lifestyle, or
environmental factors, might change the phenotype of an organism [1]. These alterations, e.
g. copy number variations or mutational modifications of RNA molecules, affect the
regulation of biological activities within individual cells [2, 3]. The entirety of all coding
and non-coding RNAs derived from a cell at a certain point in time is referred to as the
transcriptome [4]. Apparently, any change in the transcriptome culminates into functional
alterations at both cellular and organismic level. Therefore, quantifying transcriptome
variations and/or gene expression profiling remains crucial for understanding phenotypic
alterations associated with disease and development [5, 6].
In the past, quantitative polymerase chain reaction (qPCR) was used as the tool of
choice for quantifying transcripts and for performing gene expression analyses. Although
qPCR remains a cheap and accurate technique for analyzing small sets of genes or groups
of genes, it fails to scale to genome-wide level [7]. The introduction of DNA helped to scale
transcriptomic studies to a genome-wide level due to their ability to accurately analyze
thousands of transcripts at low cost [8, 9]. However, requirements of a priori knowledge of
genome sequence, cross-hybridization errors, presence of artifacts, and the inability to
analyze alternate splicing and non-coding RNAs limit the usage of microarrays [10, 11].
Currently, next-generation sequencing (NGS) has revolutionized the transcriptomic analysis landscape due to higher coverage, detection of low abundance and novel transcripts,
144
R. Bharti and D. G. Grimm
In this chapter, we introduce the concept of RNA-Seq analyses. First, we start to
provide an overview of a typical RNA-Seq experiment that includes extraction of
sample RNA, enrichment, and cDNA library preparation. Next, we review tools for
quality control and data pre-processing followed by a standard workflow to perform
RNA-Seq analyses. For this purpose, we discuss two common RNA-Seq strategies,
that is a reference-based alignment and a de novo assembly approach. We learn how
to do basic downstream analyses of RNA-Seq data, including quantification of
expressed genes, differential gene expression (DE) between different groups as
well as functional gene analysis. Eventually, we provide a best-practice example
for a reference-based RNA-Seq analysis from beginning to end, including all necessary tools and steps on GitHub: https://github.com/grimmlab/BookChapter-RNASeq-Analyses.
11.1 Introduction
The central dogma of molecular biology integrates the flow of information encoded in
DNA via transcription into RNA molecules that eventually translates into proteins inside a
cell (Fig. 11.1, also see Chap. 1, Section 0). Any alteration, e.g. due to genetic, lifestyle, or
environmental factors, might change the phenotype of an organism [1]. These alterations, e.
g. copy number variations or mutational modifications of RNA molecules, affect the
regulation of biological activities within individual cells [2, 3]. The entirety of all coding
and non-coding RNAs derived from a cell at a certain point in time is referred to as the
transcriptome [4]. Apparently, any change in the transcriptome culminates into functional
alterations at both cellular and organismic level. Therefore, quantifying transcriptome
variations and/or gene expression profiling remains crucial for understanding phenotypic
alterations associated with disease and development [5, 6].
In the past, quantitative polymerase chain reaction (qPCR) was used as the tool of
choice for quantifying transcripts and for performing gene expression analyses. Although
qPCR remains a cheap and accurate technique for analyzing small sets of genes or groups
of genes, it fails to scale to genome-wide level [7]. The introduction of DNA helped to scale
transcriptomic studies to a genome-wide level due to their ability to accurately analyze
thousands of transcripts at low cost [8, 9]. However, requirements of a priori knowledge of
genome sequence, cross-hybridization errors, presence of artifacts, and the inability to
analyze alternate splicing and non-coding RNAs limit the usage of microarrays [10, 11].
Currently, next-generation sequencing (NGS) has revolutionized the transcriptomic analysis landscape due to higher coverage, detection of low abundance and novel transcripts,
144
R. Bharti and D. G. Grimm
