• Choice of the RNA extraction and library preparation methods, NGS platform,
quality check, and pre-processing protocols affect the outcomes of RNA-Seq
experiments.
• RNA-Seq data processing highly depends on the presence or absence of a
reference genome that remains the pivotal determinant of the choice of a reference-based alignment or a de novo assembly for a given dataset.
• We provide a best-practice RNA-Seq analysis workflow on GitHub: https://
github.com/grimmlab/BookChapter-RNA-Seq-Analyses. This workflow is a stepwise reference-based RNA-Seq analysis example and shall help to get started with
a basic RNA-Seq analysis.
Further Reading
• The Biostar Handbook: 2nd Edition.
https://www.biostarhandbook.com/
• GNU Bash Reference Manual by Chet Ramey, Brian Fox.
https://www.gnu.org/software/bash/manual/bash.pdf
• An Introduction to Statistical Learning with Applications in R by Gareth James, Daniela
Witten, Trevor Hastie, and Robert Tibshirani.
http://faculty.marshall.usc.edu/gareth-james/ISL/
• Python for Bioinformatics by Sebastian Bassi.
• Next-generation transcriptome assembly. Martin JA, Wang Z. Nat Rev Genet. 2011 Sep
7;12(10):671-82.
• RNA-Seq: a revolutionary tool for transcriptomics. Wang Z, Gerstein M, Snyder M. Nat
Rev Genet. 2009 Jan;10(1):57–63.
• Computational and analytical challenges in single-cell transcriptomics. Stegle O,
Teichmann SA, Marioni JC. Nat Rev Genet. 2015 Mar;16(3):133–45.
• Genetic Variation and the De Novo Assembly of Human Genomes, Chaisson MJP,
Wilson RK, Eichler EE, Nat Rev Genet. 2015 Nov; 16(11):627–40.
Answers to Review Questions
Answer to Questions 1: During the read alignment the first crucial step is to determine
the point of origin of the read sequence with respect to the reference genome. For an
accurate mapping it is important to consider that reads might originate from either
cDNA corresponding to spliced transcripts or from non-spliced transcripts. In the case
of spliced transcripts, contiguous read sequences are separated by intervening splice
junction (SJ) boundaries and are split into two fragments and assigned separately. Thus,
there are two main alignment approaches for aligning reads against a reference genome,
that is a spliced unaware and spliced aware alignment. For this purpose, the RNA-Seq
reads are typically mapped to either a genome (splice aware) or a transcriptome (splice
unaware). Most of the splice unaware alignment tools align DNA against DNA and
11 Design and Analysis of RNA Sequencing Data
169
quality check, and pre-processing protocols affect the outcomes of RNA-Seq
experiments.
• RNA-Seq data processing highly depends on the presence or absence of a
reference genome that remains the pivotal determinant of the choice of a reference-based alignment or a de novo assembly for a given dataset.
• We provide a best-practice RNA-Seq analysis workflow on GitHub: https://
github.com/grimmlab/BookChapter-RNA-Seq-Analyses. This workflow is a stepwise reference-based RNA-Seq analysis example and shall help to get started with
a basic RNA-Seq analysis.
Further Reading
• The Biostar Handbook: 2nd Edition.
https://www.biostarhandbook.com/
• GNU Bash Reference Manual by Chet Ramey, Brian Fox.
https://www.gnu.org/software/bash/manual/bash.pdf
• An Introduction to Statistical Learning with Applications in R by Gareth James, Daniela
Witten, Trevor Hastie, and Robert Tibshirani.
http://faculty.marshall.usc.edu/gareth-james/ISL/
• Python for Bioinformatics by Sebastian Bassi.
• Next-generation transcriptome assembly. Martin JA, Wang Z. Nat Rev Genet. 2011 Sep
7;12(10):671-82.
• RNA-Seq: a revolutionary tool for transcriptomics. Wang Z, Gerstein M, Snyder M. Nat
Rev Genet. 2009 Jan;10(1):57–63.
• Computational and analytical challenges in single-cell transcriptomics. Stegle O,
Teichmann SA, Marioni JC. Nat Rev Genet. 2015 Mar;16(3):133–45.
• Genetic Variation and the De Novo Assembly of Human Genomes, Chaisson MJP,
Wilson RK, Eichler EE, Nat Rev Genet. 2015 Nov; 16(11):627–40.
Answers to Review Questions
Answer to Questions 1: During the read alignment the first crucial step is to determine
the point of origin of the read sequence with respect to the reference genome. For an
accurate mapping it is important to consider that reads might originate from either
cDNA corresponding to spliced transcripts or from non-spliced transcripts. In the case
of spliced transcripts, contiguous read sequences are separated by intervening splice
junction (SJ) boundaries and are split into two fragments and assigned separately. Thus,
there are two main alignment approaches for aligning reads against a reference genome,
that is a spliced unaware and spliced aware alignment. For this purpose, the RNA-Seq
reads are typically mapped to either a genome (splice aware) or a transcriptome (splice
unaware). Most of the splice unaware alignment tools align DNA against DNA and
11 Design and Analysis of RNA Sequencing Data
169
