would have to introduce a long gap in the mapping of a read to span an intron of varying
length. Thus, splice unaware aligners are commonly used if reads can be directly
mapped against a reference transcriptome or prokaryotic genome. An example of an
ungapped aligner is Bowtie, which is an alignment tool based on Burrows–Wheeler
Transforms (BWT). Note that this strategy will not support the discovery of novel
transcripts. Especially transcriptomes of alternatively spliced organisms, like
eukaryotes, are generally suitable for most reference-genome based alignments. As
the information on SJs in the reference genome is crucial, most aligners also refer to
known annotated SJ databases to confirm the presence or absence of splice sites. Spliceaware aligner like STAR incorporates a three-step alignment procedure to map the
reads. In the first step, sequencing reads are aligned to a reference genome. Next,
overlapping reads present at each locus are collated into a graph representing all
isoforms. In the final step, the graph is resolved, and all isoforms associated with
individual genes are identified.
Answer to Questions 2: Gene expression quantification can be done by either
counting reads per genes, transcripts, or exons. In the first case, gene expression is
quantified by counting the number of reads that align to each gene in the RNA-Seq data.
In the second case, quantification based on counting reads per transcript involves precise
estimation of transcript abundances followed by an expectation maximization (EM)
approach for abundance-based assignments of reads to different transcripts. In the third
case, gene expression quantification using exons involves identification and indexing
sets of nonoverlapping exonic regions. The choice of the correct quantification depends
on the study design and will have a high impact on the outcome of an experiment.
Answer to Question 3: The aim of differential expression analysis is to identify genes
with difference in expression patterns under different experimental conditions (gene(s)
versus condition(s) and vice versa) for a set of samples.
Acknowledgements We are grateful to Dr. Philipp Torkler (Senior Bioinformatics Scientist,
Exosome Diagnostics, a Bio-Techne brand, Munich, Germany) for critically reading this text. We
thank for correcting our mistakes and suggesting relevant improvements to the original manuscript.
References
1. Crick F. Central dogma of molecular biology. Nature. 1970;227(5258):561–3.
2. de Smith AJ, Walters RG, Froguel P, Blakemore AI. Human genes involved in copy number
variation: mechanisms of origin, functional effects and implications for disease. Cytogenet
Genome Res. 2008;123(1-4):17–26.
3. Jonkhout N, Tran J, Smith MA, Schonrock N, Mattick JS, Novoa EM. The RNA modification
landscape in human disease. RNA. 2017;23(12):1754–69.
4. Martin JA, Wang Z. Next-generation transcriptome assembly. Nat Rev Genet. 2011;12
(10):671–82.
170
R. Bharti and D. G. Grimm
length. Thus, splice unaware aligners are commonly used if reads can be directly
mapped against a reference transcriptome or prokaryotic genome. An example of an
ungapped aligner is Bowtie, which is an alignment tool based on Burrows–Wheeler
Transforms (BWT). Note that this strategy will not support the discovery of novel
transcripts. Especially transcriptomes of alternatively spliced organisms, like
eukaryotes, are generally suitable for most reference-genome based alignments. As
the information on SJs in the reference genome is crucial, most aligners also refer to
known annotated SJ databases to confirm the presence or absence of splice sites. Spliceaware aligner like STAR incorporates a three-step alignment procedure to map the
reads. In the first step, sequencing reads are aligned to a reference genome. Next,
overlapping reads present at each locus are collated into a graph representing all
isoforms. In the final step, the graph is resolved, and all isoforms associated with
individual genes are identified.
Answer to Questions 2: Gene expression quantification can be done by either
counting reads per genes, transcripts, or exons. In the first case, gene expression is
quantified by counting the number of reads that align to each gene in the RNA-Seq data.
In the second case, quantification based on counting reads per transcript involves precise
estimation of transcript abundances followed by an expectation maximization (EM)
approach for abundance-based assignments of reads to different transcripts. In the third
case, gene expression quantification using exons involves identification and indexing
sets of nonoverlapping exonic regions. The choice of the correct quantification depends
on the study design and will have a high impact on the outcome of an experiment.
Answer to Question 3: The aim of differential expression analysis is to identify genes
with difference in expression patterns under different experimental conditions (gene(s)
versus condition(s) and vice versa) for a set of samples.
Acknowledgements We are grateful to Dr. Philipp Torkler (Senior Bioinformatics Scientist,
Exosome Diagnostics, a Bio-Techne brand, Munich, Germany) for critically reading this text. We
thank for correcting our mistakes and suggesting relevant improvements to the original manuscript.
References
1. Crick F. Central dogma of molecular biology. Nature. 1970;227(5258):561–3.
2. de Smith AJ, Walters RG, Froguel P, Blakemore AI. Human genes involved in copy number
variation: mechanisms of origin, functional effects and implications for disease. Cytogenet
Genome Res. 2008;123(1-4):17–26.
3. Jonkhout N, Tran J, Smith MA, Schonrock N, Mattick JS, Novoa EM. The RNA modification
landscape in human disease. RNA. 2017;23(12):1754–69.
4. Martin JA, Wang Z. Next-generation transcriptome assembly. Nat Rev Genet. 2011;12
(10):671–82.
170
R. Bharti and D. G. Grimm
