the experimental design, gene expression quantification can be done by either counting
reads per genes, transcripts, or exons to obtain error-free estimations [78].
In our RNA-Seq workflow, quantifications are generated from a reference-based RNASeq alignment using the following shell script (see GitHub repository):
Following this, one of the three quantification methods can be chosen as follows:
11.11 Counting Reads Per Genes
The simplest method to quantify gene expression is to count the number of reads that align
to each gene in the RNA-Seq data. Various tools such as HTSeq [79] and BEDTools [80]
are available that can count reads per gene. The workflow on GitHub contains reads per
gene counts for both HTSeq and BEDTools. Note, we will only use HTSeq for the
following explanations.
HTSeq—is a Python-based tool that has several pre-compiled scripts for analyzing NGS
data. The HTSeq-count script takes as input the genomic read alignments in SAM/BAM
format together with genome annotations in GFF/GTF format. The algorithm matches exon
locations listed in GFF/GTF file and counts reads mapped to that location. The output is
generated by clustering the total number of exons for each gene along with information on
number of reads that were left-out. The criterion for left-out or unmapped reads include
alignments at multiple locations, low alignment qualities, ambiguous alignments, and no
alignments or overlaps. This step is implemented in the described RNA-Seq analysis as
follows:
162
R. Bharti and D. G. Grimm
reads per genes, transcripts, or exons to obtain error-free estimations [78].
In our RNA-Seq workflow, quantifications are generated from a reference-based RNASeq alignment using the following shell script (see GitHub repository):
Following this, one of the three quantification methods can be chosen as follows:
11.11 Counting Reads Per Genes
The simplest method to quantify gene expression is to count the number of reads that align
to each gene in the RNA-Seq data. Various tools such as HTSeq [79] and BEDTools [80]
are available that can count reads per gene. The workflow on GitHub contains reads per
gene counts for both HTSeq and BEDTools. Note, we will only use HTSeq for the
following explanations.
HTSeq—is a Python-based tool that has several pre-compiled scripts for analyzing NGS
data. The HTSeq-count script takes as input the genomic read alignments in SAM/BAM
format together with genome annotations in GFF/GTF format. The algorithm matches exon
locations listed in GFF/GTF file and counts reads mapped to that location. The output is
generated by clustering the total number of exons for each gene along with information on
number of reads that were left-out. The criterion for left-out or unmapped reads include
alignments at multiple locations, low alignment qualities, ambiguous alignments, and no
alignments or overlaps. This step is implemented in the described RNA-Seq analysis as
follows:
162
R. Bharti and D. G. Grimm
