Description of parameters:
-G
path to the annotation file
-b
path to the genome fasta file
-p
number threads
-o
name of the output directory
11.13 Counting Reads Per Exons
In general, a single exon can appear multiple times in a GTF file, due to exon sequence
overlaps associated with transcript isoforms. Thus, abundance estimation for exons
involves cataloguing a set of nonoverlapping exonic regions. For quantifying gene expression using read counts per exon, the Bioconductor package DEXSeq [83] is mainly used.
DEXSeq modifies the input GTF file into a list of exon counting bins that list single
exons or a part of an exon that overlaps. Alignment is performed in SAM format with data
sorted by read names and/or chromosomal coordinates and a modified GTF file with exon
counting bins is generated. The output count file contains the number of reads for every
exon counting bin. Here, a list of non-counted reads is generated based on the criterion that
includes unaligned reads, low-quality alignments, ambiguous or multiple overlaps.
DEXSeq is implemented in the described RNA-Seq analysis workflow in two steps. The
first step is the preparation of annotations:
164
R. Bharti and D. G. Grimm
-G
path to the annotation file
-b
path to the genome fasta file
-p
number threads
-o
name of the output directory
11.13 Counting Reads Per Exons
In general, a single exon can appear multiple times in a GTF file, due to exon sequence
overlaps associated with transcript isoforms. Thus, abundance estimation for exons
involves cataloguing a set of nonoverlapping exonic regions. For quantifying gene expression using read counts per exon, the Bioconductor package DEXSeq [83] is mainly used.
DEXSeq modifies the input GTF file into a list of exon counting bins that list single
exons or a part of an exon that overlaps. Alignment is performed in SAM format with data
sorted by read names and/or chromosomal coordinates and a modified GTF file with exon
counting bins is generated. The output count file contains the number of reads for every
exon counting bin. Here, a list of non-counted reads is generated based on the criterion that
includes unaligned reads, low-quality alignments, ambiguous or multiple overlaps.
DEXSeq is implemented in the described RNA-Seq analysis workflow in two steps. The
first step is the preparation of annotations:
164
R. Bharti and D. G. Grimm
