[89]. Normalization protocols together with statistical testing might have the largest impact
on the results of an RNA-Seq analysis [90, 91]. In this context, several normalization
methods have been developed based on different assumptions in RNA-Seq experiments
and corresponding gene expression analysis. There are mainly three major normalization
strategies, that is normalization by library size, normalization by distribution, and normalization by controls [92]. In case of normalization by library size, differences in sequencing
depth are removed by simply dividing by the total number of reads generated for each
sample. Similarly, normalization by distribution involves equilibrating expression levels
for non-DE genes, if the technical effects or treatments remain the same for DE and non-DE
genes. Importantly, assumptions made by any normalization method should be always
considered before choosing a preferable method that suits the study design [93]. For
instance, normalization by library size should be chosen when the total mRNA remains
equal or in the same range across different treatments/conditions despite any asymmetry. In
contrast, normalization by distribution is mostly useful if there is symmetry in sample sizes
even if there are differences in total mRNA across samples.
DE analysis involves analyzing significant differences in the quantitative levels of
transcripts or exons among various samples and conditions. Since gene regulatory mechanisms
often collate multiple genes at a time or under a single condition, obtaining statistically
significant information becomes tricky. This is attributed to quantitative differences in sample
size and expression data as a comparatively large RNA-Seq data could be generated for a
limited set of samples [94, 95]. Under this scenario, DE analysis essentially requires multivariate statistical methods that include principle component analysis (PCA), canonical correlation
analysis (CCA), and nonnegative matrix factorization (NMF) [96]. Thus, statistical software
like R and its bioconductor packages find a high utility in performing DE analysis for RNASeq data. There are several tools including Cuffdiff [81] and bioconductor packages such as
DESeq2 [97], edgeR [98], and Limma [99] that are frequently used for DE analysis.
All three packages utilize linear models for a comparative correlation between gene
expression/transcript abundance with the variables listed by the user. Additionally, these
packages need designated design matrix or experimental information containing outcomes
and parameters together with a contrast matrix that describes the comparisons to be
analyzed. The early steps in the DE analysis essentially require creating the input count
table that could be generated as described below.
BAM Files The alignment files in BAM format are converted to SAM files first (HTSeq)
or can be used directly (BEDTools). Both tools produce a count table that could be used for
DE analysis.
Individual Count Files Individual count files generated using HTSeq can be combined
either using UNIX or R commands directly. In addition, DESeq2 has a separate script for
combining the individual files into single count table. In our RNA-Seq pipeline we use a
custom R script (https://github.com/grimmlab/BookChapter-RNA-Seq-Analyses/blob/
master/Rscripts/DE_deseq_genes_bedtools.R). We first perform a rlog transformations
and then a DE analysis using DESeq2:
166
R. Bharti and D. G. Grimm
Précédent

- 174/225

Suivant