306
D. Wang
as RNAseq results in read count data rather than fluorescence intensity, different
processing and normalization procedures will be required relative to microarrays.
From studies done so far, RNAseq data can be used to model transcription-based
PODs (e.g., [23]). Though the BMD value can be quite different between whether
it is based on RNAseq or microarrays at the gene and pathway level, the PODs
from most sensitive pathways are usually consistent [24]. As more studies based on
RNAseq become available, we should have deeper appreciation of the strengths and
special characteristics for RNAseq in the context of POD modeling. In this review,
we will focus more on microarray data due to its usage in the majority of studies.
We will mention specific aspects of RNAseq data when needed.
Data processing. As common for high-throughput data analysis, extensive preprocessing for both microarray and RNAseq data is usually needed before data can
be used for BMD modeling. As mentioned, Affymetrix GeneChip is the platform
of choice for most dose response microarray experiments. It is standard practice to
transform the intensity value for each probe to log 2 scale and perform RMA normalization. RMA has been reported to remove nonbiological effects between microarrays. Though other normalization methods exist for microarrays, they have not been
extensively applied in the transcriptional modeling of PODs. For RNAseq data, the
choice is not as clear. RNAseq experiments produce data as counts of short reads for
each gene (after alignment and mapping). Commonly used software for differential
expression analysis using RNAseq often models the count data directly using Poisson
or negative binomial distributions. Popular software packages include edgeR [25],
DESeq [26], and others. For BMD modeling, however, it is most convenient if the data
can be analyzed by BMDExpress [18], which we will discuss later. As BMDExpress
was originally developed for microarray data, this means transforming the RNASeq
read count to a continuous variable, similar to microarray intensity measurement.
One commonly used approach is to calculate reads per kilobase per million mapped
reads (RPKM) followed by log 2 transformation. Other methods like kernel density
mean of M-component (KDMM) [24] have also been recommended. With more dose
response experiments being carried out with the RNAseq technology, we should gain
a better understanding of the appropriate data processing approaches in this setting.
Once the transformation has been performed, the quantities like log-transformed
RPKM are generally treated the same way as microarray intensity readings in dose
response modeling.
Modeling BMDs (BMDLs) at the gene level. For each gene, the normalized expression level from microarrays or RNAseq can be used to derive a gene-level transcriptional BMD (BMDL) as other continuous endpoints like organ weight. A general
discussion on benchmark calculations for continuous data can be found in [27].
The BMD is defined as the dose where the fitted dose response curve exceeds the
background noise level seen in control samples. As there are usually only a limited
number of doses tested, often several different models can fit the data reasonably
well. Commonly used models are linear, polynomial models of different degrees,
exponential, and Hill models. The preferred model can be selected using statistical
criteria, while some researchers advocate averaging results from multiple models.
Given that there are thousands of genes, performing model fitting and subsequent
D. Wang
as RNAseq results in read count data rather than fluorescence intensity, different
processing and normalization procedures will be required relative to microarrays.
From studies done so far, RNAseq data can be used to model transcription-based
PODs (e.g., [23]). Though the BMD value can be quite different between whether
it is based on RNAseq or microarrays at the gene and pathway level, the PODs
from most sensitive pathways are usually consistent [24]. As more studies based on
RNAseq become available, we should have deeper appreciation of the strengths and
special characteristics for RNAseq in the context of POD modeling. In this review,
we will focus more on microarray data due to its usage in the majority of studies.
We will mention specific aspects of RNAseq data when needed.
Data processing. As common for high-throughput data analysis, extensive preprocessing for both microarray and RNAseq data is usually needed before data can
be used for BMD modeling. As mentioned, Affymetrix GeneChip is the platform
of choice for most dose response microarray experiments. It is standard practice to
transform the intensity value for each probe to log 2 scale and perform RMA normalization. RMA has been reported to remove nonbiological effects between microarrays. Though other normalization methods exist for microarrays, they have not been
extensively applied in the transcriptional modeling of PODs. For RNAseq data, the
choice is not as clear. RNAseq experiments produce data as counts of short reads for
each gene (after alignment and mapping). Commonly used software for differential
expression analysis using RNAseq often models the count data directly using Poisson
or negative binomial distributions. Popular software packages include edgeR [25],
DESeq [26], and others. For BMD modeling, however, it is most convenient if the data
can be analyzed by BMDExpress [18], which we will discuss later. As BMDExpress
was originally developed for microarray data, this means transforming the RNASeq
read count to a continuous variable, similar to microarray intensity measurement.
One commonly used approach is to calculate reads per kilobase per million mapped
reads (RPKM) followed by log 2 transformation. Other methods like kernel density
mean of M-component (KDMM) [24] have also been recommended. With more dose
response experiments being carried out with the RNAseq technology, we should gain
a better understanding of the appropriate data processing approaches in this setting.
Once the transformation has been performed, the quantities like log-transformed
RPKM are generally treated the same way as microarray intensity readings in dose
response modeling.
Modeling BMDs (BMDLs) at the gene level. For each gene, the normalized expression level from microarrays or RNAseq can be used to derive a gene-level transcriptional BMD (BMDL) as other continuous endpoints like organ weight. A general
discussion on benchmark calculations for continuous data can be found in [27].
The BMD is defined as the dose where the fitted dose response curve exceeds the
background noise level seen in control samples. As there are usually only a limited
number of doses tested, often several different models can fit the data reasonably
well. Commonly used models are linear, polynomial models of different degrees,
exponential, and Hill models. The preferred model can be selected using statistical
criteria, while some researchers advocate averaging results from multiple models.
Given that there are thousands of genes, performing model fitting and subsequent
