15 In Silico Prediction of the Point of Departure (POD) …
307
pathway analysis can require significant efforts. Fortunately, this difficulty is alleviated by the availability of BMDExpress [18], which automated much of the work of
gene-level modeling and subsequent analysis.
Since its introduction, BMDExpress has been widely adopted for modeling transcriptional BMDs. With further development by a team led by Scott S Auerbach of
National Toxicology Program, BMDExpress 2.0 has been released with enhancements (https://github.com/auerbachs/BMDExpress-2/wiki). The work of Thomas
et al. [17] described above provides a typical workflow of using BMDExpress. The
user has the option to filter genes according to some criteria of differential expression
(e.g., p-values from ANOVA). If this is to be done, the threshold should be sufficiently
low in order not to exclude genes with real signal. The BMD is usually set as the dose
where response exceeds the baseline level by a value of some constant times the standard deviation (s.d.) from the control samples. Usually, 1 s.d. or 1.349 s.d. are used,
see [27] for explanation with regard to the relationship with benchmark risk (BMR).
BMDExpress can fit several commonly used dose response models. Especially with
the version 2.0, users can choose from linear, 2–4-degree polynomial, 2–5-parameter
exponential, power, and Hill models. If only nested models are chosen (e.g., linear
and polynomial models), the likelihood ratio test can be used to choose the preferred
model; otherwise, the model with the lowest AIC is preferable. Once the preferred
model is chosen, the probeset-wise BMD (BMDL) can be determined.
As BMDExpress was developed for microarray data, the probeset level results
can be automatically mapped to gene identifiers. For RNAseq data, one will need to
obtain read counts at the gene level and perform normalization as discussed before.
Then, data can be fed into BMDExpress. Though good results have been reported
using this approach (e.g., [23]), Black et al. [24] note different behaviors for genes
with high expression and low expression levels. This is likely due to the well-known
mean–variance relationship of read count data (genes with high mean expression levels tend to have high variance). Going forward, it might be sensible to incorporate the
voom approach (in the limma Bioconductor package, [28]) in dose response modeling by introducing precision weights based on the predicted variance. Of course, a
significant amount of work will be involved to produce a software package for easy
practical use.
BMD modeling at the pathway or gene category level. The expression level for
a single gene is notoriously noisy. Thus, integrating information from a biological
pathway or gene category is more preferable in deriving the PODs with transcription data. This obviously leads to the question of how to define pathways or gene
categories. Currently, there are a number of sources providing extensive listings of
biological pathways or gene categories, mostly compiled through literature reviews
to link genes with biological functions. BMDExpress 2.0 provides analysis with Gene
Ontology (http://www.geneontology.org/), Reactome (http://www.reactome.org/), or
categories defined by the user. In published studies, Ingenuity Pathways (QIAGEN),
the Molecular Signature Database (MSigDB) [29], and other sources have also been
used. Usually, only pathways with a certain number of genes (e.g., >5) with BMDs
meeting quality requirements will be analyzed. Though it is common to use the
Précédent

- 315/416

Suivant