146
D. L. Svoboda et al.
Chemical and drug structures are the data type of greatest abundance housed in the
DrugMatrix database. Approximately, 8000 chemical structures are curated in the
database (ftp://anonftp.niehs.nih.gov/ntp-cebs/datatype/Drug_Matrix/Drugmatrix_
Curated_Chemicals_With_SMILES.txt), with ~2000 of these having some degree
of baseline curation and ~800 with full curation. Curation is defined as a collection
of facts (e.g., pharmacokinetics, toxicity, pharmacology) describing each compound,
recorded in the literature. An integral part of the curation process was the creation of
an ontology consistent with terms and logic used in the field of toxicology describing chemical and biological properties. Chemical mapping to ontology terms can
be found elsewhere (ftp://anonftp.niehs.nih.gov/ntp-cebs/datatype/Drug_Matrix/
Compound%20Literature%20Annotations/COMPOUND_ANNOTATIONS.txt).
To allow users of DrugMatrix to explore associations between gene expression
and effects of drugs and chemicals, a select set of 130 commercially available pharmacological targets were screened using competitive binding assays. The results of
these studies can be found elsewhere (ftp://anonftp.niehs.nih.gov/ntp-cebs/datatype/
Drug_Matrix/DM_invitro_assay_data.xlsx). These data have also been integrated in
number of public HTS databases including ChEMBL and Pubchem (search “DrugMatrix” in these databases to find results).
Distinct species can respond to a chemical challenge in different ways. Most public
resources have focused on pathway curation in human or mouse (most common nonhuman model system for academic studies). Hence, to most effectively interpret rat
toxicogenomics data, 137 pathways were curated focusing on a rat biology review
of the literature. Detailed citations are provided for the inclusion of genes/proteins,
and their linkage to other components of the pathways. All the curated pathways are
available for download (ftp://anonftp.niehs.nih.gov/ntp-cebs/datatype/Drug_Matrix/
DrugMatrix%20Pathways.zip).
One of the primary goals of acquiring the assets associated with DrugMatrix was
to make all data freely available to the research community for mining from a variety
of perspectives, using a diversity of computational and bioinformatic approaches. All
data resources noted above in addition to additional DrugMatrix data, resources, and
information can be found elsewhere (ftp://anonftp.niehs.nih.gov/ntp-cebs/datatype/
Drug_Matrix/).
8.3 DrugMatrix Database
A highly-integrated relational database holds the algorithmically extracted experimental data described above (Fig. 8.3). A detailed description of the database
can be found in CEBS (ftp://anonftp.niehs.nih.gov/ntp-cebs/datatype/Drug_Matrix/
DrugMatrixDataWarehouse.pdf). The DrugMatrix data warehouse architecture is a
highly denormalized, modified star-schema. The “hubs” of the schema are the six
main information domains, or schema dimensions: GENE, COMPOUND, EXPRESSION EXPERIMENT, EXPRESSION STUDY, PATHWAY, and ASSAY (Table 8.1).
These hubs represent the main information domains in the DrugMatrix user interface.
Précédent

- 158/416

Suivant