330
A. Zakharov and A. Lagunin
by the certain computer software and is used for the safety assessment and risk
analysis of industrial compounds, research and development products in the field
of biomedical and toxicological sciences [1]. Therefore, toxicology databases have
to be constantly updated by cheminformatic resources, which are useful for the creation of secondary data sets, e.g. training sets for the (Q)SAR modeling of toxicity.
A huge amount of information from toxicity databases has become freely available recently. It has played an important role in the development of (Q)SAR models
and computer programs for the toxicity prediction. Unfortunately, the content of
freely available databases is still different from compound libraries used for development of drugs and from industrial compounds. Recent initiatives of regulatory
agencies require to develop the toxicology database with a free access and promote the usage of a computer technology [6]. Table 11.1 shows the list of publicly
available toxicity databases, which describe the effect of substances on the human
health, and electronic resources, which are useful for risk assessment and safety of
chemical compounds [1].
The private toxicity databases offer more accurate toxicity data and the extended
chemical space of representative structures in comparison with public databases.
Despite expansion of the chemical space and various numbers of proposed descriptors, the private databases have limitations related to the models selection, types of
algorithms and the content of data, which is probably a part of confidential business information, such as a proprietary structure of pharmaceutical molecules. The
private toxicity databases may also provide internal systems created by industry or
government agencies [21]. These databases may not be suitable for a commercial
usage, but are useful for the internal analysis. Therefore, publication of the scientific research based on these data is often difficult to evaluate independently. The
most known commercial databases associated with toxicity are Accelrys Toxicity
Database (contains information about the structure and different types of toxicity
for more than 150,000 compounds from RTECS and other sources) and Leadscope
Toxicity Database. Standardization of toxicity databases is designed to facilitate
integration between different sources and to provide their quality. Since databases
are often not compatible with each other, standardization initiatives (e.g., controlled
vocabularies) can help to combine their data [22].
11.2.3 Descriptors
Appropriate description of the chemical structure is a major component and limitation for creation of high-quality (Q)SAR models. Molecular descriptors are important for the toxicity modeling technology because their numeric representation is
the basis for construction of structure-activity relationships by computational models. Therefore, if the selected descriptors do not reflect aspects influencing on manifestation of the molecule toxicity, the developed model may show a poor accuracy.
There are both commercial and public software which allow generating different
A. Zakharov and A. Lagunin
by the certain computer software and is used for the safety assessment and risk
analysis of industrial compounds, research and development products in the field
of biomedical and toxicological sciences [1]. Therefore, toxicology databases have
to be constantly updated by cheminformatic resources, which are useful for the creation of secondary data sets, e.g. training sets for the (Q)SAR modeling of toxicity.
A huge amount of information from toxicity databases has become freely available recently. It has played an important role in the development of (Q)SAR models
and computer programs for the toxicity prediction. Unfortunately, the content of
freely available databases is still different from compound libraries used for development of drugs and from industrial compounds. Recent initiatives of regulatory
agencies require to develop the toxicology database with a free access and promote the usage of a computer technology [6]. Table 11.1 shows the list of publicly
available toxicity databases, which describe the effect of substances on the human
health, and electronic resources, which are useful for risk assessment and safety of
chemical compounds [1].
The private toxicity databases offer more accurate toxicity data and the extended
chemical space of representative structures in comparison with public databases.
Despite expansion of the chemical space and various numbers of proposed descriptors, the private databases have limitations related to the models selection, types of
algorithms and the content of data, which is probably a part of confidential business information, such as a proprietary structure of pharmaceutical molecules. The
private toxicity databases may also provide internal systems created by industry or
government agencies [21]. These databases may not be suitable for a commercial
usage, but are useful for the internal analysis. Therefore, publication of the scientific research based on these data is often difficult to evaluate independently. The
most known commercial databases associated with toxicity are Accelrys Toxicity
Database (contains information about the structure and different types of toxicity
for more than 150,000 compounds from RTECS and other sources) and Leadscope
Toxicity Database. Standardization of toxicity databases is designed to facilitate
integration between different sources and to provide their quality. Since databases
are often not compatible with each other, standardization initiatives (e.g., controlled
vocabularies) can help to combine their data [22].
11.2.3 Descriptors
Appropriate description of the chemical structure is a major component and limitation for creation of high-quality (Q)SAR models. Molecular descriptors are important for the toxicity modeling technology because their numeric representation is
the basis for construction of structure-activity relationships by computational models. Therefore, if the selected descriptors do not reflect aspects influencing on manifestation of the molecule toxicity, the developed model may show a poor accuracy.
There are both commercial and public software which allow generating different
