18 Integrating QSAR, Read-Across, and Screening Tools …
371
Using this similarity measurement, VEGA shows the six most similar compounds
to the target chemical found in the training and test set together with their property
values. Thus, VEGA immediately provides a simple way to make read-across.
The similarity search is also used to evaluate the reliability of the model prediction.
VEGA applies a separate software for this, different from the QSAR model used
to make the prediction. VEGA performs a sophisticated procedure to evaluate the
reliability of the model, which refers to the applicability domain of the model. This
assessment is quantitative and is provided through the applicability domain index
(ADI). The ADI is a single overall measurement. ADI combined several evaluations
depending on the model, since some models are quantitative (QSAR) and others are
only categorical (SAR). Below we list the components of the overall assessment.
• Similarity of the most similar compounds. This parameter checks how similar are
the most similar compounds.
• Chemometric check of the descriptor space. This verifies if the target compound is
inside the range of the descriptors used within the model. In addition, the software
checks if the molecular weight of the target compound is within the range of the
molecular weights of the chemicals at the basis of the model.
• Check of the descriptor sensitivity. The software artificially modifies of up to
10% the values of the molecular descriptors and verifies if dramatic changes are
recorded, indicating an unstable situation.
• Check for outliers based on specific fragments. For some models, outliers have
been identified, which share a common fragment.
• Identification of the presence of rare fragments. This is done using an atomcentered fragment tool.
• Accuracy of the prediction. In this case, the software looks for the three most
similar compounds and compares the predicted and experimental values for these
compounds. This provides an evaluation on how well the model behaves for chemicals, which are close to the target compound.
• Concordance with the experimental values of the similar compounds and the predicted value of the target compound. This checks if there is agreement between
the read-across and the QSAR prediction.
18.2.3 The ToxRead Software
In addition to the use of VEGA for read-across, as described above, there is a specific
program in VEGAHUB, which is devoted to read-across: ToxRead [16]. ToxRead
works by applying the same software for similarity as used in VEGA. Again, the
most similar compounds are shown, as in VEGA. However, in this case the user can
choose the number of the most similar compounds. Furthermore, the library of the
similar compounds within ToxRead is not limited to one single model, but integrates
the different collections of chemicals for the same endpoint. The real novelty of
ToxRead is that, in addition to the most similar compounds, the software provides
371
Using this similarity measurement, VEGA shows the six most similar compounds
to the target chemical found in the training and test set together with their property
values. Thus, VEGA immediately provides a simple way to make read-across.
The similarity search is also used to evaluate the reliability of the model prediction.
VEGA applies a separate software for this, different from the QSAR model used
to make the prediction. VEGA performs a sophisticated procedure to evaluate the
reliability of the model, which refers to the applicability domain of the model. This
assessment is quantitative and is provided through the applicability domain index
(ADI). The ADI is a single overall measurement. ADI combined several evaluations
depending on the model, since some models are quantitative (QSAR) and others are
only categorical (SAR). Below we list the components of the overall assessment.
• Similarity of the most similar compounds. This parameter checks how similar are
the most similar compounds.
• Chemometric check of the descriptor space. This verifies if the target compound is
inside the range of the descriptors used within the model. In addition, the software
checks if the molecular weight of the target compound is within the range of the
molecular weights of the chemicals at the basis of the model.
• Check of the descriptor sensitivity. The software artificially modifies of up to
10% the values of the molecular descriptors and verifies if dramatic changes are
recorded, indicating an unstable situation.
• Check for outliers based on specific fragments. For some models, outliers have
been identified, which share a common fragment.
• Identification of the presence of rare fragments. This is done using an atomcentered fragment tool.
• Accuracy of the prediction. In this case, the software looks for the three most
similar compounds and compares the predicted and experimental values for these
compounds. This provides an evaluation on how well the model behaves for chemicals, which are close to the target compound.
• Concordance with the experimental values of the similar compounds and the predicted value of the target compound. This checks if there is agreement between
the read-across and the QSAR prediction.
18.2.3 The ToxRead Software
In addition to the use of VEGA for read-across, as described above, there is a specific
program in VEGAHUB, which is devoted to read-across: ToxRead [16]. ToxRead
works by applying the same software for similarity as used in VEGA. Again, the
most similar compounds are shown, as in VEGA. However, in this case the user can
choose the number of the most similar compounds. Furthermore, the library of the
similar compounds within ToxRead is not limited to one single model, but integrates
the different collections of chemicals for the same endpoint. The real novelty of
ToxRead is that, in addition to the most similar compounds, the software provides
