134
G. Idakwo et al.
altering the physical meaning of the features. More analysis may be required to interpret the selected features. The stability of selected features and proper feature subset
validation methods are often overlooked. Feature selection bias can be avoided by
embedding the feature selection process within the inner loop of a cross-validation
process to avoid an overly optimistic performance value. Although dimensionality reduction has been shown to improve model performance, there is still room
for improvement when it comes to evaluating and validating feature selection and
extraction methods and their stability. For the sake of reproducibility, researchers
are encouraged to publish important parameters for feature selection or extraction
methods they employed, such as the threshold for a variance score. Regardless of
the choice of features (molecular descriptors, fingerprints or a combination) used for
modeling, SAR models can benefit from dimensionality reduction techniques.
References
1. Lavecchia A (2015) Machine-learning approaches in drug discovery: methods and applications.
Drug Discov Today 20(3):318–331
2. Raies AB, Bajic VB (2016) In silico toxicology: computational methods for the prediction of
chemical toxicity. Wiley Interdiscip Rev Comput Mol Sci 6(2):147–172
3. Greene N, Pennie W (2015) Computational toxicology, friend or foe? Toxicol Res
4(5):1159–1172
4. Kruhlak NL, Benz RD, Zhou H, Colatsky TJ (2012) (Q)SAR modeling and safety assessment
in regulatory review. Clin Pharmacol Ther 91(3):529–534
5. Tropsha A (2010) Best practices for QSAR model development, validation, and exploitation.
Mol Inform 29(6–7):476–488
6. Yang H, Sun L, Li W, Liu G, Tang Y (2018) In silico prediction of chemical toxicity for drug
design using machine learning methods and structural alerts. Front Chem 6:30. https://doi.org/
10.3389/fchem.2018.00030
7. Danishuddin Khan AU (2016) Descriptors and their selection methods in QSAR analysis:
paradigm for drug design. Drug Discov Today 21(8):1291–1302
8. Leach AR, Gillet VJ (2007) Molecular descriptors. An introduction to chemoinformatics.
Springer, Dordrecht, pp 53–74
9. Todeschini R, Consonni V (2000) Handbook of molecular descriptors. Wiley-VCH, Weinheim
10. Duan J, Dixon SL, Lowrie JF, Sherman W (2010) Analysis and comparison of 2D fingerprints:
insights into database screening performance using eight fingerprint methods. J Mol Graph
Model 29(2):157–170
11. National
Institutes
of
Health
(2009)
PubChem
substructure
fingerprint.
ftp://ftp.ncbi.nlm.nih.gov/pubchem/specifications/pubchem_fingerprints.txt. Accessed 10
Oct 2018
12. Rogers D, Hahn M (2010) Extended-connectivity fingerprints. J Chem Inf Model
50(5):742–754
13. Huang R, Xia M, Nguyen D-T et al (2016) Tox21Challenge to build predictive models of nuclear
receptor and stress response pathways as mediated by exposure to environmental chemicals
and drugs. Front Environ Sci 3:85. https://doi.org/10.3389/fenvs.2015.00085
14. Mayr A, Klambauer G, Unterthiner T, Hochreiter S (2016) DeepTox: toxicity prediction using
deep learning. Front Environ Sci 3:80. https://doi.org/10.3389/fenvs.2015.00080
15. Subramanian J, Simon R (2013) Overfitting in prediction models—Is it a problem only in high
dimensions? Contemp Clin Trials 36(2):636–641
G. Idakwo et al.
altering the physical meaning of the features. More analysis may be required to interpret the selected features. The stability of selected features and proper feature subset
validation methods are often overlooked. Feature selection bias can be avoided by
embedding the feature selection process within the inner loop of a cross-validation
process to avoid an overly optimistic performance value. Although dimensionality reduction has been shown to improve model performance, there is still room
for improvement when it comes to evaluating and validating feature selection and
extraction methods and their stability. For the sake of reproducibility, researchers
are encouraged to publish important parameters for feature selection or extraction
methods they employed, such as the threshold for a variance score. Regardless of
the choice of features (molecular descriptors, fingerprints or a combination) used for
modeling, SAR models can benefit from dimensionality reduction techniques.
References
1. Lavecchia A (2015) Machine-learning approaches in drug discovery: methods and applications.
Drug Discov Today 20(3):318–331
2. Raies AB, Bajic VB (2016) In silico toxicology: computational methods for the prediction of
chemical toxicity. Wiley Interdiscip Rev Comput Mol Sci 6(2):147–172
3. Greene N, Pennie W (2015) Computational toxicology, friend or foe? Toxicol Res
4(5):1159–1172
4. Kruhlak NL, Benz RD, Zhou H, Colatsky TJ (2012) (Q)SAR modeling and safety assessment
in regulatory review. Clin Pharmacol Ther 91(3):529–534
5. Tropsha A (2010) Best practices for QSAR model development, validation, and exploitation.
Mol Inform 29(6–7):476–488
6. Yang H, Sun L, Li W, Liu G, Tang Y (2018) In silico prediction of chemical toxicity for drug
design using machine learning methods and structural alerts. Front Chem 6:30. https://doi.org/
10.3389/fchem.2018.00030
7. Danishuddin Khan AU (2016) Descriptors and their selection methods in QSAR analysis:
paradigm for drug design. Drug Discov Today 21(8):1291–1302
8. Leach AR, Gillet VJ (2007) Molecular descriptors. An introduction to chemoinformatics.
Springer, Dordrecht, pp 53–74
9. Todeschini R, Consonni V (2000) Handbook of molecular descriptors. Wiley-VCH, Weinheim
10. Duan J, Dixon SL, Lowrie JF, Sherman W (2010) Analysis and comparison of 2D fingerprints:
insights into database screening performance using eight fingerprint methods. J Mol Graph
Model 29(2):157–170
11. National
Institutes
of
Health
(2009)
PubChem
substructure
fingerprint.
ftp://ftp.ncbi.nlm.nih.gov/pubchem/specifications/pubchem_fingerprints.txt. Accessed 10
Oct 2018
12. Rogers D, Hahn M (2010) Extended-connectivity fingerprints. J Chem Inf Model
50(5):742–754
13. Huang R, Xia M, Nguyen D-T et al (2016) Tox21Challenge to build predictive models of nuclear
receptor and stress response pathways as mediated by exposure to environmental chemicals
and drugs. Front Environ Sci 3:85. https://doi.org/10.3389/fenvs.2015.00085
14. Mayr A, Klambauer G, Unterthiner T, Hochreiter S (2016) DeepTox: toxicity prediction using
deep learning. Front Environ Sci 3:80. https://doi.org/10.3389/fenvs.2015.00080
15. Subramanian J, Simon R (2013) Overfitting in prediction models—Is it a problem only in high
dimensions? Contemp Clin Trials 36(2):636–641
