Machine Learning Predictions of Adsorption Energies …
143
trees used in RFR, GBR, and ETR was fixed at 200; other hyperparameters were
set to the default values given by scikit-learn. ETR and RFR are less sensitive to
hyperparameters than is GBR, and thus do not require demanding hyperparameter
tuning. Moreover, the training time for ETR is shorter than those for RFR and GBR
because ETR is based on random-splitting decision trees. Adsorption energies can
thus be readily predicted using the ETR method without DFT calculations or hyperparameter tuning. We therefore conclude that the ETR model is the best choice for
the prediction of the adsorption energies considered in this study. The computation
times for the DFT calculations and ML analysis were compared. A DFT calculation
of the adsorption energy of CH 3 on a Cu monometallic surface took about 10 h on
our 32-core workstation, whereas each corresponding ML prediction took less than
1 s on our 1-core laptop computer. The computation time for ML prediction depends
on the applied ML model, not the system of interest. For many ML methods that
include tree ensembles, the predicted values can be calculated instantly because the
function form for prediction is explicitly obtained by fitting an ML model function
to the given training data. For prediction, the corresponding descriptor values can be
simply substituted into the existing function. This demonstrates the potential of ML
as a fast alternative to DFT calculations.
For a better understanding of the ML analysis, we now briefly discuss the prediction results. As shown in our pre-evaluations (Fig. 4), ETR and RFR tend to slightly
outperform GBR. Our problem consists of only 46 system examples, some of which
are unique, and thus the difficulty of prediction greatly depends on the training and
test split. If the training and test sets contain similar systems, the adsorption energies
can easily be predicted. In contrast, if the sets contain dissimilar systems, they are
difficult to predict. In this situation, it is more important to reduce prediction variance
than prediction bias. Both ETR and RFR fit this purpose, with ETR outperforming
RFR because it is better at reducing prediction variance. Our current ML model gives
some errors even with ETR. This is most likely due to the small dataset size used
in this study. For more accurate prediction, it is generally desirable to collect larger
datasets.
Next, we show the relevance (or redundancy) of each of the 12 descriptors used
in the ETR model. ETR models take a form of an additive ensemble of simple
regression-tree models that recursively partition the data points into two using a
single selected descriptor for each split. It provides a feature importance score for
each descriptor computed as the normalized total reduction of the impurity for the
generated partitions of the data by that descriptor. This impurity-based score can
be used as a standard tool to assess the relative importance of a descriptor with
respect to the predictability of adsorption energies. Since this score obtained from
the training set can be sometimes misleading, another more careful choice would be
the permutation importance score, a decrease in the prediction accuracy for a given
trained model when a single feature value in the test set is randomly shuffled. Note
that although the statistical importance of descriptors varies with the ML method
used, these results are only for ETR. Figure 6 shows the feature importance scores
and permutation importance scores (averaged with fourfold cross validation) of all
12 descriptors for predicting the adsorption energies of CH 3 . The most important
Précédent

- 146/167

Suivant