102
Z. Liu et al.
Feature Selection
In the converter data from steel enterprises, the high dimension and high correlation
of feature variables make the useful information carried highly redundant and carry
more data noise [14, 15]. For the 12 feature variables analyzed above, it is necessary
to select the appropriate combination of feature variables in each feature subspace
containing data subset, to carry more useful information and less data noise as far as
possible. And it is conducive to the performance prediction of the model. Here, the
appropriate combination of feature variables is selected based on RFE.
The feature variable selection algorithm based on RFE is shown in Table 4. The
input is a training set with 12-dimensional feature variables and the outputs are the
optimal feature variable combination and the corresponding prediction model. Based
on the current feature variables, an independent prediction model is trained, and its
prediction performance will also be evaluated. And the importances of the current
feature variables are calculated and sorted, and the feature variable with the lowest
importance will be eliminated. For the remaining feature variables, repeat the above
process until the number of feature variables is zero. In the above process, the number
of feature variables gradually decreases from 12 to 1, so 12 different combinations
of feature variables and 12 corresponding prediction models will be produced. The
model with the best prediction performance on the training set is selected as the final
prediction model of the subspace, and the corresponding feature variable combination
is the selected feature variable. The Mean Absolute Error (MAE) is selected as the
evaluation indicator in this article.
In the RFE algorithm, to select the optimal combination of feature variables, the
importance calculation of feature variables not only needs to consider the correlations
between feature variables and oxygen consumption, but also the correlation between
feature variables. Here, Baptiste Gregorutti’s method [16] is used to calculate the
importance of feature variables. The feature variables and oxygen consumption can
be regarded as following a Gaussian distribution, namely
(X, y) ∼ N n+1
0,
C τ
τ
T
σ
2
y
(5)
Table 4 Feature variable selection algorithm based on RFE
Input:
Output:
Training set {X, Y}with n feature variables, n = 12
Optimal feature variable combinations and corresponding prediction model
Step 1:
Step 2:
Step 3:
Step 4:
Step 5:
Train an independent prediction model with n feature variables;
Calculate the importance of each feature variable;
Sort the importances of feature variables and eliminate the least important feature, then
n = n – 1;
Repeat Step 1–Step 3 for the remaining features until the feature amount is 0;
Select the best in n combinations with different number of feature variables and n
corresponding prediction models
Z. Liu et al.
Feature Selection
In the converter data from steel enterprises, the high dimension and high correlation
of feature variables make the useful information carried highly redundant and carry
more data noise [14, 15]. For the 12 feature variables analyzed above, it is necessary
to select the appropriate combination of feature variables in each feature subspace
containing data subset, to carry more useful information and less data noise as far as
possible. And it is conducive to the performance prediction of the model. Here, the
appropriate combination of feature variables is selected based on RFE.
The feature variable selection algorithm based on RFE is shown in Table 4. The
input is a training set with 12-dimensional feature variables and the outputs are the
optimal feature variable combination and the corresponding prediction model. Based
on the current feature variables, an independent prediction model is trained, and its
prediction performance will also be evaluated. And the importances of the current
feature variables are calculated and sorted, and the feature variable with the lowest
importance will be eliminated. For the remaining feature variables, repeat the above
process until the number of feature variables is zero. In the above process, the number
of feature variables gradually decreases from 12 to 1, so 12 different combinations
of feature variables and 12 corresponding prediction models will be produced. The
model with the best prediction performance on the training set is selected as the final
prediction model of the subspace, and the corresponding feature variable combination
is the selected feature variable. The Mean Absolute Error (MAE) is selected as the
evaluation indicator in this article.
In the RFE algorithm, to select the optimal combination of feature variables, the
importance calculation of feature variables not only needs to consider the correlations
between feature variables and oxygen consumption, but also the correlation between
feature variables. Here, Baptiste Gregorutti’s method [16] is used to calculate the
importance of feature variables. The feature variables and oxygen consumption can
be regarded as following a Gaussian distribution, namely
(X, y) ∼ N n+1
0,
C τ
τ
T
σ
2
y
(5)
Table 4 Feature variable selection algorithm based on RFE
Input:
Output:
Training set {X, Y}with n feature variables, n = 12
Optimal feature variable combinations and corresponding prediction model
Step 1:
Step 2:
Step 3:
Step 4:
Step 5:
Train an independent prediction model with n feature variables;
Calculate the importance of each feature variable;
Sort the importances of feature variables and eliminate the least important feature, then
n = n – 1;
Repeat Step 1–Step 3 for the remaining features until the feature amount is 0;
Select the best in n combinations with different number of feature variables and n
corresponding prediction models
