The gas emission from the mining face is complicated and affected by many factors
including gas content, gas permeability, pressure, and buried depth. The highdimensional characteristics of these influencing factors are likely to lead to dimensional
disaster and over-exposure, which will influence the accuracy of prediction. Thus,
dimension, accuracy, and representativeness of the factors affecting the amount of gas
emission are crucial to the prediction. In this article, a feature selection method based
on Lasso algorithm is proposed. On the basis of the original feature space, an optimal
feature subset is selected by eliminating irrelevant and redundant features with better
readability; meantime, the feature meaning of the original dataset has not changed. The
main features of the influence factors of gas emission screened from data perspective
are used to establish the prediction model, which can solve the multi-collinearity
problem between variables so that the changing law of gas emission from the mining
face can be accurately tracked.
2 Principle of Lasso Algorithm
The least absolute shrinkage and selection operator (Lasso) is a regularized sparse
model with penalty, first proposed by statistician Tibshirani in 1996. In order to provide
effective algorithm support for Lasso, Efron et al. [4] proposed the least angle
regression (LARS) algorithm. Friedman proposed algorithm which can solve calculation problem of Lasso as well, but glmnet ignores the difference between the two
adjacent regression coefficients, making the volatility of the estimated value large. Zou
and Trevor [5] proposed the elastic net method, which adds a two-norm constraint on
the basis of LARS to solve the over-fitting problem of high-dimensional small sample.
The basic idea of Lasso regression is to perform L1 norm constraint on regression
coefficients, so that the regression coefficients of some independent variables can be
automatically compressed to zero since the residual sum of square is minimized. In
other words, based on the least square estimation of the traditional linear regression
method, it adds the penalty term of the absolute value situation to select variables and
obtain an interpretable model.
For multiple linear regression models:
y ¼ a þ b 1 x 1 þ b 2 x 2 þ Á Á Á þ b p x p þ e;
ð1Þ
The Lasso estimates for the constant term and the regression coefficient are as
follows:
ð^ a; ^
bÞ ¼ arg min
X n
i¼1
y i À a i À
X p
j¼1
b j x ij
! 2
s:t:
X p
j¼1
b j
t; ðt ! 0Þ
ð2Þ
Lasso regression uses the correlation among data to compress the regression
coefficient b i which has little influence on y by controlling the parameter t. The smaller
the t is, the higher the degree of compression, so that more regression coefficients will
be compressed to 0 and the corresponding variables will be removed from the model to
166
Q. Chen and L. Huang
Précédent

- 178/679

Suivant