268
F. Firouzi et al.
L1 and L2 regularization parameters to cost function. Let us call the regularization
function g(β):
For lasso regression : g(W ) = λ (|w 1 | + |w 2 |)
For ridge regression : g(W ) = λ
w
2
1 + w
2
2
In the above equation, λ is the penalization parameter, and w1 and w2 are the
coefficients of x1 and x2, respectively. g(W) for lasso and ridge are depicted by
the blue diagram in Fig. 5.22. In the cost function of lasso and ridge regression,
we need to minimize f(w1, w2) + g(w1, w2). This is equivalent to find those points
that two contour plots (red and blue diagrams) meet each other. In other words, we
should calculate the minimum of f (W) + g(W), which is the intersection of two
functions (f(W) and g(W)). As shown in Fig. 5.22, in lasso regression, two contour
diagrams can meet at a point where either w1 or w2 is zero. Therefore, the solution
of lasso can be sparse. On the other hand, the contour plots in ridge regression do
not have any tips, and thus it cannot result in any sparse solution.
5.2.2.2 Elastic Net Regularization
Zou and Hastie introduced the concept of the elastic net to overcome the weaknesses
of L1 and L2 regularizations in 2005. When the number of independent variables is
more than the sample size (p > n), only one independent variable can be selected
from any set of highly correlated independent variables using the lasso regression
algorithm (up to n independent variables). In addition, if the number of independent
variables is less than the sample size, the ridge regression method would have a
better performance.
Most of the times, highly correlated independent variables have similar regression coefficients. This situation is called the grouping effect. In real-world applications, the grouping effect can be beneficial for building the model. For example,
in gene identification of diseases, the researchers are intended to find associated
independent variables rather than only one from each set (which happens in
lasso). Additionally, selecting a single independent variable from a set of highly
correlated independent variables could result in a less robust model, which increases
the precision error. This fact demonstrates why ridge regression performs more
efficiently than lasso in this situation.
The elastic net algorithm is a combination of both L1 and L2 norms, in which
some coefficients are shrunk (similar to ridge regression) and some are set to zero
(like the lasso regression method). This method has two shrinkage parameters:
w
∗
= argminy − xw
2
2 + λ 2 w
2
2 + λ 1 w 1
F. Firouzi et al.
L1 and L2 regularization parameters to cost function. Let us call the regularization
function g(β):
For lasso regression : g(W ) = λ (|w 1 | + |w 2 |)
For ridge regression : g(W ) = λ
w
2
1 + w
2
2
In the above equation, λ is the penalization parameter, and w1 and w2 are the
coefficients of x1 and x2, respectively. g(W) for lasso and ridge are depicted by
the blue diagram in Fig. 5.22. In the cost function of lasso and ridge regression,
we need to minimize f(w1, w2) + g(w1, w2). This is equivalent to find those points
that two contour plots (red and blue diagrams) meet each other. In other words, we
should calculate the minimum of f (W) + g(W), which is the intersection of two
functions (f(W) and g(W)). As shown in Fig. 5.22, in lasso regression, two contour
diagrams can meet at a point where either w1 or w2 is zero. Therefore, the solution
of lasso can be sparse. On the other hand, the contour plots in ridge regression do
not have any tips, and thus it cannot result in any sparse solution.
5.2.2.2 Elastic Net Regularization
Zou and Hastie introduced the concept of the elastic net to overcome the weaknesses
of L1 and L2 regularizations in 2005. When the number of independent variables is
more than the sample size (p > n), only one independent variable can be selected
from any set of highly correlated independent variables using the lasso regression
algorithm (up to n independent variables). In addition, if the number of independent
variables is less than the sample size, the ridge regression method would have a
better performance.
Most of the times, highly correlated independent variables have similar regression coefficients. This situation is called the grouping effect. In real-world applications, the grouping effect can be beneficial for building the model. For example,
in gene identification of diseases, the researchers are intended to find associated
independent variables rather than only one from each set (which happens in
lasso). Additionally, selecting a single independent variable from a set of highly
correlated independent variables could result in a less robust model, which increases
the precision error. This fact demonstrates why ridge regression performs more
efficiently than lasso in this situation.
The elastic net algorithm is a combination of both L1 and L2 norms, in which
some coefficients are shrunk (similar to ridge regression) and some are set to zero
(like the lasso regression method). This method has two shrinkage parameters:
w
∗
= argminy − xw
2
2 + λ 2 w
2
2 + λ 1 w 1
