266
F. Firouzi et al.
which it also should be minimized. Among regularization methods, L2 and L1
are the most frequently used methods, which modify the cost function by adding
a generalization term as follows:
Cost function = Loss + Regularization term
Adding the regularization term results in smaller weights/coefficients, which
leads to smaller overfitting. In other words, the regularization term punishes the
cost function. The utilized regularization term is different in L1 and L2 methods. In
L2, the cost function would be as follows:
Cost function = Loss + λ
w
2
where parameter λ is the regularization parameter, a hyperparameter that must be
optimized for better performance. L2 regularization is also called weight decay
because it forces the weights toward zero, however not exactly zero. A regression
analysis method that performs L2 regularization is called ridge regularization. Ridge
regularization is one of the well-known techniques to overcome overfitting.
Similarly, in L1, the cost function is written in this way:
Cost function = Loss + λ
w
in which the cost function penalizes the absolute value of the weights. However,
unlike L2, the weights can be forced to be absolute zero. Since those input variables
(features) with zero coefficients can be dropped from the regression model, the
L1 regularization is useful for feature selection and reducing the complexity of
the model. L1 regularization methods are also called lasso regularization. In other
words, L1 regularization provides sparse solutions of the model by removing unimportant input variables. Getting sparse solutions could decrease the computational
complexity of the model due to the presence of features with a coefficient of zero.
In case of highly correlated features, the ridge generalization distributes the
coefficients among all of the features depending on the correlation, but lasso
regularization chooses the features selectively and makes the coefficient of other
features zero.
5.2.2.1 Geometric Interpretations of Regularization
One can define the L1 norm as the cumulative summation of absolute values of
a vector’s components. As an example, the L1 norm of the vector [x1, x2] is
|x1| + |x2|. With this definition in mind, we plot all the points whose L1 norms
are equal to a constant value (c), as presented by a blue line in Fig. 5.21.
The geometry of L1 norm in Fig. 5.21 looks like a rotated square (an octahedron
in higher dimensional space), in which the points on the tips are sparse (which
F. Firouzi et al.
which it also should be minimized. Among regularization methods, L2 and L1
are the most frequently used methods, which modify the cost function by adding
a generalization term as follows:
Cost function = Loss + Regularization term
Adding the regularization term results in smaller weights/coefficients, which
leads to smaller overfitting. In other words, the regularization term punishes the
cost function. The utilized regularization term is different in L1 and L2 methods. In
L2, the cost function would be as follows:
Cost function = Loss + λ
w
2
where parameter λ is the regularization parameter, a hyperparameter that must be
optimized for better performance. L2 regularization is also called weight decay
because it forces the weights toward zero, however not exactly zero. A regression
analysis method that performs L2 regularization is called ridge regularization. Ridge
regularization is one of the well-known techniques to overcome overfitting.
Similarly, in L1, the cost function is written in this way:
Cost function = Loss + λ
w
in which the cost function penalizes the absolute value of the weights. However,
unlike L2, the weights can be forced to be absolute zero. Since those input variables
(features) with zero coefficients can be dropped from the regression model, the
L1 regularization is useful for feature selection and reducing the complexity of
the model. L1 regularization methods are also called lasso regularization. In other
words, L1 regularization provides sparse solutions of the model by removing unimportant input variables. Getting sparse solutions could decrease the computational
complexity of the model due to the presence of features with a coefficient of zero.
In case of highly correlated features, the ridge generalization distributes the
coefficients among all of the features depending on the correlation, but lasso
regularization chooses the features selectively and makes the coefficient of other
features zero.
5.2.2.1 Geometric Interpretations of Regularization
One can define the L1 norm as the cumulative summation of absolute values of
a vector’s components. As an example, the L1 norm of the vector [x1, x2] is
|x1| + |x2|. With this definition in mind, we plot all the points whose L1 norms
are equal to a constant value (c), as presented by a blue line in Fig. 5.21.
The geometry of L1 norm in Fig. 5.21 looks like a rotated square (an octahedron
in higher dimensional space), in which the points on the tips are sparse (which
