3 Uncertainty Quantification in Lasso-Type Regularization Problems
93
3.3.2 Cross-Validation
Cross-validation is a commonly used method to identify the optimal value of a
tuning parameter, which is in our case the penalty parameter λ. It is based on
minimizing an estimate of the prediction error. In cross-validation, we use one part
of the data to fit the LASSO model and the other part of the data to validate it [15].
We fix initially a dense grid of values of λ, i.e., λ is discretized with small
step-sizes over a suitable range which reflects the scope of the regularization tradeoff that we are willing to consider. The dataset is then divided into K equally
sized partitions. We assume for simplicity that K is a divisor of n so that each
partition contains n/K elements. For each fixed value of λ of the grid, and the k’th
partition, k = 1, . . . , K, we fit the LASSO model using the remaining K − 1 parts
and calculate the prediction error of the fitted model. Specifically, denote ˆ
β
−k
λ the
parameter vector obtained under a penalty of λ when omitting the k’th partition, so
that x T
i
ˆ
β
−k
λ is the corresponding fitted model under predictor x i . Then the prediction
error for the k’th partition is
P k (λ) =
K
n
n /K
i=1
L(y i , x
T
i
ˆ
β
−k
λ )
(3.35)
where, for the linear model (Eq. (3.3)), the loss function L is just the squared error.
We repeat this step for every k = 1, 2, . . . , K and combine the values of P k (λ) to
find the average prediction error, P (λ) = K −1 K
k=1 P k (λ). This is then repeated
for every value of λ in the grid, and we choose the value of λ which minimizes P (λ)
[16].
For smaller values of λ, the LASSO estimators contain more predictors which
may lead to an over-fitted model. However, for larger values of λ, the model
has fewer predictors leading to sparsity and producing a more easily interpretable
model.
To avoid misunderstandings, it is noted that the problem of finding the optimal λ
(in the sense of minimal prediction error), as discussed in this subsection, is very
different from, and entirely unrelated to, the problem of maximizing over λ as,
for instance, in Eq. (3.27). The latter is a purely formal operation which ensures
mathematical equivalence of the two dual versions of the LASSO optimization
problem and does not imply any statement on the best choice of λ.
3.3.2.1 Example: Gaia Dataset
Figure 3.5 represents the cross-validation curve for the Gaia dataset. Here we have
taken normalized data to get rid of scalability. The graph is consistent with the
property of cross-validation, i.e., we can see that for smaller values of λ the number
of predictors is higher and for larger values of λ the number of predictors gets
Précédent

- 98/568

Suivant