88
T. Basu et al.
Therefore, if X T X is invertible (this requires that the number of observations, n, is
larger or equal than the total number of predictors, p), then the OLS estimator is
given by
ˆ
β
OLS = (X
T X)
−1 X
T Y ,
(3.19)
where (X T X) −1 X T is the Moore–Penrose inverse of X.
The Gauss–Markov theorem states that when the errors are uncorrelated with
expectation zero and constant variance, then the OLS estimate is the best linear
unbiased estimator.
Two issues that often arise are:
1. If p > n then X T X is singular; hence Eq. (3.18) has no unique solution.
2. Even if p ≤ n, p may still be much larger than needed, and we may wish
to identify sparse solutions where unnecessary parameters are set to zero. In
other words, we may wish to perform variable selection as part of our statistical
inference.
3.2.2 Non-Negative Garrote
The non-negative garrote was introduced by Breiman [3]. It is a two-stage procedure
that gives a sparse solution. It has a close relationship to the LASSO; however
as a starting point of the problem, the OLS estimates are needed. Given the
initial estimate ˆ
β
OLS ∈ R p , we solve the following optimization problem over
c = (c 1 , c 2 , . . . , c p ) T :
ˆ
c = arg min
c≥0
c 1 ≤t
Y − XC ˆ
β
OLS
2
2
(3.20)
where C := diag(c) ∈ R p×p , and where . 1 denotes the l 1 -norm; that is c 1 =
p
i=1 |c i |. We get the final non-negative garrote parameter estimate ˆ
β by setting
ˆ
β i = ˆ
c i ˆ
β OLS
i
for each i ∈ {1, 2, . . . , p}.
Equivalently, we can solve the dual problem, by introducing a Lagrangian
multiplier λ for the constraint c 1 − t ≤ 0 [16], similar to what we discussed
in Sect. 3.1.4:
max
λ≥0
min
c≥0
Y − XC ˆ
β
OLS
2
2 + λ(c 1 − t)
(3.21)
Effectively, we thus need to solve
Précédent

- 93/568

Suivant