82
T. Basu et al.
emerging in the eighteenth century, following advances in the theory of probability
[14, p. 176].
In this chapter, we will discuss statistical regularization and uncertainty quantification problems using Least Absolute Shrinkage and Selection Operator (LASSO)
estimators [24, 25]. The LASSO estimator is a popular regularization method due
to its variable selection property. After Tibshirani introduced LASSO in 1996 [24],
numerous authors contributed further to the theory, including Osborne, Presnell,
and Turlach [21] and Efron et al. [6]. Friedman et al. [10] discussed computational
aspects of the LASSO. Park and Casella [22] introduced the Bayesian approach for
LASSO estimators, using a hierarchical mixture model for parameter estimation.
Other notable works deal with the specification of shrinkage parameter by Lykou
and Ntzoufras [18], the Dirichlet LASSO by Das and Sobel [4], and the spike and
slab LASSO by Ro˘ cková [23].
First, we will introduce the basic notions behind statistical modeling and
regularization. In Sect. 3.2, we will look at some important concepts of parameter
estimation with and without regularization. Eventually, we will introduce the
LASSO estimators in Sect. 3.3. In Sect. 3.4, we will discuss different uncertainty
quantification methods for the LASSO followed by an extension to the logistic
model in Sect. 3.5. Section 3.6 concludes the chapter.
3.1.1 Statistical Modeling
To make statistical inferences from data, first, we need variables and a model
describing the relations between those variables. We can categorize variables into
response variables and predictor variables:
1. Predictor (or independent) variables are characteristics of the system which
directly control the properties of the system.
2. Response (or dependent) variables are characteristics of the system which depend
on the predictor variables. In other words, they respond to a change of values of
the predictors in some systematic fashion.
Assume we have a dataset containing n independent and identically distributed
(i.i.d.) observations of real-valued responses y 1 , . . . , y n ∈ R, along with corresponding vector-valued predictors x 1 , . . . , x n ∈ R p . We consider each x i to be a
column vector.
Example 3.1 (Gaia Dataset) Gaia is a mission by the European Space Agency
(ESA) to formulate a three-dimensional map of our galaxy [8]. The data depicted
in Fig. 3.1 are part of a dataset which was simulated prior to the launch of the
mission from computer experiments [1, 7]. The data contain essentially spectral
information divided into p = 16 wavelength bands (intervals), along with certain
stellar parameters which are to be inferred from the spectral data. That is, each
observation in the data set represents a stellar object, and the measurement for
Précédent

- 87/568

Suivant