84
T. Basu et al.
where φ is a function that depends on a parameter vector β. For instance, as will be
described in Sect. 3.1.3 in more detail, in a linear regression context, one typically
has φ(x i , β) = x T
i β. There also exist non-parametric approaches which do not
assume an explicit parametric shape, but most of such approaches achieve this by
simply introducing a large number of parameters, so that they still can be expressed
as in Eq. (3.1).
3.1.2 Statistical Inference
Statistical inference is the process by which we use the available data to gain
knowledge about the model parameters, such as β in Eq. (3.1), as well as their
uncertainties. In a wider sense, it will also include methods by which we quantify
and validate our assumptions on the model. Statistical inference deals with the estimation of parameters that are used to specify the family of probability distributions
which underlie the statistical model for y i |x i . Inference has several applications in
science and engineering. Generally, there are two conceptually different approaches
to statistical inference: the frequentist approach and the Bayesian approach. There
are some other concepts available which are beyond the scope of this chapter but
are addressed in other articles in this volume.
The frequentist approach is the most widely used estimation method. Sometimes
it is referred to as the “classical” approach. The estimation can be a point estimate
where we simply try to find the best guess for the parameter of the parametric model.
Alternatively, we seek an interval which covers the unknown parameter value with
high probability (generally 0.95). We call this a 95% confidence interval.
While several point estimators are available, the maximum likelihood estimator
(or MLE) is among the most popular because of its simple and wide implementability and its consistency properties. It finds the parameter value which maximizes
the probability density of the sample given the parameter, i.e., the likelihood. For
linear regression models under normal errors, MLE is equivalent to the ordinary
least squares (OLS).
The Bayesian approach starts from Bayes’ rule for conditional probability.
Denote the data by Y . For example, in our setting, Y is simply the vector of
observed response values (y 1 , . . . , y n ) T . The statistical model is specified through a
likelihood function p(Y | β). In the context of the regression model in Eq. (3.1), this
likelihood would be considered conditional on the observed values of the predictors,
i.e., the observed values of the predictors are considered as fixed. Finally, we need a
prior distribution p(β) for the model parameters β to incorporate our prior knowledge. Bayes’ rule then tells us that the posterior distribution p(β | Y ) is given by
p(β | Y ) ∝ p(β) × p(Y | β).
(3.2)
The normalization constant can be calculated from the law of total probability if
necessary. However, this calculation may not be always trivial so that simulation
Précédent

- 89/568

Suivant