270
F. Firouzi et al.
Data
Prior
Information
Bayes’ Theorem
Posterior
Distribution
Fig. 5.24 Bayesian linear regression
In the Bayesian viewpoint, not only the output (y) has a distribution, but also all
the model parameters (weights) have a distribution. In Bayesian regression, the
following terms are defined:
Priors Priors are the initial or guess value of the model parameters that a domain
expert can put into the model prior to training. If there is no knowledge about the
parameters, non-informative priors, such as a normal distribution, could be used
instead.
Maximum likelihood Maximum likelihood estimation is a method to determine
the model parameter values in a way that the produced data (output of the model)
is equal to the actual observed data. For example, for a Gaussian distribution curve,
which has two parameters to be optimized (the mean, μ, and the standard deviation,
σ ), the maximum likelihood method can find the model parameters in a manner that
the generated curve best fits the data.
Posterior Posterior indicates the output distribution of Bayesian linear regression
based on the model parameters and priors. For a given dataset, one can estimate the
posterior probability distribution by the Bayesian rule:
posterior =
likelihood × prior
Normalization
Bayesian regression algorithm enables us to compute the distribution of possible
model parameters (posterior) based on the training dataset and the prior (see
Fig. 5.24). Note that when we have infinite data, the posterior converges to the output
of OLS linear regression. On the other hand, when we do not have enough data to
train the model, the distribution of posterior spreads out.
5.3 Feature Selection
Feature selection (also called variable or attribute selection) is the process of
selecting a subset of the input features to construct a high-performance machine
learning model. In other words, feature selection enables us to implement a potential
Précédent

- 276/647

Suivant