72
D. Krpelík and T. Basu
in an apriori-selected set of candidate distributions, say P := {P θ , θ ∈ Θ}. Based
on the axiomatic theory of probability, precise methods mainly comprise of two
competing methodologies—frequentist and Bayesian. Nevertheless, the common
inference scheme for constructing distributional point estimates, i.e. selecting a
single best-fitting probability distribution, is simply as follows:
1. Choose (subjectively) a set of plausible sampling distributions P.
2. Construct the likelihood function L(θ ; ;
x), which models the probability of
observing the collection
x for each parameter θ .
3. Select ˆ
θ that best fits the observations and approximate ˆ
P by P = P ˆ
θ (the
frequentist approach), or construct a mixture of distributions from the chosen
family with mixing weights w(θ) ∝ L(θ ; ;
x)π 0 (θ ) and approximate ˆ
P by
P = w(θ)P θ (the Bayesian approach; π 0 is called the prior distribution).
4. Evaluate the approximations of the desired probabilities from inferred distribution P .
The samples may come in various forms. Most commonly, they are considered
precisely specified (e.g. real values for a real RV), in which case the likelihood
function for inference from a set of independent samples will take the form
L(θ ; ;
x) =
n
i=1
f θ (x i ),
(2.26)
where f θ is the probability density function of distributions from the chosen family
P indexed by θ .
Example 2.6 Let us assume that we have a set of observations
x :=
{x 1 , . . . , x N } of a positive RV X. We choose the set of admissible sampling
distributions of the RV X to be the set of all exponential distributions
(F (x; θ) = 1 − exp(−θx)). The frequently used frequentist method is the
so-called maximum-likelihood estimate (MLE). Here, we seek such value of
θ which maximises the likelihood function (Eq. (2.26)). Thus,
θ MLE = argmax θ∈Θ L(θ ; ;
x),
and construct P = P θ MLE .
The resulting CDF is depicted in Fig. 2.7 with a label “MLE”.
Example 2.7 Let us assume the same scenario as in Example 2.6. We again
select the set of admissible sampling distributions of the RV X, P, to be
(continued)
D. Krpelík and T. Basu
in an apriori-selected set of candidate distributions, say P := {P θ , θ ∈ Θ}. Based
on the axiomatic theory of probability, precise methods mainly comprise of two
competing methodologies—frequentist and Bayesian. Nevertheless, the common
inference scheme for constructing distributional point estimates, i.e. selecting a
single best-fitting probability distribution, is simply as follows:
1. Choose (subjectively) a set of plausible sampling distributions P.
2. Construct the likelihood function L(θ ; ;
x), which models the probability of
observing the collection
x for each parameter θ .
3. Select ˆ
θ that best fits the observations and approximate ˆ
P by P = P ˆ
θ (the
frequentist approach), or construct a mixture of distributions from the chosen
family with mixing weights w(θ) ∝ L(θ ; ;
x)π 0 (θ ) and approximate ˆ
P by
P = w(θ)P θ (the Bayesian approach; π 0 is called the prior distribution).
4. Evaluate the approximations of the desired probabilities from inferred distribution P .
The samples may come in various forms. Most commonly, they are considered
precisely specified (e.g. real values for a real RV), in which case the likelihood
function for inference from a set of independent samples will take the form
L(θ ; ;
x) =
n
i=1
f θ (x i ),
(2.26)
where f θ is the probability density function of distributions from the chosen family
P indexed by θ .
Example 2.6 Let us assume that we have a set of observations
x :=
{x 1 , . . . , x N } of a positive RV X. We choose the set of admissible sampling
distributions of the RV X to be the set of all exponential distributions
(F (x; θ) = 1 − exp(−θx)). The frequently used frequentist method is the
so-called maximum-likelihood estimate (MLE). Here, we seek such value of
θ which maximises the likelihood function (Eq. (2.26)). Thus,
θ MLE = argmax θ∈Θ L(θ ; ;
x),
and construct P = P θ MLE .
The resulting CDF is depicted in Fig. 2.7 with a label “MLE”.
Example 2.7 Let us assume the same scenario as in Example 2.6. We again
select the set of admissible sampling distributions of the RV X, P, to be
(continued)
