192
Appendix A
where
1
( )
( ),
N
N
k
k
t
tY
=
= ∑
Y
an m × 1 vector of complete-data suffi cient
statistics. The formula in Equation A.53 was provided by Dempster
et al. (1977) for the general incomplete-data problem. For our incomplete-data problem, explicit expressions for the covariances are given
by Equations A.49 and A.50. These are in forms that can readily be
computed.
The Fisher information matrix I (␪) is the expectation of
2 −
2
/ −u r −u s logL with respect to the joint distribution of X 1 , . . . , X n .
From Equation A.52, this matrix is equal to the covariance matrix of
the “estimated” complete-data score functions:
( )
1
E
l o g
, ,
1 ,2 , ,
N
k
n
k
r
f Y
r
m
u
=


∂
=
…


∂


∑
x
␪
␪
At the maximum-likelihood estimate u ˆ
n
, the Fisher information
matrix I(u ˆ
n
) may be estimated by I 0 (u ˆ
n
), with entries, evaluated at u ˆ
n
,
that are given by Equation A.52; hence, an estimate of the asymptotic
covariance matrix is given by the inverse of I 0 (u ˆ
n
). Tests and confi dence
procedures can be obtained by the usual normal approximation.
Inference for ␪ and N
In the previous section we demonstrated how maximum-likelihood
estimates for ␪ can be obtained when N is known. In this section, still
assuming the weight function w( y) is given, we are interested in estimating both ␪ and N. One approach is to do an (m11)-dimensional
grid search of a likelihood function L (␪, N | x n ) based on Equation A.8.
Another approach is to solve the likelihood equations in Equation A.18
via the EM algorithm to fi nd ␪ ˆ (N ) for different values of N, then determine N ˆ that maximizes the log-likelihood profi le, log L (␪ ˆ (N ), N | x n ).
The trouble with both of these approaches is that they are computationally expensive. In petroleum resource applications, our experience
with the log-likelihood profi le is that it is a rather “fl at” function of N
and frequently produces N ˆ 5 n, and on occasion it produces an unacceptably large estimate of N.
A third approach is to ignore the superpopulation part completely
and estimate N based on a method suggested by Gordon (1993) and
then estimate ␪ conditional on N. Gordon’s idea is to split a successive
sample from a fi nite population into two parts to approximate the
Précédent

- 215/257

Suivant