200
Let F denote an underlying superpopulation distribution that is
assumed to be discrete (i.e., that gives mass p k to z k , k 5 1, ... , K ). In this
appendix, we explain the procedure for estimating F without making
any assumptions about its shape. Suppose there are N pools in a play
with magnitudes (such as pool sizes) Y 1 , ... , Y N . This model assumes that
the N values are generated independently of an identical distribution, F.
Let ( y 1 , ... , y n ) denote the magnitudes of the n discovered pools, in
order of discovery. Let N k be the unknown number of Y i ’s that have
masses of z k , and let n k be the observed number of y i ’s in the sample that
have masses of z k , k 5 1, ... , K. It is assumed that sampling is executed
proportional to the size measure w( y).
Let b i 5 w( y i ) 1 · · · 1 w( y n ). It can then be shown that the probability
of observing the ordered sample ( y 1 , ... , y n ) under the successive sampling discovery model is proportional to
( )
0
1
1
( )
k
k
N n
K
K
n
t w z
k
k
n
k
k
L
p
p e
g t dt
−
∞
−
=
=
∝
∑
∏ ∫
(B.1)
where g n (t ) is the density of T 5 « 1 / b 1 1 · · · 1 « n / b n and the « i ’s are independent and identical standard exponential random variables (Wang
Appendix B: Nonparametric Procedure for
Estimating Distributions
Let F denote an underlying superpopulation distribution that is
assumed to be discrete (i.e., that gives mass p k to z k , k 5 1, ... , K ). In this
appendix, we explain the procedure for estimating F without making
any assumptions about its shape. Suppose there are N pools in a play
with magnitudes (such as pool sizes) Y 1 , ... , Y N . This model assumes that
the N values are generated independently of an identical distribution, F.
Let ( y 1 , ... , y n ) denote the magnitudes of the n discovered pools, in
order of discovery. Let N k be the unknown number of Y i ’s that have
masses of z k , and let n k be the observed number of y i ’s in the sample that
have masses of z k , k 5 1, ... , K. It is assumed that sampling is executed
proportional to the size measure w( y).
Let b i 5 w( y i ) 1 · · · 1 w( y n ). It can then be shown that the probability
of observing the ordered sample ( y 1 , ... , y n ) under the successive sampling discovery model is proportional to
( )
0
1
1
( )
k
k
N n
K
K
n
t w z
k
k
n
k
k
L
p
p e
g t dt
−
∞
−
=
=
∝
∑
∏ ∫
(B.1)
where g n (t ) is the density of T 5 « 1 / b 1 1 · · · 1 « n / b n and the « i ’s are independent and identical standard exponential random variables (Wang
Appendix B: Nonparametric Procedure for
Estimating Distributions
