Section 9.3: Permutation Procedures
167
9.3.2 Permutation (PP)
and Bootstrap (BP) Procedures
For both pp and BP the null hypothesis of D 2 = 0 allows the two samples
to be pooled, i.e. (Xl" .Xn, Y1" .Yn)' Two new n samples of X and Y
are created by random resampling from the pooled sample, either without
replacement in the case of pp or with replacement in the case of BP.
For pp this amounts to reordering the 20 vectors and relabeling the first
10 as Xl to Xn and the second 10 as Y 1 to Y n; in the Livezey and ehen (1983)
example in the last section this amounts to reordering the SOl series. In contrast, for BP the original samples may not be used proportionately, indeed
parts may not be used at all. However, this should not be construed as a deficiency, because BP is equivalent to random sampling from a monotonically
nondecreasing step function with steps at the observations. As the number of
observations increases, this empirical distribution function converges to the
true distribution.
Empirical distribution functions for each case considered (PP or BP, test
statistic, m) are built up with 200 different resamplings of the original samples
to be tested. The original sample is then tested against the 5% tail of the
empirical distribution function and the results noted.
9.3.3 Properties
The performance of PP and BP under various conditions (test statistic, m)
was studied by generating 1000 sets of paired samples (Xl" ,Xn'Y1" .Yn)
for every m and then conducting field tests using both procedures described
above with both test statistics (D 2 and k) on every one ofthese twin samples.
The true 5% value of both chi-squared and the counting norm are known
and the expected rejection rate for the exact parametric tests is 5%, so two
summary quantities that are of particular interest are the proportion of test
statistic estimates that exceed the true values (i.e. the bias in estimating
the true values) and the rate at which the null hypothesis is rejected in the
1000 tests. Perfect performance implies values of 0.50 and 0.05 for these two
quantities respectively. Results are displayed in Table 9.I.
Generally the procedures perform weIl for most m, but overall BP performs
a bit better in terms of rejection rates and PP in terms of bias sizes (the
bias for BP / D 2 tests is particularly large for large m). There really are no
significant differences between the performance of the D2 and counting norm
tests for PP. 2
2The counting norm has the advantage of highlighting the locations in the field that
lead to field significance when it is present. Of course, there is still no automatie guarantee
in this case that any particular location that is locally significant is not so by accident.
The confidence that can be placed on local significance increases with both its nominal
value and the level of the field's significance.
Précédent

- 176/336

Suivant