52 Is a Model’s Scatter Really “Very Small” …
331
To estimate confidence limits on differences in a performance measure (PM)
between two or more models, we start with a master table containing N rows, each of
which represents one time period and one concentration sampler. Bootstrap or jackknife resampling methods are used to determine the difference PM (i, j) between
two models (i and j). The values of PM (i, j) are ranked and used to define 95%
confidence limits. If the 95% confidence limits on PM (i, j) overlap 0.0, then it
can be concluded, with 95% confidence, that the difference in PM between the two
models is not significantly different from 0.0.
52.3 JU2003 Data Set and Models Evaluated
The JU2003 field data involved releases of puffs or plumes of SF 6 in the Oklahoma
City domain in 2003 [1]. Three to six puffs were released on each of the ten IOP
days [3, 6]. Observations of SF 6 concentration were made by 10 real-time samplers
(resolution of 0.5 s) at distances of a few hundred meters. In the UDINEE model comparison study [5], predictions of ten models from several countries were compared
with the puff observations. In this paper, the dosage/Q outputs from the UDINEE
data set are compared, where Q is the total mass of SF 6 in a puff. Puff-sampler
pairs are included in the analysis only if both predicted and observed max 0.5 s
concentrations exceed 400 ppm. The UDINEE model comparison reports and output
files use anonymous designations (e.g., Model M1). In the comparisons discussed in
this paper, three of the models are compared for two of the IOPs (3 and 8). BOOT
was applied to these models and data and the performance measures listed above
were calculated. In addition, confidence limits on the performance measures and the
differences in performance measures were calculated by BOOT.
52.4 Results
Of the many tables and figures that show the statistical results, we select one figure
and one table for discussion. There were four puffs released during each of IOP3
and IOP8 and, for each puff, six or seven samplers produced valid data. A total of
48 puff-sampler data pairs during IOP3 and IOP8 are analyzed. We consider only
Models M3, M5, and M11, since they produced model predictions for all 48 puffsampler pairs. Here we present results for dosages (concentrations integrated over
time, with units ppt-s) divided by mass emitted, Q.
As an aid to understanding the quantitative calculations by BOOT, it is useful
to look at scatter plots such as Fig. 52.1, where predicted and observed dosage/Q
are plotted for models M3, M5, and M11. The viewer’s first impression is that the
points are a “shotgun blast pattern” with little skill by the models. The scatter covers
a range of plus and minus about a factor of ten. Although models M3 and M5 appear
to overpredict more often than model M11, the points for the three models overlap.
331
To estimate confidence limits on differences in a performance measure (PM)
between two or more models, we start with a master table containing N rows, each of
which represents one time period and one concentration sampler. Bootstrap or jackknife resampling methods are used to determine the difference PM (i, j) between
two models (i and j). The values of PM (i, j) are ranked and used to define 95%
confidence limits. If the 95% confidence limits on PM (i, j) overlap 0.0, then it
can be concluded, with 95% confidence, that the difference in PM between the two
models is not significantly different from 0.0.
52.3 JU2003 Data Set and Models Evaluated
The JU2003 field data involved releases of puffs or plumes of SF 6 in the Oklahoma
City domain in 2003 [1]. Three to six puffs were released on each of the ten IOP
days [3, 6]. Observations of SF 6 concentration were made by 10 real-time samplers
(resolution of 0.5 s) at distances of a few hundred meters. In the UDINEE model comparison study [5], predictions of ten models from several countries were compared
with the puff observations. In this paper, the dosage/Q outputs from the UDINEE
data set are compared, where Q is the total mass of SF 6 in a puff. Puff-sampler
pairs are included in the analysis only if both predicted and observed max 0.5 s
concentrations exceed 400 ppm. The UDINEE model comparison reports and output
files use anonymous designations (e.g., Model M1). In the comparisons discussed in
this paper, three of the models are compared for two of the IOPs (3 and 8). BOOT
was applied to these models and data and the performance measures listed above
were calculated. In addition, confidence limits on the performance measures and the
differences in performance measures were calculated by BOOT.
52.4 Results
Of the many tables and figures that show the statistical results, we select one figure
and one table for discussion. There were four puffs released during each of IOP3
and IOP8 and, for each puff, six or seven samplers produced valid data. A total of
48 puff-sampler data pairs during IOP3 and IOP8 are analyzed. We consider only
Models M3, M5, and M11, since they produced model predictions for all 48 puffsampler pairs. Here we present results for dosages (concentrations integrated over
time, with units ppt-s) divided by mass emitted, Q.
As an aid to understanding the quantitative calculations by BOOT, it is useful
to look at scatter plots such as Fig. 52.1, where predicted and observed dosage/Q
are plotted for models M3, M5, and M11. The viewer’s first impression is that the
points are a “shotgun blast pattern” with little skill by the models. The scatter covers
a range of plus and minus about a factor of ten. Although models M3 and M5 appear
to overpredict more often than model M11, the points for the three models overlap.
