238
R. Ménard and M. Deshaies-Jacques
in scales being sampled in a single observation vs that of the model grid box, subgrid
scale variability that may be captured by the observation but not by the model,
missing modeling processes, etc., that collectively we call representativeness error.
The observation error is thus not well known and need to be estimated. Although from
the analysis algorithm we can compute an analysis error variance, this value depends
only on the input observation and model error statistics used in the algorithm and may
not represent the true analysis error if they are incorrectly specified. Furthermore,
misspecified error statistics generally leads to a degradation of the analysis accuracy.
One way to evaluate the true analysis error is with a forecast, but the forecast accuracy
depends on many other things than the analysis accuracy. It depends on the accuracy
of emissions, the coherence between observed and unobserved species, the boundary
layer mixing and ventilation, etc. The impact of observations may also be limited
as the AirNow observations are confined to the surface. Furthermore, to estimate
and tune the input errors statistics based on forecasts is computationally expensive.
Thus, our objective is to develop a method to evaluate the true analysis error but
without relying on a model forecast. Doing so, we also contribute in maximizing the
information content provided by the observations.
Our method to evaluate the analysis error variance is based on cross-validation.
To perform cross-validation we separate the observation data set into N spatially
random-distributed observation sites (here we use N = 3), performing an analysis
using N–1 subsets and verifying the analysis with the remaining single subset. This
leads to an objective measure of analysis accuracy [8, 9] where the process is done N
times by permutation of all subsets used for evaluation. An example of three subsets of
PM 2.5 for cross-validation is illustrated in Fig. 37.1 of Ménard and Deshaies-Jacques
[8].
One of the main sources of the true observation and model errors comes in
observation-minus-model residuals, often noted as O-B (B for background). As the
observations are compared with the (interpolated) model values, they thus contains
errors of representativeness. The variance of the O-B can also be considered as the
sum of observation and model error variances (this is evident when observation and
model errors are uncorrelated). We could obtain an estimate of observation error
variance and an estimate of model error variance, if we would know their ratio and
that is exactly what optimization using cross-validation is capable of obtaining, i.e.
the ratio of error variances that minimizes the analysis error variance (see Ménard
and Deshaies-Jacques [9] for a full discussion). An example of this for PM 2.5 is
provided in Fig. 37.1.
37.2 Main Results of Cross-Validation
In Fig. 37.1, the variance of the independent observation-minus-analysis using the
(N-1) subsets, i.e. var(O-A), for each subsets is displayed with circles. The mean
variance over the 3 subsets is displayed with a solid black line. Verifying against
the same set of observation as those used for the analysis is also displayed along the
R. Ménard and M. Deshaies-Jacques
in scales being sampled in a single observation vs that of the model grid box, subgrid
scale variability that may be captured by the observation but not by the model,
missing modeling processes, etc., that collectively we call representativeness error.
The observation error is thus not well known and need to be estimated. Although from
the analysis algorithm we can compute an analysis error variance, this value depends
only on the input observation and model error statistics used in the algorithm and may
not represent the true analysis error if they are incorrectly specified. Furthermore,
misspecified error statistics generally leads to a degradation of the analysis accuracy.
One way to evaluate the true analysis error is with a forecast, but the forecast accuracy
depends on many other things than the analysis accuracy. It depends on the accuracy
of emissions, the coherence between observed and unobserved species, the boundary
layer mixing and ventilation, etc. The impact of observations may also be limited
as the AirNow observations are confined to the surface. Furthermore, to estimate
and tune the input errors statistics based on forecasts is computationally expensive.
Thus, our objective is to develop a method to evaluate the true analysis error but
without relying on a model forecast. Doing so, we also contribute in maximizing the
information content provided by the observations.
Our method to evaluate the analysis error variance is based on cross-validation.
To perform cross-validation we separate the observation data set into N spatially
random-distributed observation sites (here we use N = 3), performing an analysis
using N–1 subsets and verifying the analysis with the remaining single subset. This
leads to an objective measure of analysis accuracy [8, 9] where the process is done N
times by permutation of all subsets used for evaluation. An example of three subsets of
PM 2.5 for cross-validation is illustrated in Fig. 37.1 of Ménard and Deshaies-Jacques
[8].
One of the main sources of the true observation and model errors comes in
observation-minus-model residuals, often noted as O-B (B for background). As the
observations are compared with the (interpolated) model values, they thus contains
errors of representativeness. The variance of the O-B can also be considered as the
sum of observation and model error variances (this is evident when observation and
model errors are uncorrelated). We could obtain an estimate of observation error
variance and an estimate of model error variance, if we would know their ratio and
that is exactly what optimization using cross-validation is capable of obtaining, i.e.
the ratio of error variances that minimizes the analysis error variance (see Ménard
and Deshaies-Jacques [9] for a full discussion). An example of this for PM 2.5 is
provided in Fig. 37.1.
37.2 Main Results of Cross-Validation
In Fig. 37.1, the variance of the independent observation-minus-analysis using the
(N-1) subsets, i.e. var(O-A), for each subsets is displayed with circles. The mean
variance over the 3 subsets is displayed with a solid black line. Verifying against
the same set of observation as those used for the analysis is also displayed along the
