102
General Data Collection and Sampling Design Considerations
that have not been sampled. Predictive statistical
techniques, such as generalized linear models
(GLMs) (McCullagh and NeIder, 1989), generalized additive models (GAMs) (Hastie and Tibshirani, 1990), and classification and regression trees
(CARTs) (Breiman et aI., 1984), are powerful and
flexible tools for generating quantitative relationships between species or communities and environmental factors (driving variables) using locationspecific data sets (Nicholls, 1989, 1991a, b; Yee
and Mitchell, 1991; Franklin 1998). Such models
can be used to predict biotic distributions in unsampled areas if the environmental characteristics
of these areas are known or can be derived from
maps or simulation models. The reliability of spatial prediction is increased when the frequency of
observations is evenly distributed across the environmental space (Nicholls, 1989).
Data collection should include explicit consideration of how well the full range of environmental
factors is sampled when the use of statistical models for prediction to unsampled areas is anticipated.
The SSRS and MSSRS schemes, including gradsects, with their purpose of covering the full range
of biotic variability over the range of environmental variability, can generate the required data. Logistical constraints may make complete coverage
unattainable, but the evenness of coverage determines the extent to which the prediction is an interpolation within the range of the data, rather than
an extrapolation outside its range (Margules and
Stein, 1989). The statistical properties needed to
implement the data for a specific application will
constrain the selection of a sampling strategy (see
Chapter 6 and Section 7.4).
Another example of data extrapolation is He et
aI.' s (1998) integration of TM data and PIA data
in a GIS environment (see Section 7.3.4) to represent the distribution of dominant tree species by
age classes and their associated species.
7.9 Conclusions
Several attributes of IREAs restrict the use of classical sampling design methodologies: (1) the assessment areas are large and accessibility may be
restricted in some locations; (2) data need to be collected for multiple components of ecosystems at
multiple spatial scales; (3) data are used for multiple purposes; and (4) strict deadlines and budgetary
and political constraints limit the resources and
time allocated to de novo sampling design and field
survey. IREA data collection must focus on the
multi scale spatial heterogeneity of multiple ecosystem components, in contrast to the purpose of classical sampling design. In this chapter, we have reviewed and discussed four elements of data
collection for IREAs: their general data requirements, appropriate standard data collection procedures, the use of existing data, and the trade-offs
among statistical theory, logistical benefits, and efficiency. Identification of specific sampling strategies, including the number, shape, and configuration of samples, arises from the purpose of the
IREA and from cost-effectiveness criteria. The following are general practical guidelines for the data
collection phase of IREAs:
1. Define the scale and the purpose of the survey
as clearly as possible. This will determine the
grain, interval, and extent of the sampling design. The lREA characterization defined by the
sampling extent may differ from the initial assessment area.
2. Review existing data and use them where possible for analysis and as templates for designing new surveys. Recognize that the gain realized by establishing databases for existing data
may very well offset the cost.
3. Consider the potential of SS schemes, including
SSRSs and MSSRSs, as cost-effective and efficient sampling designs. Systematic comparison
of the pros and cons of simple RSs versus SSs
should be conducted in the light of specific applications.
4. If an SS technique is used, match the scale of
the stratifying variables with the scale and purpose of the IREA. Selection of inappropriate
variables or scale of the data will diminish the
efficiency of the design.
5. Choose sample size and configuration to detect
nested spatial structures.
6. Assess the representativeness of the data collected and test the efficiency of the design, as
appropriate.
7. Use suitable models (e.g., GLMs, GAMs, and
CARTs) or other methods to extrapolate the collected data.
Any of the references on the topics discussed in
this chapter can provide additional sources of information. Good comparisons and discussions of
the various sampling strategies are found in Neldner et aI. (1995) and Neave et aI. (1997). Further
details on nested designs and references are available in Fortin et aI. (1989), Bellehumeur and Legendre (1998), and Legendre and Legendre (1998).
General Data Collection and Sampling Design Considerations
that have not been sampled. Predictive statistical
techniques, such as generalized linear models
(GLMs) (McCullagh and NeIder, 1989), generalized additive models (GAMs) (Hastie and Tibshirani, 1990), and classification and regression trees
(CARTs) (Breiman et aI., 1984), are powerful and
flexible tools for generating quantitative relationships between species or communities and environmental factors (driving variables) using locationspecific data sets (Nicholls, 1989, 1991a, b; Yee
and Mitchell, 1991; Franklin 1998). Such models
can be used to predict biotic distributions in unsampled areas if the environmental characteristics
of these areas are known or can be derived from
maps or simulation models. The reliability of spatial prediction is increased when the frequency of
observations is evenly distributed across the environmental space (Nicholls, 1989).
Data collection should include explicit consideration of how well the full range of environmental
factors is sampled when the use of statistical models for prediction to unsampled areas is anticipated.
The SSRS and MSSRS schemes, including gradsects, with their purpose of covering the full range
of biotic variability over the range of environmental variability, can generate the required data. Logistical constraints may make complete coverage
unattainable, but the evenness of coverage determines the extent to which the prediction is an interpolation within the range of the data, rather than
an extrapolation outside its range (Margules and
Stein, 1989). The statistical properties needed to
implement the data for a specific application will
constrain the selection of a sampling strategy (see
Chapter 6 and Section 7.4).
Another example of data extrapolation is He et
aI.' s (1998) integration of TM data and PIA data
in a GIS environment (see Section 7.3.4) to represent the distribution of dominant tree species by
age classes and their associated species.
7.9 Conclusions
Several attributes of IREAs restrict the use of classical sampling design methodologies: (1) the assessment areas are large and accessibility may be
restricted in some locations; (2) data need to be collected for multiple components of ecosystems at
multiple spatial scales; (3) data are used for multiple purposes; and (4) strict deadlines and budgetary
and political constraints limit the resources and
time allocated to de novo sampling design and field
survey. IREA data collection must focus on the
multi scale spatial heterogeneity of multiple ecosystem components, in contrast to the purpose of classical sampling design. In this chapter, we have reviewed and discussed four elements of data
collection for IREAs: their general data requirements, appropriate standard data collection procedures, the use of existing data, and the trade-offs
among statistical theory, logistical benefits, and efficiency. Identification of specific sampling strategies, including the number, shape, and configuration of samples, arises from the purpose of the
IREA and from cost-effectiveness criteria. The following are general practical guidelines for the data
collection phase of IREAs:
1. Define the scale and the purpose of the survey
as clearly as possible. This will determine the
grain, interval, and extent of the sampling design. The lREA characterization defined by the
sampling extent may differ from the initial assessment area.
2. Review existing data and use them where possible for analysis and as templates for designing new surveys. Recognize that the gain realized by establishing databases for existing data
may very well offset the cost.
3. Consider the potential of SS schemes, including
SSRSs and MSSRSs, as cost-effective and efficient sampling designs. Systematic comparison
of the pros and cons of simple RSs versus SSs
should be conducted in the light of specific applications.
4. If an SS technique is used, match the scale of
the stratifying variables with the scale and purpose of the IREA. Selection of inappropriate
variables or scale of the data will diminish the
efficiency of the design.
5. Choose sample size and configuration to detect
nested spatial structures.
6. Assess the representativeness of the data collected and test the efficiency of the design, as
appropriate.
7. Use suitable models (e.g., GLMs, GAMs, and
CARTs) or other methods to extrapolate the collected data.
Any of the references on the topics discussed in
this chapter can provide additional sources of information. Good comparisons and discussions of
the various sampling strategies are found in Neldner et aI. (1995) and Neave et aI. (1997). Further
details on nested designs and references are available in Fortin et aI. (1989), Bellehumeur and Legendre (1998), and Legendre and Legendre (1998).
