10.2 Data Appropriateness Concerns
Data appropriateness does not ask the question can these data be used for
this purpose, but instead asks the question should these data be used for
this purpose. To determine whether data should be used for a model, it is
necessary to understand the original purpose of data collection. There are
no cookbooks for determining when to use a certain data set except to say
that the data must not violate the assumptions of the model. The objectives
for collecting the data may not have been clearly articulated even if they
seemed clear to the project managers at initiation (Conquest et al. 1993).
Knowing whether a particular data set is appropriate for a given model is
based upon a clear understanding of the assumptions underlying both the
data and the model. The key to this understanding is achieved by evaluating the metadata, or “data about data.”
Metadata describe the “what, who, when, why, and how” of the data. The
importance of providing and using complete metadata with models and
data cannot be overestimated. Metadata are used to match up the assumptions and limitations of the data with those same factors for the model.
For spatial data, the Federal Geographic Data Committee (FGDC;
http://www.fgdc.gov) has developed minimum standards defining metadata.
Good metadata should contain a clear definition of source, units, underlying assumptions, variability, scale, and resolution. Data limitations and
qualifiers should be clearly stated or easily inferred from the metadata or
supplemental “README” documentation and should be visible to the
data user when accessing the data.
10.2.1 Source
Metadata should address the following questions: Who created the data?
For what purpose were the data created? Were the data recorded in the
field? If recorded in the field, how were the data collected? Were the data
derived from remotely sensed imagery or other GIS data? Were the data
simulated output from a model? These questions help determine the appropriateness of a data set for a specific use, with the purpose of data creation
being the key constraint on wise data use. For example, the U.S. Environmental Protection Agency’s (USEPA’s) EPA 303d (Impaired Water Quality
streams) data provide national coverage but were not designed to be a
nationally consistent data set. Individual states were not required to use the
same protocols and methods for identifying impaired waters, and so the
resulting data set is a mix of different reporting methods and different
criteria for classifying stream reaches. Analysis of these data, therefore,
should not be used to describe national trends, and even summary information could be very misleading. Furthermore, data with a spatial component (e.g., latitude and longitude coordinates) may not have been created
184
David Hohler et al.
Précédent

- 190/327

Suivant