development of models is also based on ecological theory and systems
principles, which may help to account for interactions and multivariate
relationships that may be hidden in linear representations of field data.
Model coefficients based solely on linear field data may force the model
to behave in regimes for which it was not initially designed. At a minimum,
the time, space, and characterization of the data must be understood in
the context in which the data were acquired. This placing of the data in a
contextual setting will provide a beginning to understand the caveats
of the data/model/decision information flow. Know your data, know your
model!
In addition, at USEPA we find that more data are not necessarily better
data. With the Internet, we can easily become saturated with data and information. Web sites abound with access to databases that can easily be downloaded to an individual’s computer. With such an abundance of data,
determination of the applicability of individual data sets to the modeling
and decision-making process can be difficult. It is important to understand
the ramifications of making decisions without good data. Is such a situation
any better than making wrong decisions with good data?
In the flow of data from field measurements to the resource manager or
decision maker, it is imperative that contextual meaning is carried along
with the flow of data. Many databases that have been in existence for a long
period of time may appear to be useful for things other than that for which
they were developed. An example is the STORET database that the
USEPA maintains. At first glance, the data would appear to be useful in a
geographic information system (GIS) context and could be used to display
water-quality information spatially.The database has locational information
in the attributes of sites where water-quality measurements have been
made. The accuracy of the locational information is often not well documented. This drawback, however, is not the major problem of using databases like this in a GIS framework. The major problem is that the data do
not always fit any rational sampling model that would allow them to be spatially mapped. The data are representative of individual sites with no spatially integrated sampling scheme. Data are often measured to fulfill permit
requirements, enforcement activities, background sampling, and other ad
hoc schemes. Thus, it would be very easy for uninformed GIS modelers to
use STORET data in a manner that would give undesirable results. Other
legacy databases have similar problems.
The necessity of metadata cannot be overemphasized. At each step in the
data-gathering and modeling process, adequate documentation must be
recorded to provide a firm foundation for the data processing and the following decision-making process. Metadata can be streamlined by using a
form-based structure to record pertinent information at each step. The
metadata should always be carried with the data. These data about data are
often at least as important as the base data because they provide the context
of the base data.
9. Data and Information Issues: Communication Is the Key
173
principles, which may help to account for interactions and multivariate
relationships that may be hidden in linear representations of field data.
Model coefficients based solely on linear field data may force the model
to behave in regimes for which it was not initially designed. At a minimum,
the time, space, and characterization of the data must be understood in
the context in which the data were acquired. This placing of the data in a
contextual setting will provide a beginning to understand the caveats
of the data/model/decision information flow. Know your data, know your
model!
In addition, at USEPA we find that more data are not necessarily better
data. With the Internet, we can easily become saturated with data and information. Web sites abound with access to databases that can easily be downloaded to an individual’s computer. With such an abundance of data,
determination of the applicability of individual data sets to the modeling
and decision-making process can be difficult. It is important to understand
the ramifications of making decisions without good data. Is such a situation
any better than making wrong decisions with good data?
In the flow of data from field measurements to the resource manager or
decision maker, it is imperative that contextual meaning is carried along
with the flow of data. Many databases that have been in existence for a long
period of time may appear to be useful for things other than that for which
they were developed. An example is the STORET database that the
USEPA maintains. At first glance, the data would appear to be useful in a
geographic information system (GIS) context and could be used to display
water-quality information spatially.The database has locational information
in the attributes of sites where water-quality measurements have been
made. The accuracy of the locational information is often not well documented. This drawback, however, is not the major problem of using databases like this in a GIS framework. The major problem is that the data do
not always fit any rational sampling model that would allow them to be spatially mapped. The data are representative of individual sites with no spatially integrated sampling scheme. Data are often measured to fulfill permit
requirements, enforcement activities, background sampling, and other ad
hoc schemes. Thus, it would be very easy for uninformed GIS modelers to
use STORET data in a manner that would give undesirable results. Other
legacy databases have similar problems.
The necessity of metadata cannot be overemphasized. At each step in the
data-gathering and modeling process, adequate documentation must be
recorded to provide a firm foundation for the data processing and the following decision-making process. Metadata can be streamlined by using a
form-based structure to record pertinent information at each step. The
metadata should always be carried with the data. These data about data are
often at least as important as the base data because they provide the context
of the base data.
9. Data and Information Issues: Communication Is the Key
173
