8.2 Background Issues
Project management issues include, in the broadest sense, a wide range of scientific and practical
issues associated with actually conducting an ecological study at medium to large scales. We assume
that the primary purpose of such studies is to use
computer processing to perform some type of
analysis on existing digital data sets, which in tum
produces new derived data sets and ultimately produces final results that are (or are based on) digital data products. From this base follows a set of
information-related project management concerns,
such as what input data sets will be used and how
they will be collected, digitized if necessary, and
spatially coregistered (i.e., have elements in each
data set matched to common spatial points of reference). A related set of concerns exists for the project's output data products. Data processing implies
an additional set of concerns, including what computer hardware and software are used, what processing parameters are used with that hardware and
software, and how appropriate facilities are obtained, maintained, and scheduled. Thus the concerns in a modem ecological study of even modest
size must include both those pertinent to ecological and scientific issues and those pertinent to data
management and processing. The latter are more
traditionally associated with the domain of the centralized computer center. However, technological
advances allowing computing resources to be
widely distributed have also shifted the responsibility of addressing data management concerns to
the users of the distributed computing facilities.
The problems associated with maintaining general data repositories are not new. A wide variety
of literature is associated with this topic, starting
with a vast body of work describing how to construct a general-purpose database, typically in conjunction with a specific database modeling technique or commercial database system. Examples of
this work include standard texts such as Kroenke
(1998) and McFadden et al. (1999). Also pertinent
here is work describing general quality metrics for
data repositories (Ballow and Tayi, 1999). Another
common approach is to look specifically at the data
management issues associated with a particular
type of data set. Work related to the storage and
manipulation of image data sets is particularly
widespread, ranging from standard texts (Russ
1995) to entire conference proceedings (Rosenholm and Osterlund, 1997).
Only recently has work appeared that focuses on
the specific requirements of maintaining a variety
of specialized data sets to support ecosystem modeling and management. Examples include reports
from the Sequoia 2000 (Stonebraker, 1994; Frew,
109
1996) and GEMS (Bruegge et aI., 1995) projects.
Both are large-scale research projects involving the
design of information systems to support a diverse
collection of data sets, researchers, and ecosystem
analysis projects. Sequoia 2000 is notable for its
goals of utilizing high-speed network capabilities
to support data sets and users distributed across a
wide geographic area (Stonebraker, 1994; Frew,
1996). GEMS is notable for placing user requirements at the heart of its design, rather than implementation efficiency or efficacy (Bruegge et aI.,
1995). Although the software system architectures
developed in projects such as these provide interesting examples of how such systems can be structured, neither provides much detail on the complex
issues of data management faced by laboratories
and agencies with large legacy data collections.
More pertinent in this context are systems that address the specific problems faced in integrating existing data sets and databases with expanded, more
ambitious goals in modeling and information dissemination. Examples here include two ecosystem
information systems designed with both data managers and data clients in mind (Ford et aI., 1994;
Cowan et aI., 1996). These systems focus on allowing data managers to organize their existing
data collections by constructing new data indices,
rather than requiring the conversion of massive
amounts of data to a new database format. The systems also utilize Web facilities to provide access
to data collections through these indices. The former system is discussed in detail in Section 8.5.
If we evaluate the most recent literature, looking
specifically for evidence of how recent technological advances have forced data managers to view
the data management and dissemination problem in
different terms, several common themes emerge.
As a starting point for this discussion, it is helpful
to first divide the set of issues into subclasses that
reflect their importance in historical terms.
8.2.1 Data Processing
An obvious special class of concerns in any project is that associated with the actual processing, focusing on the details of the computing hardware
and software used to implement a particular form
of data analysis. In historical terms, computational
resources were relatively scarce. Software for spatial data analysis was both highly constrained and
highly nonstandard, existing primarily in the form
of locally written, nonportable code that was designed to be executed on locally available, highly
constrained hardware. As constraints on both hardware and software have receded due to technolog-
Précédent

- 119/539

Suivant