8.4 The Solution, by Analogy
ferences. That is, when organizing a small information resource, as is typical in an informationtechnology-poor environment, it is easy to overlook
the importance of details such as how (or it) the
metadata are made available and how the metadata
and data entities are organized. However, as each
data set is added, both the complexity of the organization and the need to provide effective organization increase. At some point, simple, implicit, or
monolithic management schemes no longer suffice
and must be replaced by more carefully planned hierarchical schemes. Currently, most ecosystem research and development laboratories have recognized how radical advances in basic computing,
networking, and storage technology can effectively
remove the low-level issues in entity storage, processing, and dissemination. However, the same labs
have generally been somewhat slower to recognize
the attendant problems with data organization. The
simple ability to compute faster or to post Web
pages for public access does not solve the problem
of data dissemination; instead, rapid proliferation
of and public access to information make the need
for coherent organization all the more critical.
To continue the analogy, note that the problem
of effectively managing a rapidly growing information resource to facilitate internal and external
use is hardly new. This is essentially the same problem that has been faced by libraries ever since technological advances made it easy to cheaply and reliably produce books. Librarians learned very early
on that simple, monolithic lists of books did not
represent an organized collection. Instead, they began to organize collections based on multiple indices of information collected about the books, resulting in the familiar by title, by author, and by
subject indices that virtually every library uses today. The information in these indices represents
metadata about the collection. Organizing these
metadata in multiple indices is important because
no single index meets all obvious access requirements of a disparate collection of users. Projects
that have attempted to modernize and computerize
libraries in recent years have expanded these indices and made them accessible on line, permitting
potential users to rapidly discern what information
is available and how it can be found. The parallels
with data dissemination should be obvious. The
first and most important goal should be to collect
the metadata systematically, organize it into one or
more easily searchable forms, and then, if possible,
put the results on line. Even if it is not feasible or
desirable to put the actual data sets on line, the primary goal in dissemination is to tell potential users
exactly what is available. This goal is best met by
113
providing convenient access to systematically indexed metadata.
To complete the analogy with traditional libraries, the amount of primary material that is
available to libraries in electronic or on-line from
other repositories is rapidly growing through subscription CD-ROMs, government publications, and
electronically published journals. The experience
in libraries indicates that having more source material available on line increases the dependency on
and need for accessible on-line indexing, as well
as the need to index on more content-specific keys.
The message for laboratories and organizations
that use and produce spatial data sets should be
clear: simply making data sets available on line is
not enough to help potential users select the correct data set or to understand the implications of
the selection in terms of key spatial attributes, such
as data location, source scale, or sampling error.
Furthermore, simply translating free-form textual
descriptions to on-line form and then autoindexing
the descriptive free text is not an effective way to
document spatial data holdings. Just as libraries
have devised specific indices (name, author, subject) to support the types of queries that dominate
their use patterns, effective indexing schemes must
be developed to support spatial data resources.
Several major initiatives are in progress to address issues of spatial data indexing on a large
scale, along with several other smaller projects designed to address specific technical issues. The sections that follow first review one such major initiative, the Federal Geographic Data Committee's
(FGDC) series of related programs to define standards for metadata content and format, to link individual spatial information repositories into recognized data clearinghouses, and to link the
holdings of these clearinghouses into contentspecific indices accessible to normal Web search
engines. Then the evolution of a smaller, more specialized data repository construction tool, the University of Montana Ecosystem Information System
(EIS) (Righter and Ford, 1994) is described. EIS
provides an example of how customized repositories can adapt and evolve to retain viability in the
current Web-dominated data dissemination environment.
As stated on their Web home page (http://
www.fgdc.gov), the primary purpose of the FGDC
is to coordinate the development of the National
Spatial Data Infrastructure (NSDI) through policies, standards, procedures, and initiatives. The
FGDC is a collaborative effort supported by 16 federal agencies, along with a variety of other state,
local, and tribal governments, universities, and pri-
Précédent

- 123/539

Suivant