112
Whether phrased in terms of processing methodology or content and context, the problem is that a
complication arises. As the information resource
expands, it becomes increasingly difficult for the
owner of the information resource to handle the
methodological context effectively, but it also becomes more important that this context be maintained. Implicit ways of handling this information,
such as using file-naming conventions or "normal
practices," break down over time with staff
turnover and expanding information resources. It
becomes essential to somehow make the context
information explicit as an inherent part of the attendant data sets. The idea of maintaining data
about data, or metadata, in explicit form is hardly
new. What is still fairly new is the idea of making
metadata about digital data sets an inherent, distinguishable, and accessible part of these data sets.
Ideally, the metadata have characteristics such as
the following. First, metadata creation should be an
explicit part of the data derivation and data import
processes so that modelers and data managers will
understand the full impact of their most common
data management actions (i.e., actions that create
new data sets). Second, metadata should be easily
distinguishable from the core data set and accessible on line in a manner that is at least as easy as
accessing the data set itself. Finally, metadata
should be used to build content-based indices that
are searchable over the Internet in a manner analogous to the now-familiar Web search. Together,
these three characteristics can guarantee that metadata are reliable, readily available for index and
search engines, and usable as the basis for Weblike queries that are far more content specific and
thus yield a much higher quality of hits than the
simple-text indexing most common today.
The concept of maintaining some form of metadata is as old as the scientific method's explicit
focus on experimental methodology, but the implementation of this concept in our informationtechnology-rich environment has been spotty at
best. Too often the creation of metadata about digital data sets has been driven by specific dissemination requirements, rather than recognition on the
part of the deriving organization that their own interests are best served if they appropriately, routinely, and carefully document the context in which
derivations are performed. In the simplest cases,
limited plans for dissemination may result in minimal data annotation, by which information is included only in the form of written reports about and
summaries of the derived data. Note that such annotations are typically isolated from the data sets
themselves. At an intermediate level are explicit
Ecological Data Storage. Management, and Dissemination
techniques for metadata collection that involve both
annotation and a certain amount of data organization; the annotation is again separate from the data,
but the organization of the collection of data sets
provides implicit information about derivation
methodology. In contrast, at the most ambitious
level is a system that automatically collects and
electronically catalogs all metadata needed to document the processing used throughout the entire
study. For example, this best-practice system would
provide the complete set of records necessary to alIowan independent group to duplicate a study and
confirm or refute its results. There is clearly a wide
gap between normal and best possible practices.
The development of a system for automatic metadata collection, verification, and cataloging is essential; as rapidly as new data sets may be derived
on line, any nonautomated scheme for collecting
and cataloging metadata is unlikely to meet even
modest dissemination goals.
8.4 The Solution, by Analogy
To make the problem of trying to reuse data that
contain neither explicit nor implicit (organizational) metadata more concrete, consider a simple
Web dissemination scenario. Suppose that a data
manager builds a Web site with only a single short
page (say 30 lines long) containing actual links to
ten data sets. The single page per site, short textual
description, and small number of data sets involved
all combine to provide an implicit context for potential users that provides them with some information about what they actually get when they
download the site's data. Contrast this with a data
collection that has grown to the point where the
Web site must describe and link to a collection of
1000 datasets. It is not feasible to simply list or
show 1000 links on a page; furthermore, a simple,
monolithic list of 1000 entities is clearly inadequate
to allow a potential user to find any specific item
of interest. Although a user can rapidly download
any of the 1000 alternatives, his or her problem is
that of determining which alternative to pick. That
is, how can a potential user easily distinguish two
or more data sets with names or short descriptions
that are similar?
Creating an effective Web distribution for a
larger set of items requires addressing two concerns. First, it is critical to provide some organization of the entities into smaller groups. Second, this
organization must reflect key characteristics of the
entities, as represented by one or more metadata
characteristics that highlight similarities and/or dif-
Whether phrased in terms of processing methodology or content and context, the problem is that a
complication arises. As the information resource
expands, it becomes increasingly difficult for the
owner of the information resource to handle the
methodological context effectively, but it also becomes more important that this context be maintained. Implicit ways of handling this information,
such as using file-naming conventions or "normal
practices," break down over time with staff
turnover and expanding information resources. It
becomes essential to somehow make the context
information explicit as an inherent part of the attendant data sets. The idea of maintaining data
about data, or metadata, in explicit form is hardly
new. What is still fairly new is the idea of making
metadata about digital data sets an inherent, distinguishable, and accessible part of these data sets.
Ideally, the metadata have characteristics such as
the following. First, metadata creation should be an
explicit part of the data derivation and data import
processes so that modelers and data managers will
understand the full impact of their most common
data management actions (i.e., actions that create
new data sets). Second, metadata should be easily
distinguishable from the core data set and accessible on line in a manner that is at least as easy as
accessing the data set itself. Finally, metadata
should be used to build content-based indices that
are searchable over the Internet in a manner analogous to the now-familiar Web search. Together,
these three characteristics can guarantee that metadata are reliable, readily available for index and
search engines, and usable as the basis for Weblike queries that are far more content specific and
thus yield a much higher quality of hits than the
simple-text indexing most common today.
The concept of maintaining some form of metadata is as old as the scientific method's explicit
focus on experimental methodology, but the implementation of this concept in our informationtechnology-rich environment has been spotty at
best. Too often the creation of metadata about digital data sets has been driven by specific dissemination requirements, rather than recognition on the
part of the deriving organization that their own interests are best served if they appropriately, routinely, and carefully document the context in which
derivations are performed. In the simplest cases,
limited plans for dissemination may result in minimal data annotation, by which information is included only in the form of written reports about and
summaries of the derived data. Note that such annotations are typically isolated from the data sets
themselves. At an intermediate level are explicit
Ecological Data Storage. Management, and Dissemination
techniques for metadata collection that involve both
annotation and a certain amount of data organization; the annotation is again separate from the data,
but the organization of the collection of data sets
provides implicit information about derivation
methodology. In contrast, at the most ambitious
level is a system that automatically collects and
electronically catalogs all metadata needed to document the processing used throughout the entire
study. For example, this best-practice system would
provide the complete set of records necessary to alIowan independent group to duplicate a study and
confirm or refute its results. There is clearly a wide
gap between normal and best possible practices.
The development of a system for automatic metadata collection, verification, and cataloging is essential; as rapidly as new data sets may be derived
on line, any nonautomated scheme for collecting
and cataloging metadata is unlikely to meet even
modest dissemination goals.
8.4 The Solution, by Analogy
To make the problem of trying to reuse data that
contain neither explicit nor implicit (organizational) metadata more concrete, consider a simple
Web dissemination scenario. Suppose that a data
manager builds a Web site with only a single short
page (say 30 lines long) containing actual links to
ten data sets. The single page per site, short textual
description, and small number of data sets involved
all combine to provide an implicit context for potential users that provides them with some information about what they actually get when they
download the site's data. Contrast this with a data
collection that has grown to the point where the
Web site must describe and link to a collection of
1000 datasets. It is not feasible to simply list or
show 1000 links on a page; furthermore, a simple,
monolithic list of 1000 entities is clearly inadequate
to allow a potential user to find any specific item
of interest. Although a user can rapidly download
any of the 1000 alternatives, his or her problem is
that of determining which alternative to pick. That
is, how can a potential user easily distinguish two
or more data sets with names or short descriptions
that are similar?
Creating an effective Web distribution for a
larger set of items requires addressing two concerns. First, it is critical to provide some organization of the entities into smaller groups. Second, this
organization must reflect key characteristics of the
entities, as represented by one or more metadata
characteristics that highlight similarities and/or dif-
