7 A Quality-enabled Spatial Integration System
135
the global schema, and query rewriting is performed with the Bucket algorithm [25].
Another algorithm, MiniCon [34], extended Bucket in using the input/output common variables between the query subgoals to reduce the bucket size. MiniCon reduces the query rewriting time and returns some solutions ignored by Bucket. The
Styx algorithm is an ontology-based algorithm which uses the parent–child dependencies of query variables for query decomposition.
As for the GIS community, essential work has focused on interoperability aspects [10, 19]. Most of the approaches propose to enrich the data models in order
to conform to a “unified model”, and the creation of the open geospatial consortium
(OGC) [32] is the most visible output of this trend. Other research work has focused
on schema integration, for example [12], object fusion [3], or ontology-based GIS
integration: Chapter 6 provides a clear description of ontology usage in the field of
data integration.
7.2.2 Data Quality
Data quality descriptions are crucial for the development of an organized business
with geographical data. Potential users of data must understand beforehand the quality of the data they intend to acquire. One of the issues is to assign a meaning to the
geographical data in a context—a standard problem now lumped together with other
similar problems in metadata description. The standardized descriptions, however—
in an equally long tradition—describe the quality of data by giving details of the data
collection process, so-called lineage data (e.g. Dublin Core). For some aspects of
spatial data quality, quantified measures are often used: for instance, the ISO 191135 standards [20, 21] propose about 15 (mostly quantifiable) quality elements and
sub-elements. The issue is exacerbated when it relates to the manipulation of multiple heterogeneous GIS, each of them providing (or not) some quality criteria and
information.
There is a growing awareness of the data quality problem in both DB and GIS
research communities. Work in DB projects have been interested in several aspects
of quality-based data integration techniques [27], [29], [17]. Within the GIS community, several geographical data quality models have been suggested by organizations
such as FGDC [13], IGN [11] and converge to an ISO/TC 211 model [22].
Classical integration approaches assume that data stored at different sources have
the same high quality and correctness. Actually, data sources are of different qualities
but sufficient for the applications they were initially dedicated to. However, they may
lead to low-quality integrated data in the context of a heterogeneous application.
Furthermore, the “high quality” definition is subjective.
7.3 Motivating Example
The example is drawn from a real geographical data integration problem being
studied in the REV!GIS project [35].
135
the global schema, and query rewriting is performed with the Bucket algorithm [25].
Another algorithm, MiniCon [34], extended Bucket in using the input/output common variables between the query subgoals to reduce the bucket size. MiniCon reduces the query rewriting time and returns some solutions ignored by Bucket. The
Styx algorithm is an ontology-based algorithm which uses the parent–child dependencies of query variables for query decomposition.
As for the GIS community, essential work has focused on interoperability aspects [10, 19]. Most of the approaches propose to enrich the data models in order
to conform to a “unified model”, and the creation of the open geospatial consortium
(OGC) [32] is the most visible output of this trend. Other research work has focused
on schema integration, for example [12], object fusion [3], or ontology-based GIS
integration: Chapter 6 provides a clear description of ontology usage in the field of
data integration.
7.2.2 Data Quality
Data quality descriptions are crucial for the development of an organized business
with geographical data. Potential users of data must understand beforehand the quality of the data they intend to acquire. One of the issues is to assign a meaning to the
geographical data in a context—a standard problem now lumped together with other
similar problems in metadata description. The standardized descriptions, however—
in an equally long tradition—describe the quality of data by giving details of the data
collection process, so-called lineage data (e.g. Dublin Core). For some aspects of
spatial data quality, quantified measures are often used: for instance, the ISO 191135 standards [20, 21] propose about 15 (mostly quantifiable) quality elements and
sub-elements. The issue is exacerbated when it relates to the manipulation of multiple heterogeneous GIS, each of them providing (or not) some quality criteria and
information.
There is a growing awareness of the data quality problem in both DB and GIS
research communities. Work in DB projects have been interested in several aspects
of quality-based data integration techniques [27], [29], [17]. Within the GIS community, several geographical data quality models have been suggested by organizations
such as FGDC [13], IGN [11] and converge to an ISO/TC 211 model [22].
Classical integration approaches assume that data stored at different sources have
the same high quality and correctness. Actually, data sources are of different qualities
but sufficient for the applications they were initially dedicated to. However, they may
lead to low-quality integrated data in the context of a heterogeneous application.
Furthermore, the “high quality” definition is subjective.
7.3 Motivating Example
The example is drawn from a real geographical data integration problem being
studied in the REV!GIS project [35].
