6 Automated Geographical Information Fusion
125
a partition of space (i.e. the broad boundaries constitute an overlap between two or
more neighboring regions). Thus, the formal mechanisms currently being developed
for fusing data-containing regions with broad boundaries are generalizations of those
formalizations already discussed.
Whatever the formal structures used, the goal is to infer crisp semantic relationships between vague categories based on indeterminate spatial extents. For example,
we may be certain that a “Copse” is a sub-category of “Woodland,” even if both categories are vague. In the case of Fig. 6.6, we might devise new inference rules like
those in Sect. 6.3 that only consider the core of the extent of each category (those
parts of space that are classified as definitely belonging to the category). Conversely,
a weaker inference system could be developed by allowing semantic relationships to
be inferred where the core of one category is contained within the entirety of another
category.
6.5.4 Computation with Uncertain Data
From the discussion above, we can begin to suggest simple mechanisms for incorporating inaccuracy and imprecision into the automated information fusion process
(vagueness is the topic of current research). One such mechanism for incorporating
inaccuracy and granularity arises from noting that sliver polygons (resulting from
inaccuracy) or regions of fine-grained detail (resulting from fine granularity) are expected to make up a relatively small proportion of the entire regions being fused. For
example, if the overlap between two regions is smaller than 5% of the total area of
either regions, this might constitute evidence that the overlap arises from inaccuracy
in the input regions. Similarly, if the overlap between region A and region B is less
than 5% of the total area of region A and more than, say, 95% of the total area of
region B, this might constitute evidence that the overlap arises from heterogeneous
granularity in the data sets (i.e. region B is at a finer granularity than region A).
Consequently, setting thresholds for the proportion of overlap between two extensions of a category provides a basis for detecting spatial relationships that can be
attributed to inaccuracy or heterogeneous granularity. Spatial relationships that are
attributed to inaccuracy or imprecision then can be omitted from premises for the
inductive inference process. The effect of using such an approach is illustrated for
our example rosetta system in Fig. 6.7, based on Fig. 6.4. Here, the small sliver
overlap between the extents of Built-up area and Woodland comprises less than 5%
of the total area of these extents. This overlap is omitted from the inductive inference process, leading to a fused taxonomy as for Fig. 6.1. However, in the fused data
set, the omitted region (black region) then becomes unclassifiable (has no category
associated with it).
The thresholds can be set arbitrarily, or by a human user. As the thresholds increase, more overlaps are omitted from the inference process, usually leading to
more direct subsumption relationships in the fused taxonomy (cf. the taxonomies in
Figs. 6.4 and 6.7). Fused taxonomies containing more direct subsumption relationships are generally more desirable because they provide more new information about
the relationships between categories in the source taxonomies (the fused taxonomies
125
a partition of space (i.e. the broad boundaries constitute an overlap between two or
more neighboring regions). Thus, the formal mechanisms currently being developed
for fusing data-containing regions with broad boundaries are generalizations of those
formalizations already discussed.
Whatever the formal structures used, the goal is to infer crisp semantic relationships between vague categories based on indeterminate spatial extents. For example,
we may be certain that a “Copse” is a sub-category of “Woodland,” even if both categories are vague. In the case of Fig. 6.6, we might devise new inference rules like
those in Sect. 6.3 that only consider the core of the extent of each category (those
parts of space that are classified as definitely belonging to the category). Conversely,
a weaker inference system could be developed by allowing semantic relationships to
be inferred where the core of one category is contained within the entirety of another
category.
6.5.4 Computation with Uncertain Data
From the discussion above, we can begin to suggest simple mechanisms for incorporating inaccuracy and imprecision into the automated information fusion process
(vagueness is the topic of current research). One such mechanism for incorporating
inaccuracy and granularity arises from noting that sliver polygons (resulting from
inaccuracy) or regions of fine-grained detail (resulting from fine granularity) are expected to make up a relatively small proportion of the entire regions being fused. For
example, if the overlap between two regions is smaller than 5% of the total area of
either regions, this might constitute evidence that the overlap arises from inaccuracy
in the input regions. Similarly, if the overlap between region A and region B is less
than 5% of the total area of region A and more than, say, 95% of the total area of
region B, this might constitute evidence that the overlap arises from heterogeneous
granularity in the data sets (i.e. region B is at a finer granularity than region A).
Consequently, setting thresholds for the proportion of overlap between two extensions of a category provides a basis for detecting spatial relationships that can be
attributed to inaccuracy or heterogeneous granularity. Spatial relationships that are
attributed to inaccuracy or imprecision then can be omitted from premises for the
inductive inference process. The effect of using such an approach is illustrated for
our example rosetta system in Fig. 6.7, based on Fig. 6.4. Here, the small sliver
overlap between the extents of Built-up area and Woodland comprises less than 5%
of the total area of these extents. This overlap is omitted from the inductive inference process, leading to a fused taxonomy as for Fig. 6.1. However, in the fused data
set, the omitted region (black region) then becomes unclassifiable (has no category
associated with it).
The thresholds can be set arbitrarily, or by a human user. As the thresholds increase, more overlaps are omitted from the inference process, usually leading to
more direct subsumption relationships in the fused taxonomy (cf. the taxonomies in
Figs. 6.4 and 6.7). Fused taxonomies containing more direct subsumption relationships are generally more desirable because they provide more new information about
the relationships between categories in the source taxonomies (the fused taxonomies
