122
Matt Duckham and Mike Worboys
We can imagine what might have happened if some of the words on the Rosetta
Stone had been incorrectly drafted or inscribed. It is possible that such inaccuracies
would lead to incorrect lexicographic inferences, especially in the case of systematic
inaccuracies. To guard against inaccuracy, it is important to ensure the extensions
used in the inference process are large enough such that examples of incorrect correspondences due to random inaccuracies will be greatly under-represented when
compared with examples of correct correspondences. In terms of a rosetta system,
the situation is a little more complex. However, in principle, the possibility of random inaccuracies is another reason why the rosetta systems are fundamentally data
hungry: the more examples used in the inference process, the more likely it is that
these examples will provide a basis for valid inferences.
In addition to spatial inaccuracy, inaccuracies may occasionally occur within the
taxonomy itself (e.g. where one category is incorrectly labeled or incorrectly positioned within the taxonomy). Since the taxonomy is central to the fusion process, it
is difficult to see how the automated fusion process described here (or indeed any
of the fusion systems encountered in this chapter) could hope to effectively combat
such inaccuracies.
6.5.2 Imprecision
Imprecision, a lack of detail in information, is another intrinsic feature of geographical information. Imprecision leads to granularity: the existence of “clumps” or
“grains” in the data. The granularity at which geographical phenomena are represented strongly influences what features are observed. Like inaccuracy, heterogeneous levels of granularity degrade the reliability of the inductive inference process.
For example, imagine that land cover data set A has been collected at a coarser level
of spatial granularity than data set B. Then it will be likely that the detailed features found in data set B will simply not be represented in data set A (such as small
pockets of Woodland within the predominately Urban area that are represented in
data set B, but have no correspondent pockets in data set A). As a result, a naive
inductive inference process may again incorrectly infer that Woodland and Built-up
area are semantically overlapping, as in Fig. 6.5 (similar to the effects of inaccuracy in Fig. 6.4). As for inaccuracy, the fusion product in Fig. 6.5 is not particularly
informative, as it is essentially a simple overlay of the data.
It is difficult to say how the efforts to decipher Egyptian hieroglyphics would
have fared if the different versions of the official decree on the Rosetta Stone contained different levels of detail about the official declarations. The structure of natural
language does not make it easy to automatically infer relationships between texts at
different levels of detail. However, the spatial structure of geographical information
does make inferences between information sources at different levels of detail more
feasible (e.g. Sects. 6.5.4 and [20]).
In addition to spatial imprecision, it may be important to also consider the possibility of heterogeneity in taxonomic granularity. In this case, semantic differences
that are distinguished apart in the taxonomy for one data set may not be distinguished
in the taxonomy for a different data set. For example, the category Woodland is at
Précédent

- 117/317

Suivant