2009; Hu and Weng, 2011). These fine-spatial-resolution images contain rich
spatial information, providing a greater potential to extract much more detailed
thematic information (e.g., LULC), cartographic features (buildings and roads), and
metric information with stereo-images (e.g., height and area). These information
and cartographic characteristics are highly beneficial to estimating and mapping
impervious surfaces in the urban areas. However, some new problems come with
these image data, notably shadows caused by topography, tall buildings, or trees
(Dare, 2005) and the high spectral variation within the same land cover class (Hsieh
et al., 2001). Shadows obscure impervious surfaces underneath and thus increase the
difficulty to extract both thematic and cartographic information. These disadvantages
may lower image classification accuracy if classifiers used cannot effectively handle
them (Irons et al., 1985; Cushnie, 1987). In order to make full use of the rich spatial
information inherent in fine-spatial-resolution data, it is necessary to minimize the
negative impact of high intraspectral variation. Algorithms that use the combined
spectral and spatial information may be especially effective for impervious surface
extraction in the urban areas (Lu and Weng, 2007).
Per-pixel classifications prevail in the previous remote sensing literature, in which
each pixel is assigned to one category and land cover (or other themes) classes are
mutually exclusive. Per-pixel classification algorithms are sometimes referred to as
“hard” classifiers. Due to the heterogeneity of landscapes (particularly in urban
landscapes) and the limitation in spatial resolution of remote sensing imagery, mixed
pixels are common in medium- and coarse-spatial-resolution data. However, the
proportion of mixed pixels is significantly reduced in a high-resolution satellite image
scene. The presence of mixed pixels has been recognized as a major problem affecting
the effective use of per-pixel classifiers (Fisher, 1997; Cracknell, 1998). The mixedpixel problem results from the fact that the observational scale (i.e., spatial resolution)
fails to correspond to the spatial characteristics of the target (Mather, 1999). Strahler
et al. (1986) defined H- and L-resolution scene models based on the relationship
between the size of the scene elements and the resolution cell of the sensor. The scene
elements in the L-resolution model are smaller than the resolution cells and are thus
not detectable. When the objects in the scene become increasingly smaller than the
resolution cell size, they may no longer be regarded as individual objects. Hence, the
reflectance measured by the sensor may be treated as the sum of interactions among
various types of scene elements as weighted by their relative proportions (Strahler
et al., 1986). This is what happens with medium-resolution imagery, such as those of
Landsat TM or ETM+ , ASTER, SPOT, and Indian satellites, applied for urban
mapping. As the spatial resolution interacts with the fabric of urban landscapes, the
problem of mixed pixels is created. Such a mixture becomes especially prevalent in
residential areas where buildings, roads, trees, lawns, and water can all lump together
into a single pixel (Epstein et al., 2002). The low accuracy of image classification in
urban areas reflects, to a certain degree, the inability of traditional per-pixel classifiers
to handle composite signatures. Therefore, the “soft” approach of image classifications has been developed, in which each pixel is assigned a class membership of each
land cover type rather than a single label (Wang, 1990). Different approaches have
been used to derive a soft classifier, including fuzzy-set theory, Dempster–Shafer
66
ON THE ISSUE OF SCALE IN URBAN REMOTE SENSING
spatial information, providing a greater potential to extract much more detailed
thematic information (e.g., LULC), cartographic features (buildings and roads), and
metric information with stereo-images (e.g., height and area). These information
and cartographic characteristics are highly beneficial to estimating and mapping
impervious surfaces in the urban areas. However, some new problems come with
these image data, notably shadows caused by topography, tall buildings, or trees
(Dare, 2005) and the high spectral variation within the same land cover class (Hsieh
et al., 2001). Shadows obscure impervious surfaces underneath and thus increase the
difficulty to extract both thematic and cartographic information. These disadvantages
may lower image classification accuracy if classifiers used cannot effectively handle
them (Irons et al., 1985; Cushnie, 1987). In order to make full use of the rich spatial
information inherent in fine-spatial-resolution data, it is necessary to minimize the
negative impact of high intraspectral variation. Algorithms that use the combined
spectral and spatial information may be especially effective for impervious surface
extraction in the urban areas (Lu and Weng, 2007).
Per-pixel classifications prevail in the previous remote sensing literature, in which
each pixel is assigned to one category and land cover (or other themes) classes are
mutually exclusive. Per-pixel classification algorithms are sometimes referred to as
“hard” classifiers. Due to the heterogeneity of landscapes (particularly in urban
landscapes) and the limitation in spatial resolution of remote sensing imagery, mixed
pixels are common in medium- and coarse-spatial-resolution data. However, the
proportion of mixed pixels is significantly reduced in a high-resolution satellite image
scene. The presence of mixed pixels has been recognized as a major problem affecting
the effective use of per-pixel classifiers (Fisher, 1997; Cracknell, 1998). The mixedpixel problem results from the fact that the observational scale (i.e., spatial resolution)
fails to correspond to the spatial characteristics of the target (Mather, 1999). Strahler
et al. (1986) defined H- and L-resolution scene models based on the relationship
between the size of the scene elements and the resolution cell of the sensor. The scene
elements in the L-resolution model are smaller than the resolution cells and are thus
not detectable. When the objects in the scene become increasingly smaller than the
resolution cell size, they may no longer be regarded as individual objects. Hence, the
reflectance measured by the sensor may be treated as the sum of interactions among
various types of scene elements as weighted by their relative proportions (Strahler
et al., 1986). This is what happens with medium-resolution imagery, such as those of
Landsat TM or ETM+ , ASTER, SPOT, and Indian satellites, applied for urban
mapping. As the spatial resolution interacts with the fabric of urban landscapes, the
problem of mixed pixels is created. Such a mixture becomes especially prevalent in
residential areas where buildings, roads, trees, lawns, and water can all lump together
into a single pixel (Epstein et al., 2002). The low accuracy of image classification in
urban areas reflects, to a certain degree, the inability of traditional per-pixel classifiers
to handle composite signatures. Therefore, the “soft” approach of image classifications has been developed, in which each pixel is assigned a class membership of each
land cover type rather than a single label (Wang, 1990). Different approaches have
been used to derive a soft classifier, including fuzzy-set theory, Dempster–Shafer
66
ON THE ISSUE OF SCALE IN URBAN REMOTE SENSING
