216
Fuzzy Statistical and Modeling Approach to Ecological Assessments
15.4.1 Fuzzy Classifications
In fuzzy modeling techniques, many methods can
be easily used for ecological data analysis (e.g., see
Zimmermann, 1991; Bandemer and Nather, 1992).
Among them, fuzzy classification (or fuzzy clustering or pattern recognition) is the most commonly
used fuzzy data analysis in clustering of ecotoxicological data (Frederichs et aI., 1996), vegetation
analysis (Banyikwa et aI., 1990), Landsat TM data
classification for soil moisture estimates (Lindsey
et aI., 1992), GIS-based soil mapping (McBratney
et aI., 1992; Odeh et aI., 1992; Zhu et aI., 1996),
the impact of winter climate on crop (Chen et aI.,
1992), downscaling from GCMs to local climate
(Bardossy, 1994), ecological site classification
(Ayoubi et aI., 1991; EI-Shishiny and Ghabbour,
1991; Owsinski and Zadrozny, 1991), and land
suitability analysis (Hall et aI., 1992).
Fuzzy classification or clustering is the search
for structure in incomplete and uncertain data.
Given any finite data set X of objects, the problem
of clustering in X is to assign object labels that identify natural subgroups in the set. Because the data
are unlabeled, this problem is often called unsupervised learning, the word learning here referring
to learning the correct labels for good subgroups.
In tum, it should be clear that good subgroups implies that we have a question or problem in mind
that finding subgroups of some kind will answer,
help describe the process, and so forth. The objective is to partition X into a certain number (c) of
natural and homogeneous subsets, where the elements of each set are as similar as possible to each
other and at the same time as different from those
of the other sets as possible. The number (c) can
be fixed beforehand or may result as a consequence
of ecological or mathematical constraints. Because
the technique of clustering is unsupervised (i.e., we
do not have any information on the class structures
or the labeled samples of classes), clustering algorithms attempt to partition X based on certain assumptions and criteria. The output partition, which
includes both number of classes (c) and the class
membership structures, depends on the criteria that
are used to control the clustering algorithm.
Many algorithms have been developed to obtained hard clusters from a given data set. Among
these, the c-means algorithms and their generalizations, the ISODAT A (Iterative Self-Organizing
Data Analysis Techniques A) clustering methods,
are probably the most widely used. The c-means
algorithms assume that c is known, whereas c is
unknown in the case of ISODATA algorithms. By
allowing algorithms to assign each object a partial
or distributed membership in each of the c clusters,
fuzzy clustering offers special advantages over conventional algorithms. Sometime we refer to fuzzy
classification as supervised learning. The reference
cited in the first paragraph basically used fuzzy cmeans algorithms.
Let us look at a hypothetical example, habitat
comparison and evaluation. Let us assume that we
have five different habitats U = {Xl,X2,X3,X4,XS}
with four different indices for evaluation. Based on
some criteria, experts may use 1 to 5 (excellent) to
rank these indices. Now we assume that we have
Xl = (Xll,X12,X13,XI4) = (5,5,3,2)
X2 = (X21,X22,X23,X24) = (2,3,4,5)
X3 = (X31,X32,X33,X34) = (5,5,2,3)
X4 = (X4j,X42,X43,X44) = (1,5,3,1)
Xs = (XSl,XS2,XS3,XS4) = (2,4,5,1)
We can use simple distance measure to construct
membership function as follows:
{
I,
if i = j
rij = 1 - o.I±lxik - xjkl, ifi * j i,j = 1, ... ,5
k=i
to establish the fuzzy similarity matrix:
0.1
1
0.8 0.5
0.1 0.2
1 0.3
1
0.31
0.4
0.1
0.6
1
Mathematically, to have partitioning, we need to
obtain the fuzzy equivalent matrix through, R ~
R2 ~ R4 ~ R 8 , that is,
R' ~R'~R" {
0.4 0.8 0.5 0.5
1 0.4 0.4 0.4
1 0.5 0.5
1 0.6
1
If we set the a-level equal to 0.6, we have three
classes, {Xl,X3}, {X2}, and {X4,XS}; when you
change the a-level, you can have different classes.
Based on these results, we can say that habitats 1
and 3, and 4 and 5 have very similar quality, respectively; but habitat 2 has its own unique condition.
The work done by Hall et a1. (1992) provides a
good example for using the fuzzy classification
technique to analyze the land-use suitability of the
Cimanuk watershed in northwest Java, Indonesia.
It is not too difficult to extend their work to general ecological landscape assessment by using an
Fuzzy Statistical and Modeling Approach to Ecological Assessments
15.4.1 Fuzzy Classifications
In fuzzy modeling techniques, many methods can
be easily used for ecological data analysis (e.g., see
Zimmermann, 1991; Bandemer and Nather, 1992).
Among them, fuzzy classification (or fuzzy clustering or pattern recognition) is the most commonly
used fuzzy data analysis in clustering of ecotoxicological data (Frederichs et aI., 1996), vegetation
analysis (Banyikwa et aI., 1990), Landsat TM data
classification for soil moisture estimates (Lindsey
et aI., 1992), GIS-based soil mapping (McBratney
et aI., 1992; Odeh et aI., 1992; Zhu et aI., 1996),
the impact of winter climate on crop (Chen et aI.,
1992), downscaling from GCMs to local climate
(Bardossy, 1994), ecological site classification
(Ayoubi et aI., 1991; EI-Shishiny and Ghabbour,
1991; Owsinski and Zadrozny, 1991), and land
suitability analysis (Hall et aI., 1992).
Fuzzy classification or clustering is the search
for structure in incomplete and uncertain data.
Given any finite data set X of objects, the problem
of clustering in X is to assign object labels that identify natural subgroups in the set. Because the data
are unlabeled, this problem is often called unsupervised learning, the word learning here referring
to learning the correct labels for good subgroups.
In tum, it should be clear that good subgroups implies that we have a question or problem in mind
that finding subgroups of some kind will answer,
help describe the process, and so forth. The objective is to partition X into a certain number (c) of
natural and homogeneous subsets, where the elements of each set are as similar as possible to each
other and at the same time as different from those
of the other sets as possible. The number (c) can
be fixed beforehand or may result as a consequence
of ecological or mathematical constraints. Because
the technique of clustering is unsupervised (i.e., we
do not have any information on the class structures
or the labeled samples of classes), clustering algorithms attempt to partition X based on certain assumptions and criteria. The output partition, which
includes both number of classes (c) and the class
membership structures, depends on the criteria that
are used to control the clustering algorithm.
Many algorithms have been developed to obtained hard clusters from a given data set. Among
these, the c-means algorithms and their generalizations, the ISODAT A (Iterative Self-Organizing
Data Analysis Techniques A) clustering methods,
are probably the most widely used. The c-means
algorithms assume that c is known, whereas c is
unknown in the case of ISODATA algorithms. By
allowing algorithms to assign each object a partial
or distributed membership in each of the c clusters,
fuzzy clustering offers special advantages over conventional algorithms. Sometime we refer to fuzzy
classification as supervised learning. The reference
cited in the first paragraph basically used fuzzy cmeans algorithms.
Let us look at a hypothetical example, habitat
comparison and evaluation. Let us assume that we
have five different habitats U = {Xl,X2,X3,X4,XS}
with four different indices for evaluation. Based on
some criteria, experts may use 1 to 5 (excellent) to
rank these indices. Now we assume that we have
Xl = (Xll,X12,X13,XI4) = (5,5,3,2)
X2 = (X21,X22,X23,X24) = (2,3,4,5)
X3 = (X31,X32,X33,X34) = (5,5,2,3)
X4 = (X4j,X42,X43,X44) = (1,5,3,1)
Xs = (XSl,XS2,XS3,XS4) = (2,4,5,1)
We can use simple distance measure to construct
membership function as follows:
{
I,
if i = j
rij = 1 - o.I±lxik - xjkl, ifi * j i,j = 1, ... ,5
k=i
to establish the fuzzy similarity matrix:
0.1
1
0.8 0.5
0.1 0.2
1 0.3
1
0.31
0.4
0.1
0.6
1
Mathematically, to have partitioning, we need to
obtain the fuzzy equivalent matrix through, R ~
R2 ~ R4 ~ R 8 , that is,
R' ~R'~R" {
0.4 0.8 0.5 0.5
1 0.4 0.4 0.4
1 0.5 0.5
1 0.6
1
If we set the a-level equal to 0.6, we have three
classes, {Xl,X3}, {X2}, and {X4,XS}; when you
change the a-level, you can have different classes.
Based on these results, we can say that habitats 1
and 3, and 4 and 5 have very similar quality, respectively; but habitat 2 has its own unique condition.
The work done by Hall et a1. (1992) provides a
good example for using the fuzzy classification
technique to analyze the land-use suitability of the
Cimanuk watershed in northwest Java, Indonesia.
It is not too difficult to extend their work to general ecological landscape assessment by using an
