Chapter 2 . Unsupervised Artificial Neural Networks
29
U-matrix method is to compute the distance between two VUs located in two
adjacent hexagons. High value distances will be used as an indication of cluster
boundaries. To visualize the distances, new hexagons will be used and from the
initial grid a new one can be constructed, inserting a new hexagon between each
adjacent hexagon (Fig. 2.8 - b). So, if c columns and r rows make up the initial
grid of the output layer, the U-matrix is a matrix with (2c - 1) columns and (2r - 1)
rows where grey levels show the distance values (Fig. 2.8 - c). But distance values
are only available for the new hexagons so the unified distance matrix has to be
completed. For this purpose, in each hexagon including a virtual unit, a distance
has been added, calculated as the minimum of its adjacent hexagons (Fig. 2.8 - d).
If dark colours are used for large distances and light colours for short distances,
the U-matrix can be seen as alandscape displaying the distances between the VUs.
This landscape is formed with light plains separated by dark ravines. When, SUs
are mapped, the units in the plains are close to each other in the input layer so
these SUs are similar (for species abundance) and clusters becomes apparent (Fig.
2.8 - e).
An enhancement of this representation can be made: a triangle-based cubic
interpolation (Watson 1992) has been applied to the U-matrix. Then, a smooth
surface is obtained and displayed with the possibility of changing the brightness of
the figure. The brighter the display, the lower the number of clusters becoming
apparent.
The outcome of this process for the Wisconsin forests can be seen in Fig. 2.9.
In this way, on Fig. 2.9 - a, 4 clusters can be made out: BI (SUs 1,2, 3 and 4), B II
(SU 5), B", (SUs 7, 9 and 10), and B IV (SUs 6 and 8). Then on Fig. 2.9 - b, the SUs
can be grouped in 2 clusters CI (SUs 1,2,3,4 and 5) and C II (SUs 6, 7, 8, 9 and
10).
If we consider the U-matix computed with the Euclidean distance (Fig. 2.10), 3
clusters can be made out: EI (SUs 1,2,3,4 and 5), Eil (SU 7) and E", (SUs 6, 8, 9
and 10).
So the clustering method with the SOM can be summarized as folIows:
Teaching the SOM: the species abundance is computed for each VU.
Computing the U-matrix.
Mapping the SUs onto the U-matrix.
Making the clustering structure apparent for the human expert of the dataset by
selecting the brightness of the display.
2.4
Discussion
As already mentioned by Chon et al. (1996), high similarity between the SOM
results and the dendrogram (Fig. 2.1) may be observed. The 8 clusters (AI to A VIII )
previously defined with the SOM are those identified on the dendrogram by a
dotted line at distance 3, except for the SU 4 which is grouped with the SUs land
2 on the dendrogram. The U-matrix enhances the SOM and brings more accurate
29
U-matrix method is to compute the distance between two VUs located in two
adjacent hexagons. High value distances will be used as an indication of cluster
boundaries. To visualize the distances, new hexagons will be used and from the
initial grid a new one can be constructed, inserting a new hexagon between each
adjacent hexagon (Fig. 2.8 - b). So, if c columns and r rows make up the initial
grid of the output layer, the U-matrix is a matrix with (2c - 1) columns and (2r - 1)
rows where grey levels show the distance values (Fig. 2.8 - c). But distance values
are only available for the new hexagons so the unified distance matrix has to be
completed. For this purpose, in each hexagon including a virtual unit, a distance
has been added, calculated as the minimum of its adjacent hexagons (Fig. 2.8 - d).
If dark colours are used for large distances and light colours for short distances,
the U-matrix can be seen as alandscape displaying the distances between the VUs.
This landscape is formed with light plains separated by dark ravines. When, SUs
are mapped, the units in the plains are close to each other in the input layer so
these SUs are similar (for species abundance) and clusters becomes apparent (Fig.
2.8 - e).
An enhancement of this representation can be made: a triangle-based cubic
interpolation (Watson 1992) has been applied to the U-matrix. Then, a smooth
surface is obtained and displayed with the possibility of changing the brightness of
the figure. The brighter the display, the lower the number of clusters becoming
apparent.
The outcome of this process for the Wisconsin forests can be seen in Fig. 2.9.
In this way, on Fig. 2.9 - a, 4 clusters can be made out: BI (SUs 1,2, 3 and 4), B II
(SU 5), B", (SUs 7, 9 and 10), and B IV (SUs 6 and 8). Then on Fig. 2.9 - b, the SUs
can be grouped in 2 clusters CI (SUs 1,2,3,4 and 5) and C II (SUs 6, 7, 8, 9 and
10).
If we consider the U-matix computed with the Euclidean distance (Fig. 2.10), 3
clusters can be made out: EI (SUs 1,2,3,4 and 5), Eil (SU 7) and E", (SUs 6, 8, 9
and 10).
So the clustering method with the SOM can be summarized as folIows:
Teaching the SOM: the species abundance is computed for each VU.
Computing the U-matrix.
Mapping the SUs onto the U-matrix.
Making the clustering structure apparent for the human expert of the dataset by
selecting the brightness of the display.
2.4
Discussion
As already mentioned by Chon et al. (1996), high similarity between the SOM
results and the dendrogram (Fig. 2.1) may be observed. The 8 clusters (AI to A VIII )
previously defined with the SOM are those identified on the dendrogram by a
dotted line at distance 3, except for the SU 4 which is grouped with the SUs land
2 on the dendrogram. The U-matrix enhances the SOM and brings more accurate
