9.2. Problems of imbalance in graph data
127
This means we only need to compute a single specific eigenvector of the symmetric
matrix for each class.
9.2 Problems of imbalance in graph data
There are two ways in which data can be imbalanced. First, the number of records in
one class might be much larger than the number in the other class(es). Second, the
number of available labelled records in one class might be larger than the number
in other class(es). And of course both might happen simultaneously. These imbalances cause performance difficulties for models based on some form of spreading
activation.
We have shown that our methods are mathematically well behaved and motivated. Comparing our results with two existing spectral kernel SSL methods: LGC
[118] and LapSVMp [59], and a method based on probability: TACO [73] which was
shown to outperform LP [3], MAD [94] and AM [92], we show the advantages of our
approach.
(a) Toy data with two circles
(b) Toy data with two moons
Figure 9.2: Applying GBE to two toy datasets with only one labelled node (square)
for each class
First, we start with toy data. The nodes of two toy datasets are generated in
two-dimensional space, and adjacency matrices are constructed by connecting each
point to its five nearest neighbors. One node of each class is labelled. Figure 9.2
shows that the results of our GBE approach work as well as the algorithms of [5, 118].
If the structure of the original graph is consistent with the class labels, only
a small repulsive force is needed to model the graph structure. Hence only a small
value for the negative edge weight anw is needed. Both of these examples have such
consistent structure, so only a small value of anw leads to good classification results.
A more substantial synthetic data was created, with 0/1 edge weights. Edges
between two nodes in the same class were created with probability 10% and between
nodes in different classes with probability 2%. Class sizes of 300 were used for
balanced datasets; sizes of 250 and 1000 for two-class datasets; and 250, 500, 750,
and 1000 for four-class datasets.
Précédent

- 148/231

Suivant