9.2. Problems of imbalance in graph data
129
The balanced case
First, we use SSL for a two-class problem, comparing all four approaches. The
number of labelled nodes in each group is varied, increasing gradually. Figure 9.4
shows that the error rates of LGC and TACO are worse than those of LapSVMp and
GBE for the datasets when the number of labelled nodes is small. However, the
difference is negligible when there are 10% or more labelled nodes in each group.
Furthermore, the error percentages tend to be stable after more than 10% labels for
all three approaches.
The LapSVMp approach has a worse performance on the synthetic dataset, but
better performances on the real-world datasets. Our GBE has good performance on
the three datasets.
We also use four equal-sized classes, with the same number of labelled nodes
in each group. Figure 9.5 shows that the performance of our approach GBE is better than others for the synthetic data, but is similar for the two real-world datasets.
Figure 9.5 also shows that the performances of LapSVMp and GBE are slightly better than LGC and TACO for the USPS handwritten data, but worse for the ISOLET
Spoken Letter data. When the number of labelled nodes in each group reaches 30
(10%) or more, the results of the four approaches tend to be consistent and stable.
All of the four approaches perform equally when the class sizes and the numbers of
labelled nodes are balanced. The performances for multiple classes are significantly
worse than for two classes since the connections among the classes become more
complicated.
Imbalance in the number of labelled nodes
Now we consider the case where the class sizes are equal but the number of labelled
nodes per class is not. For the two-class classification problem, we set the number
of labelled nodes in the second class to be 4 times the number in the first class.
Figure 9.6 shows the results of two-class classification in this setting. GBE is now
substantially better than the others, especially for small numbers of labelled nodes.
Its performance remains at the same level as the balanced case.
For the four-class problem, we label the second, third and fourth classes with
2, 3 and 4 times the number of labelled nodes as the first class. Figure 9.7 shows the
results. Our GBE approach again shows better performances for both datasets.
Figure 9.8 gives the boxplots of the distributions of error percentages in the
last column of Figure 9.7. Figure 9.8 shows that the errors of our GBE approach
are statistically significantly lower than the other two. In this case, we have 30, 60,
90 and 120 labelled instances in the first to fourth classes, respectively, more than
10% of the total instances. From the previous experiments, 10% of nodes labelled
is sufficient for a stable result, and this is true for our GBE approach in these cases.
Error percentages for the LGC, LapSVMp and TACO approaches are far higher than
in the previous experiments with the same number of labels.
129
The balanced case
First, we use SSL for a two-class problem, comparing all four approaches. The
number of labelled nodes in each group is varied, increasing gradually. Figure 9.4
shows that the error rates of LGC and TACO are worse than those of LapSVMp and
GBE for the datasets when the number of labelled nodes is small. However, the
difference is negligible when there are 10% or more labelled nodes in each group.
Furthermore, the error percentages tend to be stable after more than 10% labels for
all three approaches.
The LapSVMp approach has a worse performance on the synthetic dataset, but
better performances on the real-world datasets. Our GBE has good performance on
the three datasets.
We also use four equal-sized classes, with the same number of labelled nodes
in each group. Figure 9.5 shows that the performance of our approach GBE is better than others for the synthetic data, but is similar for the two real-world datasets.
Figure 9.5 also shows that the performances of LapSVMp and GBE are slightly better than LGC and TACO for the USPS handwritten data, but worse for the ISOLET
Spoken Letter data. When the number of labelled nodes in each group reaches 30
(10%) or more, the results of the four approaches tend to be consistent and stable.
All of the four approaches perform equally when the class sizes and the numbers of
labelled nodes are balanced. The performances for multiple classes are significantly
worse than for two classes since the connections among the classes become more
complicated.
Imbalance in the number of labelled nodes
Now we consider the case where the class sizes are equal but the number of labelled
nodes per class is not. For the two-class classification problem, we set the number
of labelled nodes in the second class to be 4 times the number in the first class.
Figure 9.6 shows the results of two-class classification in this setting. GBE is now
substantially better than the others, especially for small numbers of labelled nodes.
Its performance remains at the same level as the balanced case.
For the four-class problem, we label the second, third and fourth classes with
2, 3 and 4 times the number of labelled nodes as the first class. Figure 9.7 shows the
results. Our GBE approach again shows better performances for both datasets.
Figure 9.8 gives the boxplots of the distributions of error percentages in the
last column of Figure 9.7. Figure 9.8 shows that the errors of our GBE approach
are statistically significantly lower than the other two. In this case, we have 30, 60,
90 and 120 labelled instances in the first to fourth classes, respectively, more than
10% of the total instances. From the previous experiments, 10% of nodes labelled
is sufficient for a stable result, and this is true for our GBE approach in these cases.
Error percentages for the LGC, LapSVMp and TACO approaches are far higher than
in the previous experiments with the same number of labels.
