classified events. A human expert has manually classified each one of these events as
true. Additionally, we have a set of other K points with the same structure. This set
will be named as the calibration set. For this set we have for each point a humandefined classification. However, this set (unlike the training set) includes both points
that were classified as true and points that were classified as false. For each point of
the calibration set (called a selected point), one can calculate using the RBF
algorithm, the value of h() based on the set of neighboring points from the training
set. As it was explained previously, points which are located near the selected point
will have high influence on the value of h(), while points located far from the
selected point will have low influence on the h() value of the selected point, or
even zero influence. Based on the calculated result of the h() and the classification
policy of HRL, the selected point will be classified by the algorithm as true or false.
This means that each point in the test set will have two classifications (reference
human classification and algorithmic). The human classification can also be called
the actual (“correct”) classification since the human expert bases it on actual
observation of reality. And the model classification is the classification obtained
by the algorithm for each point in the test set.
After this process is repeated for all points from the calibration set, the results can
be assembled as a confusion matrix as shown in Table 1. Such a matrix has been
introduced, for example, by Cohen’s kappa (1960).
• True Positive – a point that was classified as true by both the algorithm and the
human expert
• False Positive – a point that was classified as true by the algorithm and as false by
the human expert
• False Negative – a point that was classified as false by the algorithm and as true
by the human expert
• True Negative – a point that was classified as false by both the algorithm and the
human expert
Each of the TP, FP, FN, and TN values is a count of the number of events satisfied
in the correlated condition as shown in Table 1 (and in the four points above). Using
these numbers, it is also possible to calculate the sensitivity and specificity of the
results. Sensitivity (also called the TP rate, the recall, or the probability of detection)
measures the proportion of actual positives that are correctly identified out of total
number of true cases – i.e., TP/(TP + FN). Specificity (also called the TN rate)
measures the proportion of actual negatives that are correctly identified out of the
total number of false cases – i.e., TN/(TN + FP). In some cases where the cost of FP
or the cost of FN is extremely high, these ratios are very important.
Table 1 Results of detection
Model classification
Actual classification
True
False
True
True positive (TP)
False negative (FN)
False
False positive (FP)
True negative (TN)
148
E. Brill
true. Additionally, we have a set of other K points with the same structure. This set
will be named as the calibration set. For this set we have for each point a humandefined classification. However, this set (unlike the training set) includes both points
that were classified as true and points that were classified as false. For each point of
the calibration set (called a selected point), one can calculate using the RBF
algorithm, the value of h() based on the set of neighboring points from the training
set. As it was explained previously, points which are located near the selected point
will have high influence on the value of h(), while points located far from the
selected point will have low influence on the h() value of the selected point, or
even zero influence. Based on the calculated result of the h() and the classification
policy of HRL, the selected point will be classified by the algorithm as true or false.
This means that each point in the test set will have two classifications (reference
human classification and algorithmic). The human classification can also be called
the actual (“correct”) classification since the human expert bases it on actual
observation of reality. And the model classification is the classification obtained
by the algorithm for each point in the test set.
After this process is repeated for all points from the calibration set, the results can
be assembled as a confusion matrix as shown in Table 1. Such a matrix has been
introduced, for example, by Cohen’s kappa (1960).
• True Positive – a point that was classified as true by both the algorithm and the
human expert
• False Positive – a point that was classified as true by the algorithm and as false by
the human expert
• False Negative – a point that was classified as false by the algorithm and as true
by the human expert
• True Negative – a point that was classified as false by both the algorithm and the
human expert
Each of the TP, FP, FN, and TN values is a count of the number of events satisfied
in the correlated condition as shown in Table 1 (and in the four points above). Using
these numbers, it is also possible to calculate the sensitivity and specificity of the
results. Sensitivity (also called the TP rate, the recall, or the probability of detection)
measures the proportion of actual positives that are correctly identified out of total
number of true cases – i.e., TP/(TP + FN). Specificity (also called the TN rate)
measures the proportion of actual negatives that are correctly identified out of the
total number of false cases – i.e., TN/(TN + FP). In some cases where the cost of FP
or the cost of FN is extremely high, these ratios are very important.
Table 1 Results of detection
Model classification
Actual classification
True
False
True
True positive (TP)
False negative (FN)
False
False positive (FP)
True negative (TN)
148
E. Brill
