143
Clustering and Classification
splits the feature space into three subspaces. For instance, assuming 2-D inputs, a
trained perceptron produces two separating lines. These two lines split the space
into three regions: a positive region (in which all examples belongs to the class under
study), a negative region (in which none of the examples belongs to the class
under study), and a dead zone. One of these separating lines that separates the
positive region from the dead formulates the positive region as follows:
w x + w x + b > T
(7.19)
1 1
2 2
The other one separates the negative region from the dead zone and forms the
negative region as follows:
w x + w x + b <− T
1 1
2 2
(7.20)
Next, we make two observations in the preceding formulations. First, note that the
earlier description of the positive and negative regions further describes the need for
bias. If T = 0, i.e., no bias, all border lines pass through the origin. This significantly
limits the capability of the method to find an optimal separation between classes.
In other words, adding bias to a network effectively increases the capabilities of the
network in separating the classes. The second observation justifies the statement
made earlier about the perceptron. As mentioned previously, a perceptron can create
correct classification only when the classes are linearly separable. From the preceding formulation, it can be seen that in a perceptron the boundary between the classes
is a line (2-D space), plane (three-dimensional [3-D] space), or hyperplane (fourdimensional [4-D] space or higher dimensions). All these boundaries are linear and
therefore cannot separate classes that are only nonlinearly separable. For nonlinearly
separable spaces, as we will see later, networks more sophisticated than a perceptron
are required.
As mentioned earlier, the perceptron algorithm can be simply modified for bipolar,
binary, and real-valued input and output vectors. The following example describes a
problem with binary inputs and binary outputs.
Example 7.6
A typical demonstration example for perceptron is using perceptron to model the
logical “OR” function. Assuming two inputs for the OR function, the output is 1 if
any of the inputs is 1; otherwise, when both inputs are 0, the output of the function
is 0. Augmenting the input space by adding the bias input, the perceptron will have
three inputs. Then, the training set for the perceptron will be as follows. For inputs
(x 1 , x 2 , b) = (1, 1, 1), (x 1 , x 2 , b) = (1, 0, 1), and (x 1 , x 2 , b) = (0, 1, 1), the target value will
be t = 1, and for the input (x 1 , x 2 , b) = (0, 0, 1), the output will be t = 0.
In training the perceptron, we assume α = 1 and θ = 0.2. Table 7.1 shows the training steps for this perceptron. In each step, we change the weights of network based on
the learning rate and the output of the network. Training continues until for a complete
epoch, no weight is changing its value. In last training step, where there is no change
in weights, the final weights are w 1 = 2, w 2 = 2, and w 3 = −1. Therefore, the positive
