310
F. Firouzi et al.
after pooling is
N −F
S + 1, where N is the dimension of input to pooling layer, F is
the dimension of filter, and S is stride.
5.6.5.5 Fully Connected Layer
Similar to a traditional neural network, where all layers are fully connected, this
layer also exists in CNN but stays right before the output layer; their activations can
hence be calculated by a matrix multiplication followed by a bias offset. This is the
last phase for a CNN network.
5.6.5.6 Well-Known CNN Architectures
LeNet, AlexNet, VGG, GoogLeNet, and ResNet are the most well-known publicly
available CNN Architectures. Take the example of famous LeNet-5 architecture
developed by Yann LeCun in 1998. This architecture consists of three convolution
layers, two pooling layers, and one fully connected layer. The architecture is used for
detecting hand-written digital recognition (MNIST) of 28 × 28 pixel images. Note
that zero padding is used to construct the input layer as 32 × 32 pixels, while the
rest of the convolutional layers do not have padding anymore. Activation function
for each layer is tanh, except the output layer. The LeNet demo and description can
be easily found on Yann LeCun’s website.
5.7 Clustering
Clustering is a well-known unsupervised learning technique which allows us to
group (cluster) data points in a way that data points in the same cluster are more
similar to each other than those in different clusters. There are multiple ways of
clustering. In this section, we overview two most popular techniques, namely, Kmeans clustering and hierarchical clustering.
5.7.1 K-Means Clustering
K-means extracts the clusters based on distance, which usually refers to Euclidean
distance in most context. K in K-means represents the number of clusters in which
we want our data to be divided into. There is a restriction in using K-means that the
data shall be in continuous values rather than category values since K-means does
not work on category data in nature. Moreover, it is advised to normalize the data
before applying the K-means method. The reason still resorts to distance calculation.
An example is that if we consider about clustering people based on both weights in
F. Firouzi et al.
after pooling is
N −F
S + 1, where N is the dimension of input to pooling layer, F is
the dimension of filter, and S is stride.
5.6.5.5 Fully Connected Layer
Similar to a traditional neural network, where all layers are fully connected, this
layer also exists in CNN but stays right before the output layer; their activations can
hence be calculated by a matrix multiplication followed by a bias offset. This is the
last phase for a CNN network.
5.6.5.6 Well-Known CNN Architectures
LeNet, AlexNet, VGG, GoogLeNet, and ResNet are the most well-known publicly
available CNN Architectures. Take the example of famous LeNet-5 architecture
developed by Yann LeCun in 1998. This architecture consists of three convolution
layers, two pooling layers, and one fully connected layer. The architecture is used for
detecting hand-written digital recognition (MNIST) of 28 × 28 pixel images. Note
that zero padding is used to construct the input layer as 32 × 32 pixels, while the
rest of the convolutional layers do not have padding anymore. Activation function
for each layer is tanh, except the output layer. The LeNet demo and description can
be easily found on Yann LeCun’s website.
5.7 Clustering
Clustering is a well-known unsupervised learning technique which allows us to
group (cluster) data points in a way that data points in the same cluster are more
similar to each other than those in different clusters. There are multiple ways of
clustering. In this section, we overview two most popular techniques, namely, Kmeans clustering and hierarchical clustering.
5.7.1 K-Means Clustering
K-means extracts the clusters based on distance, which usually refers to Euclidean
distance in most context. K in K-means represents the number of clusters in which
we want our data to be divided into. There is a restriction in using K-means that the
data shall be in continuous values rather than category values since K-means does
not work on category data in nature. Moreover, it is advised to normalize the data
before applying the K-means method. The reason still resorts to distance calculation.
An example is that if we consider about clustering people based on both weights in
